When Forecasts Know You or Just Your Disease
The Boundary Where Judgment Meets the Machine
A clinician sits down with a forecast. The model predicts the next imaging state. The clinician looks at the same scans the model saw, and has to decide whether to believe the number on the screen or the tissue in front of her. That moment — the moment where a human decides whether to trust the recommendation — is the moment Patient, Place, Prior was built to interrogate.
The paper does not ask whether the model is accurate. It asks something harder: when the model produces a patient-specific forecast, is the patient actually in the prediction, or is the patient just along for the ride?
And it is the question that determines whether the radiologist’s judgment is being augmented or quietly replaced by a number that only looks personal.
Accuracy Does Not Tell You Who the Model Is Listening To
The authors begin with a problem that has haunted machine learning in medicine since the first model claimed to predict outcomes better than a clinician. [1] Two models can achieve identical held-out performance while relying on entirely different information. One might be reading the patient’s tumor trajectory. The other might be reading the scanner brand, the acquisition protocol, or the fact that the patient was scanned at a particular site on a particular day.
These three sources have completely different scientific meanings. Aggregate error collapses them into one indistinguishable number.
So the first thing the paper does is refuse to accept accuracy as evidence of personalization. A model that predicts the future well might be doing so because it has learned the patient, or because it has learned the population and the patient happens to fit. The difference matters because only one of those is a patient-specific world model. The other is a cohort-level pattern applied to an individual — useful, perhaps, but not the thing it claims to be.
Three Questions That Separate Use From Value
The framework the authors propose — Patient, Place, Prior — is an audit, not a metric. It does not produce a single number. It produces three comparisons, each of which isolates one source of information while holding everything else fixed.
For a given patient, the Patient comparison asks whether the forecast is better when the model receives that patient’s own longitudinal imaging history rather than another patient’s history, under the same lesion mask. The Place comparison asks whether the forecast is better when the model receives the patient’s own lesion occupancy maps rather than another patient’s maps, while preserving the patient’s imaging history. The Prior comparison asks whether the forecast from the patient’s history beats a population-average prediction constructed under the same lesion occupancy maps and observed context.
Every comparison keeps the patient’s future target and evaluation procedure identical. Only the source of one input changes. This is the design that makes the audit meaningful: it does not ask whether the model is accurate, it asks which inputs the accuracy depends on.
The distinction the authors draw is between patient-conditioned and patient-specific. A model is patient-conditioned if it receives the patient’s history. That is an input property, nothing more. A model has patient-specific predictive value only if the patient’s history improves the forecast beyond what a population-average prediction achieves under matched support and context.
A Model Built to Be Audited
To test the framework, the authors built Cancer JEPA — a one-step model that forecasts frozen representations of future breast DCE-MRI examinations during neoadjuvant therapy. The architecture is deliberately factored. A patient-conditioned reduced-rank regression baseline provides a low-complexity anchor. On top of that, a neural correction is trained with an occlusion-based latent objective.
This factorization is not an architectural flourish. It is what makes the post-hoc audit possible. Because the correction is constrained to the supplied lesion region and because history and spatial support enter separately, the authors can swap one input at a time and observe what changes.
The frozen encoder provides the latent target. The model predicts future representations, not future pixels. That caution matters for the interpretation: a forecast of an encoded future examination is an observational forecast, not a causal claim about treatment response.
What the Audit Found

In a validation cohort previously used in development, the results split along the line the framework was designed to draw. [1] Forecast error is lower when the neural correction receives the patient’s own history rather than another patient’s history. Forecast error is also lower when the correction receives patient-matched lesion occupancy maps rather than substituted maps. Both the Patient and Place comparisons point in the expected direction: the model does use the patient’s information, and it does use the supplied spatial support.
But the Prior comparison — the one that asks whether the patient’s history adds value beyond a population-average prediction under matched support and context — does not resolve. The descriptive 95% interval comparing the correction computed from patient history with the population-average neural correction includes zero. [1] The advantage of using the patient’s own history over a population-level pattern is uncertain.
This is the finding that gives the paper its edge. The model uses the patient’s history. It benefits from patient-matched spatial support. But the evidence that it achieves patient-specific predictive value beyond a cohort-level pattern is not there. The audit separates input use from value, and the separation is what the audit is designed to detect.
The Radiologist’s Decision, Revisited
Go back to the radiologist. The model has produced a forecast. The audit says the forecast uses the patient’s history — swapping in another patient’s scans makes it worse. The audit also says the forecast uses the supplied lesion occupancy maps — swapping in another patient’s maps makes it worse. But the audit cannot say whether the forecast is better than what a population-average model would have produced under the same conditions.
What does the radiologist do with that?
She knows the model is reading the patient. She does not know whether reading the patient is doing anything beyond what reading the cohort would have done. The number on the screen is patient-conditioned. Whether it is patient-specific is an open question.
This is the boundary the paper draws, and it is a boundary that most clinical AI does not acknowledge. The standard evaluation reports accuracy. Accuracy is high. The model is deployed. The question of whether the accuracy comes from the patient or from the population is rarely asked. The radiologist’s judgment is asked to trust a forecast whose personalization has not been established.
The audit does not tell the radiologist to distrust the model. It tells her what the model has and has not demonstrated. That is a different kind of information than accuracy, and it is the kind that a clinician deciding whether to override a recommendation actually needs.
Why Swapping Inputs Is Not Enough
There is a temptation to think that saliency maps or feature attribution could answer the same question. The paper addresses this directly. Saliency methods characterize local sensitivity — how much the output changes when an input changes slightly. They do not tell you whether the model would perform as well with a different patient’s input. And replacing prior scans and lesion masks together confounds their contributions: you cannot tell whether the drop in performance came from losing the history or losing the spatial support.
The P3 design avoids both problems. It swaps one input at a time, holds the target fixed, and compares against a matched population-average baseline. The comparisons are not about sensitivity. They are about necessity and sufficiency: does the forecast need this input, and does this input do more than a cohort-level substitute would?
That is a harder experimental design than attribution. It requires a model whose inputs can be recombined without breaking the prediction target. Cancer JEPA was built to satisfy that requirement. Most clinical forecasting models are not.
The Population Average Wearing a Patient’s Name
The Prior comparison is the one that stings. A model can be patient-conditioned — it receives the patient’s history, it uses the patient’s history, swapping the history degrades performance — and still not beat a population-average prediction under matched support and context. The patient’s information is being used, but the model is not extracting more from it than it would from the cohort-level pattern.
This is the quiet failure mode the paper names. The forecast looks personal. The input is personal. The performance is good. But the personalization may be decorative rather than functional. The model may be learning a trajectory that is common across patients and applying it to this patient because this patient happens to fit. The patient’s history is in the input, but the predictive signal is in the population.
A clinician reading the forecast has no way to tell the difference. The number is the same. The confidence interval is the same. The only way to know is to run the audit, and the audit is rarely run.
What the Audit Does Not Claim

The authors are careful about the limits of their own framework. P3 applies when history and support can be recombined without changing the prediction target or the remaining observed context. It is a post-hoc audit of a frozen model, not a training procedure. It does not identify causal effects. The forecast is an observational prediction of an encoded future examination, not a causal response to an intervention.
It claims that the factorization permits the audit, and that the audit produces a result that accuracy alone would have hidden.
That modesty is part of the argument. The framework is not a solution to the problem of medical AI personalization. It is a way of asking whether the problem has been solved in a particular case. The answer, in this case, is: partially. The model uses the patient. Whether it uses the patient better than it uses the population is not established.
The Skill That Gets Replaced
The radiologist’s skill is not reading scans. It is knowing when to trust a number and when to override it. That skill depends on understanding what the number represents. If the number is a patient-specific forecast, it deserves more weight. If it is a population average with the patient’s name attached, it deserves less.
The audit is a tool for making that distinction. Without it, the radiologist is asked to exercise judgment about a model whose personalization has not been demonstrated. She is asked to trust the patient-specificity of a forecast on the basis of accuracy, and accuracy does not carry that information.
The paper’s contribution is not a better model. It is a way of showing that the model’s claim to patient-specificity is weaker than its accuracy suggests. The radiologist who reads the audit knows something the accuracy number hides. The radiologist who does not is making a judgment call on incomplete evidence.
Where the Boundary Moves
The boundary between human judgment and machine recommendation is not fixed. It moves as models improve and as evaluation methods improve. The P3 framework moves it in one direction: it gives the clinician a reason to withhold trust that accuracy alone would not provide.
In this cohort, the interval includes zero. In another, it might not. The framework is designed to be run again, on other models, in other settings, with other targets.
What it establishes is that the question can be asked. A forecast can be audited for patient-specific value beyond a population-level pattern. The answer may be yes or no or uncertain. But the question is now on the table, and a model that cannot answer it is a model whose personalization claim is unverified.
The Goal We Are Pursuing
The paper’s final move is the one that matters most. It does not ask whether Cancer JEPA is a good model. It asks what evidence would justify calling a medical forecast patient-specific. The answer it gives is a standard: the patient’s history must improve the forecast beyond a population-average prediction under matched support and context. Input use is not enough. Accuracy is not enough. The patient must be doing work that the population cannot do.
That standard is higher than the one most clinical AI is held to. It is also the standard that the field needs if “personalized medicine” is going to mean something other than “the model received the patient’s data.” The paper does not claim to have met the standard. It claims to have built a way of testing whether the standard is met.
The clinician at the start of this article is still sitting in front of the screen. The forecast is still there. The difference is that now there is a way to ask whether the forecast is about her patient or about the population her patient belongs to. The answer may change what she does. The question changes what she can know.
And that is what the audit makes possible: not a better prediction, but a clearer boundary between the prediction that knows the patient and the prediction that only knows the disease.
