When Forgetting Bounds Hold: Only Where R < a
A prediction that scored 0.974–0.998 and still got the future wrong
Fine-tuning a language model on new material is supposed to be the cheap path to specialization. You take a model that already knows a great deal, show it a narrow set of examples, and expect it to absorb the new skill without surrendering the old ones. The trouble is that it often does surrender them. Ask the model a question it once handled easily and it stumbles, or answers with the wrong shape entirely. The industry calls this catastrophic forgetting, and treats it as a kind of amnesia: the weights have been overwritten, the knowledge is gone.
A paper posted to arXiv re-examines that story. [1] The work asks a narrower, more testable question — whether measurements taken before a fine-tuning run can put a bound on the probability that the run makes the model lose any particular fact you wanted protected. Not whether forgetting happens, but whether its likelihood can be capped in advance, fact by fact, using only what you can observe at the starting line.
The first-order response model, estimated by finite-difference probes, predicts changes of per-fact margins with a correlation between 0.974 and 0.998. That is a strikingly tight fit. It is also, on its own, useless for the question that matters, and the paper says so plainly: predictions of forgetting built on this model failed, because forgetting requires parameter changes far outside the region in which the model was validated — roughly an 18-fold larger shift than the validated range.
The margin, the boundary, and the distance between them
To understand what the researchers actually managed to bound, you have to know what a margin is in this setting. For each fact you want the model to keep, you can define a quantity that measures how confidently the model prefers the correct answer over the alternatives. Fine-tuning nudges that quantity. If it drifts far enough in the wrong direction, the margin crosses zero and the model starts producing the wrong answer. Forgetting, in this framing, is not a memory being erased. It may be a margin crossing a boundary.
The bounds contain a term the authors call R, which measures how much the response coefficients themselves change during the run. This is the quantity that separates the two versions of the result. The simplified Freedman bound sets R equal to zero, meaning it assumes the model’s sensitivity to each fact stays put while training proceeds. And every single fact on which it was violated had R greater than or equal to a, where a is the distance of that fact’s margin to the boundary.
The pattern is not subtle. When the response coefficients moved by more than the fact’s own safety margin, the simplified bound broke. The complete Freedman bound, which keeps R in the calculation, certifies only facts where R is less than a.
What the paper delivers, then, is a certificate with a price attached. You can bound the probability that fine-tuning destroys a given fact, but only for the facts whose response coefficients change by less than their distance to the boundary. For everything else, the math goes quiet. It does not tell you the fact is doomed. It tells you that your instrument cannot see that far.

What the model appears to do versus what it is doing
There is a version of this story that stays inside machine learning and never leaves. A method has limits, the limits are characterized, the characterization is tested. Useful, unremarkable. But the shape of the finding is familiar from somewhere else, and the familiarity is the point.
Consider what the surrogate model was doing during the validation phase. It was predicting changes in per-fact margins with a correlation that would pass any reasonable sanity check. If you had stopped there — if you had taken the 0.998 as evidence that you understood the fine-tuning process — you would have walked into the next run confident and wrong. The model was not lying. It was answering a question adjacent to the one you asked. It had learned the local weather and you were asking about the climate.
That is a specific and uncomfortable kind of gap. It is not the gap between a system that works and one that fails. It is the gap between a system that works in the conditions you tested and a system whose behavior in the conditions you care about is governed by terms your test never exercised. The R term is the formal name for those terms. In plain language, R is the amount by which the model’s own sensitivities shift underneath you as training proceeds. Set it to zero and your predictions are clean. Keep it and they are honest but narrower.
The paper’s own framing makes the distinction explicit. The first-order model was validated in one region and asked to predict in another. The failure was not in the fit. The failure was in the assumption that the fit travels.
The counter-perspective that does not reverse the picture
It would be easy to read the paper as a cautionary tale about overconfidence in surrogate models, and stop there. But that reading misses what the researchers actually built. They did not conclude that first-order models of fine-tuning are useless. They concluded that first-order models can bound forgetting — on the facts whose response coefficients change by less than their distance to the boundary. The certificate is real for the linear surrogate model (held in all tested conditions). It is just conditional.
The condition is the interesting part. R greater than or equal to a is necessary, not sufficient for the simplified bound to break. It describes any fact whose safety margin is thin relative to the drift in the model’s sensitivity. In a fine-tuning run on a narrow task, the facts most at risk are exactly the ones sitting close to the boundary, and those are the ones whose R will most easily exceed their a. The bound protects the facts that were never in much danger and falls silent on the ones you were worried about.
This is not a flaw in the derivation. It is a property of the situation. The paper’s two preregistered confirmatory studies, run across 43 new conditions, weakly support this: the complete bound held in all of them, and the simplified bound failed on only 3 facts, each with R greater than or equal to a. [1] In the first study, 2 of 2 violated facts had R ≥ a; the second study remained inconclusive (minimum 5 violations needed for a decision). What makes the result worth sitting with is that the stopping point is not arbitrary. It is defined by a quantity you can measure before the run. You can compute R and a for each protected fact and know in advance which ones will receive a certificate and which ones will not. That is a different kind of knowledge than “the model will forget some things.” It is a map of your own ignorance, drawn before you start.
Where the forgetting actually lives

That distinction matters for what you do about it. If forgetting is erasure, your options are limited and expensive. If forgetting is a margin that drifted, the question becomes whether the drift is reversible, whether a small correction can push the margin back across the boundary, whether the R term that broke your bound is itself something you can influence by changing how you fine-tune.
The paper does not answer those questions. It establishes the bound and characterizes when it holds. But the framing opens a door that the erasure story keeps shut. A model that has “forgotten” a fact may be a model whose response coefficients moved too far for your instrument to track, not a model that no longer contains the fact at all. The boundary is a property of the measurement as much as of the model.
There is a further implication in the R term itself. R measures how much the response coefficients change during the run. If R is large, the model’s sensitivities are shifting, which means the fine-tuning process is not just adjusting margins — it is changing the relationship between the model and the facts. The simplified bound assumes that relationship is stable. The complete bound admits it is not. The gap between the two is the gap between a world where you can predict forgetting from a single snapshot and a world where the act of fine-tuning rewrites the terms of its own prediction.
Two distinct failure modes are at play. (1) The surrogate model fails to predict forgetting because forgetting requires ~18× larger parameter shifts than the validated region. (2) The simplified Freedman bound fails where R ≥ a.The image that summarizes why this is harder than it looks
Picture the validation phase of the experiment. The surrogate model is being checked against the real model, run after run, and the correlation is climbing toward 0.998. Everything looks aligned. The surrogate has learned the local behavior of fine-tuning, and the local behavior is smooth, predictable, well-behaved. You could build a product on this. You could ship a forgetting predictor and feel good about it.
Now picture the next run. The surrogate is asked to predict which facts will be lost, and it fails. Not because the correlation dropped — the correlation was never the problem. It fails in part because the facts that get lost are the ones whose response coefficients moved further than the validation region ever went, and in part because for other facts the simplified bound expires where R grows past a. The surrogate was never wrong about what it saw. It was blind to what it did not.
That is the shape of the problem the paper leaves you with. A system can be accurate, validated, and still unable to answer the question you brought it. The 0.998 is not a lie. It is a true statement about a region, and the forgetting happens outside that region, in the territory where the certificate expires. The model that appears to understand fine-tuning understands the part of fine-tuning that behaves. The part that forgets is the part that moved.
The paper’s contribution is not a warning about AI deception in the anthropomorphic sense. It is something more precise and more useful: a demonstration that the gap between what a model appears to do and what it is doing can be located, measured, and bounded — and that the location of the gap is itself predictable, if you are willing to compute the terms that tell you where your predictions stop working. The bound is real (for the linear surrogate). So is the silence beyond it. And the silence is not a failure of the method. It is the method telling you, in advance, exactly where it will go quiet.
Scope
The results apply to LoRA fine-tuning with SGD on models from 0.6B to 14B parameters, 40 or 80 steps, 64 protected facts per run. Under Adam the linear response breaks down after roughly ten steps; Adam and full fine-tuning are explicitly noted as open questions by the authors.
