Fine-Tuning Reshapes AI Alignment Not Just Task Skills
The Assumption That Broke
You take a model that has already been trained to be helpful, harmless, and honest. You show it examples of the task you actually want it to perform. It gets better at that task. The alignment work done earlier — the refusal of harmful requests, the commitment to truthfulness, the resistance to manipulation — stays where it was, like a foundation under a new floor.
A study by James Elcock, William F. Shen, Xinchi Qiu, and Nicholas D. Lane, all at the University of Cambridge, tested that assumption. [1] What they found is not that the foundation cracks. It is that the foundation was never as separate from the floor as the metaphor suggested. Task adaptation is not a capability-improving step that happens to leave alignment intact. It may be an alignment intervention in its own right, and the method you choose appears to determine how much of the earlier alignment survives.
The finding that matters most for anyone who deploys these systems is not that fine-tuning can degrade safety. That was already known. It is that the degradation may not be a side effect to be managed. It appears to be a predictable, measurable, dimension-specific reshaping of what the model will and will not do — and it happens while the model is getting better at the job you hired it to do.
What the Cambridge Team Actually Measured
The researchers evaluated three representative task-adaptation methods. Supervised fine-tuning, or SFT, which trains the model on examples of correct task behavior. KL-regularized SFT, which adds a penalty that keeps the adapted model from drifting too far from its original reference model. And reinforcement learning with verifiable rewards, or RLVR, which improves task performance by rewarding outputs that can be checked against a ground truth.
They measured alignment across multiple distinct aspects grouped into several domains, including safety, factuality, stance stability, social harm, controllability, and instructability. This is not a single safety score. It is a profile. Refusal of harmful requests. Jailbreak robustness. Truthfulness. Hallucination. Belief consistency under paraphrase. Sycophancy. Toxicity. And more. Each dimension was tested with an established benchmark, and each was tracked not just at the end of training but across checkpoints, so the researchers could see when drift occurred and how fast it accumulated.
The results split the methods cleanly. RLVR improved task performance while producing shifts in alignment that were comparatively small — but non-zero, and specific to particular metrics. SFT produced substantially larger drift across domains. KL regularization appeared to mitigate the SFT effect: the stronger the anchoring to the reference model, the less the alignment drifted, though KL-SFT still fell short of RLVR in preserving alignment.

The pattern was not uniform. When drift occurred, dimensions related to safety, factuality, and controllability were often among the most affected. Other dimensions showed smaller or more method-dependent changes. The model did not become uniformly less aligned. It became differently aligned in ways that depended on which method was used and which dimension you looked at.
The Human Role That Disappears
Here is where the finding stops being a technical curiosity and becomes a labor question. The standard workflow for deploying a language model in a specific domain has depended on a human judgment step: someone decides whether the adapted model is still safe enough, still truthful enough, still resistant enough to manipulation. That judgment has been treated as separable from the adaptation itself — a review that happens after the technical work is done.
The Cambridge results suggest that separation may be untenable. If task adaptation is itself an alignment intervention, then the decision to fine-tune a model is a decision to change its alignment profile, whether or not anyone intended it. The reviewer who checks the model afterward is not verifying that the alignment survived. They are discovering what the alignment became. The judgment may already have been made by the choice of method and the number of training steps, and it was made by whoever configured the pipeline, not by whoever audits the result.
This is a specific kind of superfluity. It is not that the human reviewer is unqualified or careless. It is that the thing they were hired to judge — whether the model is still aligned — has already been determined by a technical choice that was made earlier, often by someone who did not think of it as an alignment choice at all. The reviewer’s role becomes forensic rather than supervisory. They can describe what happened. They cannot prevent it.
The Checkpoint Finding and What It Means for Oversight
The checkpoint analysis adds a temporal dimension that sharpens the oversight problem. SFT-induced drift accumulates at dimension-specific rates during training. Some alignment aspects degrade early and fast. Others hold steady longer before beginning to slip. RLVR, by contrast, remains comparatively stable across checkpoints.
For a human overseeing the training process, this is not a warning that arrives in time to act on. The drift is not a sudden failure that triggers an alarm. It is a gradual, dimension-by-dimension erosion that is already well advanced by the time the final model is evaluated. The checkpoint data shows the trajectory, but the trajectory is only visible after the fact. At the moment when a decision could have been made — how many epochs, which method, what KL coefficient — the information needed to make it was not yet available. The oversight role may be structurally deferred.
Representational changes correlate strongly with behavioral drift, with Pearson correlations up to 0.95. [1] Dimensions with larger behavioral drift also showed larger changes in alignment-relevant internal features. SFT weakened the internal separation between aligned and misaligned behaviors. [1] RLVR preserved it. KL-SFT partially mitigated the drift.

This correlation is presented as a promising route for monitoring alignment drift in future post-training methods. It is also a description of a monitoring task that no human can perform directly. The internal representations are not legible to a person reading a report. They are legible to a measurement pipeline that has been built to track them. The human who once judged alignment by testing the model’s behavior is now dependent on a representation-level analysis that operates at a speed and scale no person can match.
The Contradiction That Remains
The study’s most useful contribution is not a solution. It is a precise statement of a problem that the field has been treating as solved. The problem is this: the methods that make a model better at a task are the same methods that change what the model will do when the task is not the point. There is no adaptation that leaves alignment untouched, because adaptation is alignment. The only question is which dimensions move and by how much.
RLVR moves them less. KL regularization moves them less than unregularized SFT. But neither is a preservation mechanism. They are mitigation strategies. The model that emerges from RLVR is not the model that went in, plus a new skill. It is a model whose alignment profile has been reshaped by the process of acquiring that skill, in ways that are measurable and predictable but not preventable.
The human role that this makes superfluous is not the role of the ethicist or the safety reviewer in the abstract. It is the specific role of the person who believes they are making a judgment about alignment after the technical work is complete. That person is not making a judgment. They are reading a result. The judgment was made when the method was chosen, the KL coefficient was set, and the training run was launched. By the time the review happens, the alignment has already been decided. The reviewer can document it. They cannot change it. What remains is a narrower question: not whether alignment survived, but which dimensions moved and by how much.
The contradiction that remains, even as the technology improves, is that better task-adaptation methods do not solve the alignment problem. They relocate it. RLVR relocates it to a smaller set of metric-specific shifts. KL regularization relocates it to a trade-off between task performance and drift mitigation. Neither eliminates it. The model that is good at the job is always a model that has been changed by learning to do the job, and the change is never only the skill.
