Vision Models Destabilized by Image Presence
The Label That Shouldn’t Have Moved
A vision-language model is handed a sentence and an image, then asked to classify the sentence. The instructions are explicit: decide from the sentence alone. The image is decoration. It should not matter.
It matters.
Researchers built a test called MIST — the Misleading-Image Stress Test — to measure exactly how much it matters. [1] The test comprises 200 English sentences, each containing a phrase that could be read two ways. “The proposal was dead on arrival” could describe a legislative failure or a corpse. “She broke the ice” could mean a social thaw or a frozen lake. Each sentence was paired with one of three visual contexts: an image showing the literal reading, an image showing the opposite reading, or no image at all.
The results, submitted in September 2026, are not what the researchers expected. [1] An aligned image — one that matched the sentence’s intended meaning — shifted 20.5% of the models’ labels. A misleading image — one depicting the opposite sense — shifted 19.4%. The two numbers are nearly identical. The gap between them is 1.1 percentage points.
What moved the judges was not what the image showed. It was that an image was there at all.
This is the first finding, and it is the one that should unsettle anyone who uses these systems to make decisions: the presence of visual context destabilizes a judgment even when the context carries no relevant information. The model does not evaluate the image and decide to ignore it. The model does not evaluate the image at all. It registers that something is present, and that presence alone is enough to change the answer.
What the Models Were Actually Doing
The MIST paper, authored by Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem and Lotem Peled-Cohen, tested thirteen vision-language models. [1] The setup was deliberately simple. Each model received a sentence and one of the three image conditions — aligned, misleading, or absent. The guidelines stated that the label must be decided from the sentence alone. No image should change any answer.
The models changed their answers anyway.
But here is where the result becomes stranger. When the researchers examined which labels moved, they found that only 37% of the differences between aligned and misleading conditions moved toward the sense the image depicted. If the models were genuinely reading the images — even incorrectly — you would expect a much higher proportion of shifts to follow the visual content. Instead, the majority of label changes were directionless. The image acted as noise, not as signal.
The models were not being deceived by the images. They were being destabilized by them.
This distinction matters because it changes what kind of problem we are dealing with. A model that reads a misleading image and draws the wrong conclusion is making an error of interpretation. A model that changes its answer because an image is present, regardless of what the image contains, is exhibiting something closer to a structural flaw. The input modality itself — the mere fact of multimodal presentation — introduces variance.
The Alt-Test and the Limits of Detection
The MIST paper introduces a second layer of analysis that complicates the picture further. The researchers ran what is called an “alt-test” — the alternative annotator test, which asks whether an LLM judge agrees with a panel of human annotators at least as well as a withheld member does, and which the paper cites to prior work. [1] Seven of the thirteen models passed this test. Six did not.
You might expect that the models passing the alt-test would be immune to the image-presence effect. They were not. The effect was smaller in the seven that passed — the paper reports per-model numbers, with the difference present in all of them. Every model tested, regardless of whether it passed the alt-test, showed label changes when an image was added. The alt-test predicted the magnitude of the effect, not its existence.
This finding has implications for how we evaluate AI systems. If a model passes a substitutability test — if it appears to perform as well as a human annotator on a given task — that verdict may describe the testing configuration as much as the model. The MIST paper states this directly: “a substitutability verdict describes a configuration as much as a model.” The configuration includes whether images are present, what kind of images they are, and what instructions accompany them. Change any of these variables, and the model’s behavior changes. The model is not a stable entity that can be evaluated once and then deployed. It is a system that responds to its inputs in ways that are not fully captured by its training objectives or its benchmark scores.
Agreement Without Understanding

Perhaps the most disorienting result in the MIST paper concerns agreement with human annotators. The researchers had human annotators label the same 200 sentences. They then compared the models’ agreement with those human labels across the three image conditions — absent, aligned, and misleading.
Agreement was unchanged.
The models agreed with humans at the same rate whether the image was absent, aligned, or misleading. The image shifted the models’ labels, but it did not shift them in a direction that made them more or less likely to match human judgment. The instability was orthogonal to accuracy.
This is a different kind of problem than the one we usually worry about. When we fret about AI systems being wrong, we imagine errors that degrade performance — mistakes that make the model less useful, less accurate, less trustworthy. But the MIST results describe something else: a model that is equally accurate and equally inaccurate regardless of whether irrelevant context is present, but whose individual answers change. The aggregate statistics look the same. The specific decisions do not.
For a researcher running a benchmark, this might not matter. The average score is stable. For a practitioner using the model to classify a single sentence — to make a decision about a single case — it matters enormously. The model’s answer depends on whether an image was included, even though the image should not affect the answer, even though the instructions say to ignore it, even though the image’s content is irrelevant to the task.
The model is not failing in a way that shows up in the numbers. It is failing in a way that shows up in the individual case.
The Instruction That Didn’t Help
The MIST paper includes a control condition that is easy to overlook but difficult to forget. The researchers ran the task with the “ignore the image” instruction removed, but with the image still present. In that condition, 11.6% of labels changed.
Compare that to the 20.5% shift caused by an aligned image when the instruction was present. The instruction to ignore the image did not reduce the image’s effect. It increased it. The models were more stable when they were not told to ignore the image than when they were.
This inverts the usual logic of prompt engineering. We assume that explicit instructions constrain model behavior — that telling a model what to do makes it more likely to do that thing. The MIST results suggest that in this case, the instruction to ignore the image created a kind of tension that made the image more disruptive. The model was told to disregard something that was present in its input, and that tension manifested as instability.
The paper does not offer a definitive explanation for this effect. It reports the number and moves on. But the number itself is a finding: instructions do not always work the way we expect them to. A rule that says “ignore X” may make X more salient, not less.
What This Means for Substitutability
The MIST paper is framed around a concept called substitutability — the idea that a model can replace a human annotator on a given task. If a model performs as well as a human on a benchmark, the argument goes, it can be used in place of a human in production. The model is substitutable.
The MIST results undermine this logic. If a model’s labels change by 20% when an irrelevant image is added, and if that change is not reflected in agreement with human annotators, then the model’s substitutability is not a property of the model alone. It is a property of the model-plus-configuration. Change the configuration — add an image, remove an image, change the instruction — and the substitutability may change with it.
The paper’s authors write that “what moves a judge is that an image is there, not which of the two it is.” The judge is not evaluating the image. The judge is responding to the presence of the image. The presence is the variable. The content is not.
This has implications beyond the specific task of sentence classification. Vision-language models are increasingly used in settings where images and text are both present — medical diagnosis, content moderation, document analysis, customer service. In each of these settings, the model receives multimodal input. In each of these settings, the presence of an image may introduce instability that the model’s benchmark scores do not capture.
The MIST paper does not claim that all vision-language models are unreliable. It claims that the reliability of these models is configuration-dependent in ways that current evaluation methods do not measure. A model that passes a benchmark may fail in production, not because the benchmark was wrong, but because the benchmark did not test the configuration that production uses.
The Asymmetry Between Testing and Deployment
There is a structural asymmetry in how AI systems are evaluated and how they are used. Evaluation happens in controlled conditions: fixed inputs, fixed instructions, fixed metrics. Deployment happens in uncontrolled conditions: variable inputs, variable instructions, variable contexts. The MIST paper shows that small changes in context — the presence or absence of an image — can produce large changes in output.
The asymmetry is not accidental. It is built into the way we measure model performance. A benchmark is designed to be reproducible. It fixes the conditions so that different models can be compared. But fixing the conditions means that the benchmark cannot measure how the model behaves when the conditions change. The benchmark measures the model in a specific configuration. It does not measure the model’s sensitivity to configuration.

The MIST paper introduces a method for measuring that sensitivity. The method is simple: vary the context, hold the task constant, and measure how much the output changes. The results show that the sensitivity is large — around 20% label change — and that it is not captured by agreement with human annotators.
This suggests that current evaluation practices may be systematically underestimating the instability of vision-language models. A model that scores well on a benchmark may be unstable in ways the benchmark does not detect. The benchmark measures accuracy. It does not measure robustness to irrelevant context.
The Institutional Consequence
The MIST paper does not make policy recommendations. It reports results and describes a method. But the results point toward a consequence the paper states only in passing: if substitutability is configuration-dependent, then the organizations that use AI models in place of human annotators are relying on a property that may not hold in the configurations they actually use.
The paper’s authors write that “a substitutability verdict describes a configuration as much as a model.” This is a statement about measurement. But it is also a statement about responsibility. If the verdict describes a configuration, then the configuration is part of what is being evaluated. The configuration is part of what must be controlled. The configuration is part of what must be reported.
Current practice does not report configuration in this way. A model is described by its name, its version, its training data, its benchmark scores. The conditions under which those scores were obtained — the instructions, the presence or absence of images, the specific task format — are often treated as implementation details. The MIST results suggest they are not details. They are variables that affect the outcome.
The institutional consequence is that substitutability claims may need to be qualified by the configuration in which they were tested. A model that is substitutable in one configuration may not be substitutable in another. The claim “this model can replace a human annotator” is incomplete without the specification of the conditions under which the replacement was validated.
The MIST paper suggests that AI evaluation should follow the same logic: report the configuration, not just the score.
What the Image Shows
The MIST paper’s title is “It’s Not What the Image Shows.” The image’s content is not what moves the model. The image’s presence is.
This is a specific result about a specific task, but it resonates with a broader pattern in how AI systems behave. The systems respond to features of their input that are not the features we intend them to respond to. They are sensitive to context in ways that are not captured by their training objectives. They are unstable in ways that are not captured by their benchmark scores.
The MIST paper does not explain why this happens. It does not propose a mechanism. It reports the effect and describes a method for measuring it. The explanation — whatever it is — is left for future work.
But the effect itself is a finding. It is a finding about vision-language models, about evaluation methodology, and about the gap between what these systems appear to do and what they actually do. The systems appear to read images and classify sentences. What they actually do is respond to the presence of images in ways that are not fully understood, not fully measured, and not fully controlled.
The image is there. The model changes its answer. The content of the image does not matter. The presence does.
Sources
Mentioned organisations (context, not sources)
DeepThink Search AI-generated, for reference only
