🌿freegardner

Synapse

Silent Security Failures Start Before Vulnerable Code

10 Oct 2026 · via Rss.arxiv

Silent Security Failures Start Before Vulnerable Code
Image: Eric Jones / Wikimedia Commons (CC BY-SA 2.0)

Silent Security Failures Start Before Vulnerable Code

When Success Signals Lie

A software repair agent finishes its work. The tests run green. The syntax checks pass. The patch slots cleanly into the codebase, and by every measure the system offers, the job is done. And yet the code it just wrote, or the code it just left untouched, still carries a security vulnerability.

This is the phenomenon researchers now call a silent failure: a patch that satisfies every available functional check while retaining or introducing a vulnerability. The repair looks complete. It is not. The gap between what the agent appears to have accomplished and what it actually accomplished is a central problem, and until recently nobody had a systematic way to find where in the process that gap first opened.

A study introduces a method called Security Awareness Gap Evaluation, or SAGE, designed to answer a deceptively simple question: at which turn in an agent’s execution did the repair first diverge from the security intent of the task? [1] The work is described in a paper available on arXiv (arXiv:2610.06163), and its findings suggest that the moments we blame for security failures are usually not the moments where they begin.

The Difference Between the Write and the Wrong Turn

When a vulnerable line of code appears in a file, the instinct is to treat that line as the culprit. Find the commit, find the diff, find the write that put it there. SAGE’s central insight is that this instinct is frequently wrong.

The method distinguishes between the origin of a silent failure and the introducing write. The origin is the earliest turn at which the recorded reasoning or the artefact history shows a divergence from the task’s security intent. The introducing write is merely the moment the vulnerable code lands in the file. These two events can be separated by several turns, and in some cases the introducing write does not exist at all.

The paper describes two illustrative cases. In the first, the planner agent selects AES in ECB mode at turn 1, two turns before the vulnerable code is written. A diff-based analysis would flag the write. Only the trace reveals that the inadequate defence was chosen earlier, when the agent reasoned about its approach. In the second case, the vulnerable code is inherited from the input file and never touched. There is no introducing write to locate. The trace instead shows an initial plan that omitted the relevant security requirement entirely and a later review that never revisited the omission.

In both scenarios, examining only the final patch is insufficient. The failure has a history, and that history lives in the agent’s recorded reasoning.

What the Numbers Show

The SAGE evaluation drew on 3,684 valid repair traces generated by six open-source agent frameworks and six base models, working on tasks from two established security datasets, SecurityEval and CVEfixes [1]. From these traces, the researchers identified 95 confirmed silent failures — cases where the patch passed syntactic and functional checks but was independently confirmed to contain a security vulnerability [1]

Silent Security Failures Start Before Vulnerable Code (Image 1)
AI-generated image

SAGE localized an origin in 93 of those 95 cases, a rate of 97.89 percent [1]. The distribution of origins is the most revealing result. Most identified origins involved an unaddressed security requirement or an inadequate defence choice. Only five of the 93 origins coincided with the code change itself.

Put differently: in the overwhelming majority of silent failures, the moment the agent wrote vulnerable code was not the moment the failure began. The failure began earlier, in a plan that never mentioned the requirement, or in a defence strategy that could not close the weakness.

Among the 19 cases where the agent did introduce vulnerable code, the origin preceded the introducing write in 14 [1]. The write was the symptom. The cause lay upstream.

Why Existing Tools Miss This

Failure attribution in agentic systems has become a field of its own. Taxonomies characterize agent failures by system design, inter-agent misalignment, and task verification, and attribution methods identify responsible agents or steps under different evidence conditions.

Each of these approaches assumes something that silent failures violate: that there is an observable task failure to attribute. They rely on failed runs, on step-level labels, on the signal that something went wrong. A silent failure provides no such signal. The patch passed. The tests are green. There is nothing to attribute because, by the system’s own measures, nothing failed.

Earlier work established that security deficiencies can hide behind passing patches. SAGE addresses the next diagnostic question: once you know a silent failure exists, where in the recorded trajectory did the process first go wrong? This requires separating the origin turn from code provenance, from the introducing write, and from later missed detection opportunities — a distinction that earlier categorization did not provide.

The Consistency Question

Localizing an origin is only useful if the localization holds up. The SAGE evaluation tested reliability through repeated scoring and a second judge model.

The results are nuanced. Repeated scoring and the second judge reproduced the same origin type more consistently than the exact turn. In other words, the method reliably identifies what kind of failure occurred — an unaddressed requirement, an inadequate defence — but is less consistent about pinpointing the precise turn where it happened.

Agreement with human annotators on code provenance and the introducing write ranged from moderate to substantial, with kappa values between 0.58 and 0.74, varying across task subsets [1]. Exact-turn agreement was lowest for the small group of traces that preserved only the final file.

That last finding carries a practical implication. When an agent’s execution trace is incomplete — when only the final artefact survives — localization becomes harder and less reliable. The reasoning history is not a nice-to-have. It is the evidence.

Silent Security Failures Start Before Vulnerable Code (Image 2)
AI-generated image

The Gap Between Appearance and Reality

The silent failure problem is a specific instance of a broader pattern in AI systems: the divergence between what a system appears to have done and what it has actually done. A passing test suite is a claim about functional correctness. It is not a claim about security. When an agent treats the two as equivalent, it is not lying in any intentional sense, but it is producing a misleading signal — a patch that looks done and is not.

The SAGE findings indicate that this misleading signal has a structure. The failure does not begin at the write. It begins at the plan that omitted a requirement, or the defence that was chosen without adequate reasoning, or the review that failed to revisit an earlier gap. These are moments of reasoning, not moments of code production. They are visible in the trace if anyone looks, but they produce no error, no exception, no red test.

The practical consequence is that adding more tests may not solve this. The failure is not a test failure. It is a reasoning failure that produces code which passes tests. The remedy has to operate at the level where the failure originates — in the planning, the defence selection, the review process — not at the level where it becomes visible. Whether safeguards placed at the identified origins actually prevent silent failures remains to be tested in controlled studies.

What Comes Next

The SAGE paper does not claim to solve silent failures. It claims to localize them, and it does so with a method that works on a single trace, without requiring repeated executions or labelled failure steps. That is a meaningful step because it makes the problem tractable within the constraints that real audits face: one execution, one trace, one patch that passed.

The boundary conditions the researchers identify point toward the next set of questions. Reliability varies across silent failure categories, agent frameworks, and base models. Exact-turn agreement drops when artefact availability is limited. The origin type is more reproducible than the origin turn. Each of these findings suggests where additional safeguards might be inserted — and, just as importantly, where they might not help.

The deeper implication is about how we evaluate AI systems that write code. If the only signal we collect is whether the tests pass, we will continue to miss the failures that matter most: the ones that pass. The trace — the record of what the agent reasoned, planned, and chose — is where the evidence lives. The patch tells you what was built. The trace tells you where the process first went wrong.


Sources

  1. arXiv — Paper

Mentioned organisations (context, not sources)

← back to the garden