Machine That Says It Is Listening
When the Transcript Lies
Speech recognition has spent a decade getting very good at one thing: turning a finished recording into text. Play back a podcast, feed it through a model, and the words come out clean. This is offline recognition, and by 2026 it is close to solved for major languages. The trouble starts when the audio has not finished yet.
A voice agent — the kind that answers a phone, takes an order, or sits inside a customer-service line — cannot wait for the recording to end. It has to produce words while the speaker is still speaking. That means it must guess. It must decide, at each fraction of a second, whether the sound it just heard is enough to write down, or whether more context is coming that would change the answer.
Here is the deception. A streaming system can appear to work. It emits words. The words look right. But underneath, it has made a choice that the user never sees: it has either waited too long, making the conversation feel sluggish and broken, or it has committed too early, and the transcript quietly drifts wrong. The system does not announce which error it is making. It simply produces output that looks like language and behaves like a coin flip.
The standard way to hide this problem is to tune a single global knob. Set a fixed delay. Set a fixed chunk size. Set a fixed look-ahead window. The system then behaves uniformly: every position in the audio waits the same amount of time, or commits at the same acoustic boundary. This is convenient to implement and easy to describe in a paper. It is also a fiction. The amount of future context needed to correctly transcribe a sound is not constant. Some syllables are unambiguous the moment they are spoken. Others — the ones that sound like three different words until the next phoneme arrives — need a pause. A single knob cannot express that distinction. Setting it later delays every position, including the ones that needed no wait. Setting it earlier exposes every position to the same risk. The system claims to be adaptive. It is not.
The Position-Dependent Problem
The X2Streaming-ASR paper states the issue directly: the amount of future context needed after an acoustic boundary is position-dependent. [1] That phrase is the whole argument compressed into four words. It means the decision to wait or commit is not a property of the model as a whole. It is a property of each individual output position.
Existing approaches fall into three families, and none of them solves this. The first uses global configuration — fixed chunk decoding, fixed look-ahead, fixed delay. These determine the decoding unit and how much future context is visible, but they do not decide how much future information to wait for based on the history and the current acoustics. The second family shifts emission time during training. One method encourages earlier output across the board. Another pulls emissions toward the alignment boundary. Both are global adjustments. They do not wait longer where upcoming context is specifically needed for disambiguation. The third family emits an unstable hypothesis and later revises it. These revision-based methods reduce latency by correcting earlier output, which means they are not single-pass and not hard-commit. The system says one thing, then quietly says another.
X2Streaming-ASR, described in a paper available on arXiv, takes a different route. [1] It decomposes streaming recognition into two separate questions: when to commit and what to commit. The when is handled by a policy that decides, at each position, whether to wait or emit. The what is handled by the recognizer itself. Crucially, each character is emitted only once. There is no second pass, no correction, no revision. The system lives with its decisions.
This is a harder constraint than it sounds. In a revision-based system, an early mistake can be patched later. The final transcript may look fine even if the intermediate output was wrong. A single-pass system has no such luxury. Every commit is final. The gap between what the system appears to do — produce fluent text — and what it actually does — gamble on each character — is exposed at every step.
Training the Policy to Know When to Stop
The architecture alternates between two states: listening and decoding. While the language model is decoding a previous audio segment, the causal audio encoder is simultaneously encoding the current audio. These two processes run asynchronously. The model drops the global delay parameter entirely and instead lets itself decide whether to wait or commit.
Training happens in three stages. The first stage builds a streaming recognizer that emits only characters whose acoustics have already ended. If two and a half characters of audio have arrived, it emits only the first two. This checkpoint establishes basic streaming ability and provides labels for the next stage.

The second stage is where the commit policy gets its first education. The researchers use a forced aligner to align the reference text, then randomly cut the aligned character sequence into contiguous blocks of varying length. They probe the Stage-1 model: starting at the end of the first character, the model greedily decodes incremental audio. The longest reference prefix that matches the hypothesis gets assigned a commit time. For the next mismatched character, probing moves forward to that character’s end time, or continues in 80-millisecond steps, never backward. Utterances are discarded if probing ends before every reference character receives a commit time, or if any character waits longer than 640 milliseconds past its acoustic end. Such long waits, the researchers note, teach the model to keep choosing “wait” and can prevent it from ever emitting.
This is a subtle point about deception. A model trained to be cautious will learn to wait indefinitely. It will appear to be listening carefully. In practice it is paralyzed. The 640-millisecond cap is a guardrail against a system that confuses patience with accuracy.
The third stage refines the commit policy using group-relative policy optimization. The system samples a group of wait-emit trajectories at the sentence level. If the action is “emit,” the model decodes greedily until an end-of-sequence token. A commit segment is an emit action together with the consecutive wait actions that precede it. The sentence is the sampling unit, but not the credit unit. Broadcasting one sentence-level reward to every frame would reinforce useful and superfluous waits equally, and the policy would learn to wait throughout the utterance. Instead, the system compares trajectories on the same reference character and assigns the advantage to the corresponding commit segment. Each decision in a segment shares the segment’s advantage, so the policy strengthens or suppresses wait/emit actions according to what actually helped or hurt that specific character.
The result: a system that commits characters with a mean leftover-wait latency of 32 to 109 milliseconds on Chinese and 12 to 85 milliseconds on English, measured against forced-aligned endpoints. [1] Recognition accuracy remains comparable to existing systems. On ten test sets spanning both languages, it achieves the lowest mean commit latency. This is the technical achievement. The system does what it claims: it waits when waiting helps and commits when committing is safe.
The Illusion of the Global Knob
What makes this work interesting is not just the latency numbers. It is the diagnosis of why previous systems failed. A global delay is a confession that the system does not know where it is. It treats every moment in the audio stream as equivalent. But language is not equivalent. A vowel in the middle of a common word is predictable. A consonant at the start of a name is not. A system that waits the same amount everywhere is not being careful. It is being uniform, and uniformity is a form of ignorance.
The revision-based approach is a different kind of deception. It looks accurate because it corrects itself. The final transcript may be perfect. But the intermediate output — the part that a real-time agent would act on — was wrong. If the agent starts speaking a response based on a hypothesis that later gets revised, the revision comes too late. The user has already heard the mistake. The system’s apparent accuracy is purchased with a delay that defeats the purpose of streaming.
X2Streaming-ASR avoids both traps by making the commit decision explicit and position-specific. The model does not hide behind a global parameter. It does not hide behind a later correction. It decides, at each position, whether it has enough information. And it lives with that decision.
The training procedure has its own honesty. Stage 2 probes only one-tenth of the data. The researchers argue that full annotation is costly and unnecessary because this stage only needs to teach an initial commit decision. Stage 3 then refines with rewards that are assigned at the character level, not the sentence level. The credit assignment is fine-grained because the problem is fine-grained. A system that rewarded whole sentences would learn to wait everywhere. A system that rewarded every frame equally would learn nothing. The segment-level advantage is the mechanism that makes the policy actually learn the position-dependent behavior the researchers identified as missing.
What the Numbers Do Not Say
The evaluation covers ten Chinese and English test sets. On Chinese, the system outperforms open-source models and a quoted baseline on several sets. On English, it achieves state-of-the-art performance on some datasets and ranks just behind a larger model on others, with a negligible margin. The P95 latency is modest: 240 to 480 milliseconds. Compare that to baseline systems that wait 720 milliseconds to more than a second.
These numbers are real. They are also narrow. They measure commit latency against forced-aligned endpoints. They measure recognition accuracy against reference transcripts. They do not measure how a user feels when a voice agent pauses for 300 milliseconds instead of 100. They do not measure whether a customer hangs up because the system seemed to hesitate. They do not measure the difference between a transcript that is 98 percent accurate and one that is 99 percent accurate when the missing 1 percent is a name or a number.
The system solves a technical problem: how to decide when to commit under a single-pass constraint. It does not address the deployment question: what counts as fast enough, and who gets to decide. A latency of 32 milliseconds on Chinese characters is impressive in a paper. In a phone call, it is imperceptible. A latency of 109 milliseconds is also imperceptible. The difference between them matters to a benchmark. It may not matter to a person.
The deeper issue is that the system’s honesty is invisible. A user cannot tell whether the agent waited because it was uncertain or because it was slow. A user cannot tell whether the transcript was committed early and got lucky or waited and got it right. The system’s internal decision — wait or emit — is hidden behind the fluent output. The deception is not that the system lies. It is that the system’s correctness is indistinguishable from its luck.

The Standard Nobody Set
This is where the problem leaves the paper and enters the deployment. The paper optimizes commit latency and recognition accuracy. Those are the metrics the researchers chose. They are reasonable metrics. They are also not the only possible metrics. A system could be optimized for user satisfaction, for conversational flow, for the cost of errors. None of those are measured here.
The gap between what the system measures and what matters is not a flaw in the research. It is a feature of how research works. You cannot optimize what you cannot measure, and you cannot measure what you have not defined. The researchers defined the problem as position-dependent commit decisions under a single-pass constraint. Within that definition, they solved it. The solution is elegant, the training procedure is principled, and the results are strong.
But the definition itself is a choice. The system decides when to commit based on a policy trained to maximize recognition accuracy and minimize latency. It does not decide based on whether the user is confused, whether the conversation is flowing, whether the agent’s response will be helpful. Those considerations live outside the model. They live in the application that uses the model. And the application has no way to know whether the model’s commit decision was confident or desperate.
A voice agent built on this system will appear to listen. It will produce words at the right moments. It will seem to understand. What it actually does is execute a policy that was trained on a proxy for understanding. The proxy is good. It is not the thing itself. The system that says it is listening is really saying: I have decided, at this position, that I have heard enough. That decision is a bet. The bet is well-informed. It is still a bet.
The Quiet Gap
The most important thing about X2Streaming-ASR is not its latency numbers. It is the fact that it makes the bet explicit. Previous systems hid the bet behind a global parameter or a later correction. This system puts the decision at the center of the architecture and trains it directly. The honesty is structural. The model cannot pretend it did not choose.
That honesty does not transfer to the user. The user sees a transcript. The user hears a response. The user does not see the policy that decided when to emit. The system’s internal clarity is invisible from the outside. This is the gap that no engineering can close. A system can be transparent to its developers and opaque to its users. A system can measure its own uncertainty and still present a confident face.
The researchers behind X2Streaming-ASR have built a better bet. They have shown that the amount of future context needed is position-dependent, and they have built a system that respects that fact. They have replaced a global knob with a learned policy. They have replaced sentence-level rewards with segment-level credit assignment. They have made the commit decision a first-class part of the model rather than an afterthought.
What they have not done — what no one can do — is make the user see the decision. The system says it is listening. It is really deciding, moment by moment, whether it has heard enough. That decision is now better informed than ever. It is still a decision. The gap between the appearance of understanding and the mechanism of commitment remains. It is narrower than before. It is not gone.
The technical problem is solved. The question of who defines what counts as good enough, and who bears the cost when the bet goes wrong, is not. That problem does not live in the model. It lives in the room where the model is deployed, in the conversation where the transcript is used, in the moment when a user hears a pause and wonders whether the machine is thinking or failing. The machine cannot answer that question. It can only decide, again and again, whether to wait or to speak.
