First Speaker Bias in Multi Agent AI Debate
The Hidden Cost of Sequence in Multi-Agent Debate
When several large language models debate a question in sequence — one speaks, then the next, then the next — the arrangement looks neutral. Each model contributes its reasoning; the strongest argument should win. A 2026 paper by Duofeng Xu, Bryan Hooi, and Dandan Qiao shows that this assumption is wrong in a specific, measurable way. [1] Their study, When Order Matters: First-Speaker Bias and Mitigation through Personality in Sequential Multi-Agent Debate, demonstrates that the agent who speaks first exercises disproportionate influence over the final answer. The experiments hold the participating models, the questions, and the debate protocol fixed and vary only the speaking order. [1] This is not a minor calibration issue. It means that placing a more capable model later in the sequence can substantially erode its reasoning advantage. The order of speech, not just the quality of thought, shapes the outcome.
What the Researchers Measured
Xu, Hooi, and Qiao designed experiments around sequential multi-agent debate, a format where LLMs take turns responding to a prompt. They tracked how much each agent’s contribution shaped the final consensus. The finding: first-speaker bias is pronounced and consistent. Agents who open the debate pull the group’s conclusion toward their own position at rates that exceed what their reasoning quality alone would justify. When a weaker model speaks first and a stronger model speaks second, the stronger model’s superior reasoning does not reliably correct the course. The first speaker has already framed the problem, set the terms, and anchored the discussion.
This matters because the intuitive fix — put your best model last, let it synthesize and correct — does not work as expected. The paper’s experiments show that placing the stronger agent after weaker ones substantially offsets its reasoning advantage: aggregated across four benchmarks and ten weak-strong model pairs, strong-agent influence falls by 21.01 percentage points and final accuracy by 2.34 percentage points compared with the strong-agent-first order. The reasoning advantage is real, but it competes against a structural bias that favors whoever spoke first.
Personality as an Intervention

The researchers did not stop at diagnosing the problem. They tested whether personality prompting could rebalance influence. Drawing on the Big Five model, they applied agreeableness and extraversion as behavioral interventions to either the strong or weak side of the debate. The results were trait-specific. Influence consistently shifted in the direction of lower agreeableness. [1] Assigning low agreeableness to the stronger agent while leaving the weak agents unprompted raised strong-agent influence by 9.55 percentage points and final accuracy by 1.64 percentage points relative to the no-personality baseline. Extraversion produced less systematic changes in influence and accuracy; its clearest effect appeared in the length of the agents’ justifications, not in persuasive weight.
The finding is precise: not all personality traits are equally useful as levers. Agreeableness, which governs cooperativeness and accommodation, directly affects how much an agent yields to or shapes the group’s direction. Extraversion, which governs assertiveness and social energy, mostly changes how much an agent talks, not how much it persuades. For anyone designing multi-agent debate systems, this distinction matters. If you want to correct a first-speaker bias, you need to target the trait that governs influence, not the one that governs volume.
Why Sequence Is Not Neutral
The paper’s core insight is that sequential debate is not a neutral aggregator of opinions. The order of speech is a variable, and it has measurable effects. This runs counter to a common assumption in AI system design: that if you assemble sufficiently capable models and let them reason together, the best reasoning will surface. The research shows that the best reasoning can be muted by the order in which it is presented.
The mechanism is not mysterious. In any sequential exchange, the first speaker establishes a frame. Subsequent speakers respond to that frame, often adjusting their positions toward it. The first speaker does not need to be right; they need only to be first. The bias is pronounced and systematic, not a function of any individual model’s flaw.
The Fix That Misses the Problem
The problem is not that a particular model always speaks first. The problem is that whoever speaks first gains an influence that is not proportional to their reasoning quality. The paper’s own data show that simply maximizing the strong agent’s influence is not the goal: when the strong agent was made less agreeable and the weak agents were made more agreeable, its influence rose to 71.43 percent but final accuracy fell to 67.23 percent, below the 68.08 percent achieved when only the strong agent was prompted.

The personality intervention offers a more targeted approach. By lowering the agreeableness of the stronger agent, the researchers restored some of its lost influence. This is not a cosmetic adjustment. It changes the dynamics of the debate in a way that improves final accuracy. The intervention works because it addresses the mechanism — the tendency of later speakers to defer to the first speaker’s frame — rather than just the symptom. The paper’s authors are careful to frame these prompts as controlled behavioral configurations, not as claims about stable human-like personality in language models.
The Necessary Consequence
If sequential debate is not neutral, then any system that relies on it must account for the order of speech. The first speaker’s advantage is not a bug that can be patched with a single fix. It is a feature of how sequential reasoning works, and the paper draws a direct parallel to human team decisions, where who speaks first can anchor a discussion that later participants never fully recover from. The effect is not uniform, however: the paper finds that the influence advantage of speaking first shrinks as the capability gap between the strong and weak models widens. The researchers’ contribution is to show that this feature can be measured, that its effects are substantial, and that targeted interventions can mitigate them.
The consequence follows directly: as multi-agent AI systems become more common, the design of interaction protocols will matter as much as the design of individual models. The question is not just whether an AI can reason well, but whether the system it operates in allows that reasoning to shape the outcome. Xu, Hooi, and Qiao have shown that the answer depends, in part, on who speaks first — and that targeted personality prompting can partially correct for that advantage. The paper’s own framing is narrower than the headline suggests: it is a study of one protocol, three-agent teams, and two personality traits, and its authors present the findings as a starting point for protocol design rather than a finished fix.
