GAIATrace Captures Full Token Behavior of AI Agents
The agent decides to call a search tool. It does not know, in any meaningful sense, that it is making a decision. It has simply observed a partial result, weighed it against an internal plan, and chosen the next move. Somewhere in a data center, a GPU is being provisioned for a workload that will not look like the last one. Nobody watching the screen sees the fork in the road. The fork is buried in a token stream that has never been recorded — until now.
A team of researchers has released GAIATrace and Vidur-Agent, a dataset and simulator that, for the first time, capture the full token-level behavior of two modern agentic systems — MiroThinker and OWL — as they solve a heterogeneous mix of general-purpose tasks from the GAIA benchmark. [1] The work is not a product announcement. It is a measurement instrument, and what it measures is the gap between what we assume about autonomous AI and what it actually does.
What the trace reveals that logs never did
Previous trace datasets captured only the surface: queries in, answers out. GAIATrace preserves the full reasoning tokens, the task-level structure that groups ordered queries by task, and the activity of every participating large language model — including the auxiliary summarization and extraction models that quietly do much of the work. The distinction matters because closed-source models do not reveal their reasoning tokens. When those tokens are invisible, so is the behavior that determines cost, latency, and failure modes.
The dataset covers two systems chosen to represent different architectural families. MiroThinker is ReAct-like: a main LLM generates plans, invokes tools, observes results, and refines its approach. It occasionally calls a second LLM to summarize and extract data from large web-scraped texts. Even this system involves multiple LLMs and tools in the researchers’ setup. OWL is more complex: multiple specialized agents work under a coordinator, each generating its own sub-plan, maintaining memory, and making decisions, with different LLM architectures serving different agents.
The gain here is not raw speed. It is visibility into behavior that was previously only inferable. Before GAIATrace, anyone trying to understand why an agentic system behaved a certain way was working from aggregate metrics and guesswork. Now there is a record of what actually happened, token by token, across the entire system.
The simulator that makes experiments repeatable
Alongside the dataset, the researchers built Vidur-Agent, a trace-driven simulator that extends the Vidur LLM simulator with support for agentic execution. It handles heterogeneous models, KV and prefix caching, prefill-decode disaggregation, and advanced scheduling policies. The practical effect is that GAIATrace can be replayed to perform reproducible, low-cost system evaluation across diverse simulated environments.
This is the concrete lift: experiments that once required bringing up expensive multi-model systems and running controlled trials — which the paper describes as expensive and challenging — can now be run in simulation. The barrier to entry for studying complex agentic behavior drops from building a data center to replaying a trace.
The researchers used both artifacts to characterize how modern agentic systems handle general tasks and how system design choices shape their behavior. What they found complicates several assumptions that have been carried over from chatbot serving.
The pattern that is not monotonic

In simpler ReAct-like setups, input length grows monotonically as each turn appends prior steps to the context. The GAIATrace data shows something different. Agent behaviors are far more heterogeneous than prior work observed, driven by heterogeneity in workloads, agents, models, and tools. The traces deviate substantially from the simple monotonically-increasing prefill pattern.
This matters because a core assumption in LLM serving — that input length grows predictably — underpins many optimization decisions. When that assumption fails, so do the optimizations built on it. The paper does not simply report that the assumption fails; it shows that the diversity of patterns is driven by the diversity of the systems themselves. A multi-agent system with specialized roles does not behave like a single loop. A workload mixing general-purpose tasks does not behave like a benchmark of homogeneous questions.
The gain is a more accurate model of the problem. The cost of not having it is designing infrastructure for a workload that does not exist — and paying for that mismatch in every deployment.
When better per-query metrics make the task slower
One of the paper’s sharpest findings is a misalignment that runs counter to intuition. Certain system design choices significantly worsen per-query metrics such as tail time-to-first-token or time-per-output-token, yet improve per-task latency. The metrics that serving systems have been optimized around are not the metrics that determine whether the task completes faster.
The reason is straightforward once stated: unlike chatbot LLMs, where each query faces a user and must meet strict TTFT and TPOT requirements, most tokens in agentic AI are never read by a human. The tokens are intermediate reasoning, tool outputs, and internal state. Optimizing for the experience of a reader who does not exist is optimizing for the wrong thing.
This is not a minor tuning issue. It suggests that prior optimizations targeting per-query metrics may not be optimal for agentic AI. The paper does not claim they are useless — it claims they are misaligned with the objective that actually matters, which is completing the task. The gain from recognizing this misalignment is the ability to make design choices that trade per-query performance for end-to-end speed, deliberately and with evidence.
The bottleneck that moves
Static hardware provisioning assumes that the bottleneck stays in one place. The GAIATrace study shows that the dominant bottleneck changes with task arrival rate. A GPU allocation that is near-optimal at one arrival rate can be substantially worse at another.
This finding has a concrete implication: dynamic reconfiguration is not a luxury. It is a response to a workload whose demands shift on a timescale that static planning cannot anticipate. The paper does not prescribe a specific reconfiguration strategy; it establishes that the need exists and that the variation is large enough to matter.
The gain is a clearer picture of where the system will strain under different conditions. The cost of ignoring it is provisioning for an average that never occurs, and discovering the gap only when load arrives.
Prefix caching, unevenly

Prefix caching — reusing the KV cache across queries that share a common prefix — has been reported to improve end-to-end latency by around 15.7 percent in prior work. [1] The GAIATrace study found end-to-end task latency improvements of 1.13 to 3.21 times with prefix caching. [1] That is substantially larger than previously reported.
But the benefit is uneven. Some components even experience a degraded TTFT. The paper notes that this again shows how task latency and per-query metrics can misalign. The gain is real and large on average, but it is not uniform, and treating it as uniform would lead to wrong conclusions about which parts of the system benefit — MiroThinker’s Sub-LLM, for instance, sees its TTFT degrade.
The finding is a reminder that aggregate improvements can hide local regressions. The trace data makes those regressions visible. Without it, a system designer might deploy prefix caching broadly, see the overall improvement, and never notice that a specific component got slower.
What the gaps in our knowledge cost
The paper opens with an observation that is easy to read past: the community’s understanding of agentic AI system characteristics remains thin. A recent study offered valuable early insights but examined only a narrow slice of the design space. More complex setups — multi-agent, multi-model systems tackling general tasks with a broad range of tools — remain largely uncharted.
The reason is not lack of interest. It is that bringing up such systems and running controlled experiments are expensive and challenging. The cost of knowledge has been the bottleneck, not the demand for it.
GAIATrace and Vidur-Agent address that bottleneck directly. By providing a token-level trace dataset and a trace-driven simulator, they lower the cost of asking questions about complex agentic behavior. The paper’s authors state their hope that GAIATrace, Vidur-Agent, and their initial study will be a useful foundation for future research in this direction.
The sentence left unspoken
The paper ends with a call for deeper study of complex agentic systems — multi-agent, multi-model, and multi-tool — and workloads. It does not say what happens if that study does not happen. It does not need to.
The unspoken sentence is this: we are deploying systems whose behavior we do not understand, on workloads whose patterns we have not measured, with infrastructure optimized for metrics that do not match the objective. The trace dataset does not solve that. It makes it visible. And visibility, in a field moving this fast, is the difference between building on evidence and building on assumption.
The agent will keep deciding without knowing it is deciding. The question is whether we will know what it decided, and why — and whether we will have measured the system well enough to answer.
