AI Portfolio Agents Misattribute Market Noise to Skill
The Trade That Looked Brilliant Until You Checked the Market
A fund manager sits down on a Monday morning in 2026, reviews the positions from last week, and concludes that the defensive tilt into utilities was the right call. The portfolio beat its benchmark by two percentage points. The reasoning seems sound: the manager had flagged rising bond yields, rotated into rate-sensitive sectors, and the market rewarded that thinking. There is only one problem. Every other portfolio manager in the building also beat the benchmark that week. So did the index, the equal-weight basket, and the intern who forgot to log in and left the default allocation untouched.
The credit the manager assigns to the utilities call is real in the accounting sense and fictional in the causal sense. This is the exact error that large language model agents make when they manage money, and a new paper published in 2026 by researchers at the intersection of machine learning and finance has put that error on trial. The finding is stark: the AI systems that learn from experience in portfolio management are not learning which experiences work. [1] They are learning which markets they happened to be sitting in when those experiences were retrieved. And when the market turns, the lesson they thought they had learned turns into a liability.
The paper, titled “MemTrial: Learning When to Trust Memory in LLM Portfolio Agents,” describes an agent that does what no human portfolio manager can do reliably: it isolates the contribution of a single piece of advice from the market noise that surrounds every decision [1] The method is not a better stock picker. It is a better bookkeeper of causation. And in building it, the researchers suggest that the judgment human investors prize most — the ability to know which of your past decisions actually mattered — is something an AI can now perform more rigorously than the person who lived through those decisions.
The Blind Spot in Every Learning Agent
The problem begins with a design choice that seemed reasonable. When an LLM agent makes a portfolio decision, it retrieves relevant experiences from its memory — short lessons written after previous trades, such as “avoid concentrated positions in semiconductor stocks during earnings season” or “increase allocation to short-duration bonds when the yield curve inverts.” The agent conditions its decision on these retrieved experiences. Later, when the outcome is known, the agent credits each experience with the outcome of the decision that used it. If the trade made money, the experiences involved gain importance. If it lost money, they lose importance.
This is the credit rule used by FinMem, FinCon, FinAgent, and TradingAgents, among others. It is intuitive, computationally cheap, and fundamentally broken in financial markets. The reason is that the outcome of a portfolio decision is dominated by the market move over the holding period, and every decision made on that date shares this move. An experience retrieved before a rally is credited with the rally. An experience retrieved before a sell-off is blamed for the sell-off. The credit tracks the market, not the experience.
The researchers measured this directly on three real benchmarks, using the utility of the equal-weight (1/N) portfolio over the same holding period as a proxy for the market [1]. On every date, they generated eight drafts with subsets of its retrieved experiences chosen by a fractional factorial design and compared two measures for each experience: its outcome credit, the mean utility of the drafts that used it, and its counterfactual contribution, the mean utility of those drafts minus that of the drafts without it. Outcome credit rose and fell with the market almost in lockstep. It was nearly unrelated to the contribution of the experience, which itself was unrelated to the market.
The arithmetic is simple and damning. Because each experience appears in half of the drafts, its outcome credit equals the date’s mean utility plus half its contribution. The first term dominates [1]. Agents that accumulate outcome credit therefore rank experiences by the markets in which they were retrieved. It is as if a basketball coach concluded that a player is great because the team won the championship during the season that player was on the roster, without checking whether the player was on the court, whether the team would have won anyway, or whether the player’s presence changed the outcome at all.
The consequence is not theoretical. Every experience-learning agent tested underperformed the 1/N portfolio on PortBench and InvestorBench, two standard benchmarks for LLM portfolio agents. The best experience-learning agent finished 15 to 38 percent below the equal-weight baseline across settings on these benchmarks [1]. The agents that were supposed to improve with experience were worse than an agent that never learned anything.
When the Measurement Itself Is the Problem
If the fix were as simple as measuring counterfactual contribution instead of outcome credit, the problem would have been solved years ago. The difficulty is that measuring what an experience changes requires a control group, and in portfolio management, the control group is expensive, noisy, and never repeats. Data valuation offers a remedy in principle: value a data point by its marginal contribution to a utility, averaged over subsets of the other points. But applying it to LLM portfolio agents raises two challenges that the researchers label CH1 and CH2 [1].

The first challenge is measurement. Data valuation evaluates a utility many times on a fixed dataset. For an LLM agent, every evaluation is a paid call whose output varies from sample to sample, and every decision date is a single market episode that never comes back. A contribution measured on one date is therefore mostly noise. On PortBench, the retrieved experiences changed the agent’s utility detectably on only 2 to 4 of 56 months — about as often as chance alone would produce [1]. The signal, if it exists, is buried under sampling variance and market randomness.
This is a problem that human portfolio managers face as well, though they rarely acknowledge it. A manager who makes 12 decisions a year has 12 data points. If each decision is influenced by a dozen factors, the manager cannot isolate the effect of any single factor with statistical confidence. The manager relies instead on narrative, on the story that makes sense of the outcome, on the feeling that the decision was right or wrong. That feeling is the human equivalent of outcome credit. It tracks the market, not the decision.
The second challenge is trust. Even where experiences clearly change the decision — as on about 85 percent of InvestorBench days — the contributions learned from past days do not predict those of the following days [1]. Selecting experiences by their average past contribution still trails 1/N. Existing agents always act on their learned scores. They have no notion of insufficient evidence and no fallback for when their scores do not hold. They are, in effect, portfolio managers who cannot say “I don’t know” and cannot choose to do nothing.
Putting Memory on Trial
The researchers’ answer is an agent called MemTrial, which learns when to trust its memory through a closed loop of four steps. The name is a deliberate double meaning: memory is both the object of the trial and the defendant. The agent does not assume that its experiences are valuable. It tests them, and it acts on them only after they have passed.
The first step is counterfactual attribution. On every date, the agent generates eight drafts of its decision with combinations of its four retrieved experiences chosen by a fractional factorial design. This is a statistical technique from the 1960s, developed for agricultural experiments and industrial quality control, that allows the effect of multiple factors to be estimated from a carefully chosen subset of all possible combinations. The agent estimates the contribution of each experience as its Banzhaf value — the average difference between the drafts with and without it. Because all drafts face the same market, the outcome they share cancels in their difference. The market move, the dominant term in outcome credit, simply disappears [1].
The second step is value estimation. A hierarchical Bayesian model pools the noisy contributions into values across dates and across experiences with similar content. This is the same logic used in small-area estimation, where sparse data from many small regions is combined to produce stable estimates for each region. An experience that appears to help on one date but not another, or that resembles other experiences with consistent effects, is treated accordingly [1].
The third step is the trust test. Following prequential analysis — a method from the 1980s for evaluating forecasts by comparing them to outcomes as they arrive — the agent scores its predicted values against the contributions measured once the outcome is known. It trusts its values only when these out-of-sample forward scores are significantly positive. If the values fail to predict unseen dates, the agent does not act on them [1].
The fourth step is anchored action. If the values are trusted, the agent deploys the draft with the highest predicted value. If they are not, it deploys a blend of its drafts anchored at a conservative reference, such as 1/N, and learns online how far to move from it. The anchor is the agent’s admission that it does not know. It is the equivalent of a portfolio manager saying, “I have no edge this month, so I will hold the index” [1].”
The Numbers That Matter
The evaluation covers three real benchmarks — PortBench, InvestorBench, and ClassAlloc — and one semi-synthetic environment called PlantedMem, which uses real prices and real LLM drafts but in which the quality of every experience is known. The semi-synthetic setting is crucial because it allows the researchers to check whether the agent benefits from experiences that actually matter, not just from experiences that appear to matter in hindsight.
On PlantedMem, MemTrial was the best of 15 methods, including the best experience-learning agent and several non-learning baselines. On PortBench and InvestorBench, where the quality of experiences is unknown and may be poor, MemTrial limited its losses to at most 2.2 percent below 1/N, against 15 to 38 percent for the best experience-learning agent. Averaged over five benchmark settings, its utility was 21.2 percent higher than that of the best experience-learning agent. With eight different LLMs, it beat every LLM-based baseline on InvestorBench [1].
The pattern is consistent: when experiences matter, MemTrial uses them. When they do not, it stays close to the conservative reference. The agent is not smarter than the other agents in the sense of making better predictions. It is smarter in the sense of knowing when its predictions are worthless. That is a form of judgment that human portfolio managers claim to have and frequently do not.

The Human Skill That Just Became Optional
The paper’s contribution is not a better trading strategy. It is a demonstration that the credit assignment problem in experience learning — the problem of figuring out which past lessons actually caused which outcomes — can be solved by an algorithm that treats its own memory as a hypothesis rather than a fact. The agent puts its experiences on trial, measures their effects against a control group, pools the evidence, and refuses to act until the evidence is strong enough.
This is precisely the skill that distinguishes a seasoned portfolio manager from a novice. The novice learns from every outcome, crediting the market’s rise to their own brilliance and the market’s fall to bad luck or external events. The seasoned manager knows that most of what happens is the market, that the signal is buried in noise, and that the only way to learn is to isolate the effect of the decision from the effect of the environment. The seasoned manager also knows when to stop learning and hold the index.
MemTrial does this systematically, on every date, with statistical rigor that no human can match. It does not get tired, does not get overconfident, does not confuse a bull market for skill. It also does not have the human capacity for narrative, for pattern recognition across domains, for the kind of intuitive leap that occasionally produces a genuinely novel insight. But the paper’s benchmarks do not measure those capacities. They measure the ability to allocate capital across assets without losing money to a learning process that is worse than not learning at all.
The uncomfortable implication is that the credit assignment problem in portfolio management — long considered a domain where human judgment is irreplaceable — has been reduced to a statistical procedure. The agent that runs this procedure is not a better investor than a human. It is a better epistemologist. It knows what it knows, knows what it does not know, and acts accordingly. In a field where overconfidence is the most expensive personality trait, that may be the most valuable skill of all.
What Would Have Been Possible
The paper’s final contribution is a comparison that is easy to miss. The researchers did not just build a better agent. They showed what the existing agents could have been if they had been designed differently. The 15 to 38 percent underperformance of the best experience-learning agent on PortBench and InvestorBench is not a measure of the difficulty of the problem. It is a measure of the cost of a design choice — the choice to credit experiences with outcomes rather than with contributions.
That choice was not inevitable. The tools for counterfactual attribution existed. Fractional factorial designs date to the 1960s. Banzhaf values date to 1964. Hierarchical Bayesian models date to the 1970s. Prequential analysis dates to 1984. The researchers did not invent new mathematics. They assembled existing mathematics into a loop that puts memory on trial. The result is an agent that is not just better at portfolio management but better at learning how to learn.
The comparison that matters is not between MemTrial and the other agents. It is between the agents that exist and the agents that could have existed if the field had asked the right question earlier. The question is not “did the trade make money?” The question is “did the experience change the trade?” The first question is easy to answer and leads to agents that track the market. The second question is hard to answer and leads to agents that track the truth. The difference between them, on the benchmarks tested, is 21.2 percent in utility and a much larger difference in the credibility of the claim that these systems are learning anything at all.
In 2026, the year this paper was published, the debate about AI in finance is still framed as a contest between human intuition and machine calculation. The paper suggests that the more important contest is between two kinds of calculation: the kind that confuses correlation with causation and the kind that does not. The human portfolio manager who cannot isolate the effect of their own decisions from the market’s effect is not a better investor than MemTrial. They are a worse version of the same thing. The agent that puts its memory on trial is not replacing human judgment. It is replacing the illusion of human judgment with something that can actually be tested.
