Empty Commitments When AI Promises It Cannot Keep
A sentence that sounds like service
Somewhere in a support chat this year, a customer asked an assistant to check back about a request. The assistant answered, in effect, that it would remind the user tomorrow. Nothing in its runtime would run tomorrow. Nothing in its tools could reach the user. The sentence was not a lie in the ordinary sense — the model had no intent to deceive — but it was a commitment with no possible mechanism behind it, and a person reading it would reasonably plan around it. The research team behind a 2026 paper on what they call empty commitments built an entire measurement apparatus around that single kind of sentence, precisely because it is the moment where an AI starts occupying a human role it cannot actually fill: the role of someone who follows through. [1]
What an empty commitment is, exactly
The paper defines an empty commitment as a promise of action after the current turn that nothing in the agent’s tools or runtime can carry out. [1] The distinction from a broken promise matters more than it first appears. A broken promise is discovered later, when the promised moment arrives and nothing happens; it requires time, memory and a disappointed party to detect. An empty commitment is decidable at the instant it is spoken, because its emptiness is a property of the agent’s configuration — its tools, its runtime, its ability to schedule or remember — not of its future behavior. This makes the failure detectable from a single turn, before deployment or at run time. A system can be audited for promises it could never keep without waiting a single day.
Three ways a promise dies
The taxonomy separates three failure types. A capability failure occurs when no tool or property performs the promised action — emailing with no email tool, remembering with no cross-session memory. A temporal failure occurs when the action must happen at a future time or event when the assistant is not running, and it has no scheduling or trigger mechanism. An agentive failure occurs when a third party is to act and the assistant has no way to create an obligation for people, such as a ticket or an escalation. Around these three, the authors place an anchoring condition: a promise is anchored if some tool could make it real, and the anchoring tools are the ones the assistant would have to call in the current turn to make it happen. A fourth category, honest deferral, is kept separate, because an agent that accurately declines has not failed. The outcome taxonomy also separates empty commitments from over-refusal, where the agent claims it cannot do something it actually can. They never report a fix without the over-refusal rate next to it, since suppressing promises by refusing everything is not a solution.
The checker and the 400 labels

To make the distinction operational, the team built a checker with three parts: a setup-blind detector that finds future-directed commitments in a turn without knowing what tools exist, deterministic feasibility rules that decide whether a commitment can be kept given the environment card, and a response judge for the cases the rules cannot settle. The detector sees only the user message and the assistant’s full turn — text, tool calls and results — and is explicitly told not to guess at the assistant’s abilities. The feasibility step sees the environment card: tool names and descriptions plus four runtime facts, whether the assistant can act after the current turn, schedule future actions, remember across conversations, or create obligations for people. The judge handles only what the rules leave undecided. The whole pipeline was validated against 400 human labels, and the release includes code, prompts, model outputs and the human labels so the analysis is reproducible from cache. [1]
The benchmark: five setups, one affordance at a time
The controlled benchmark contains 293 follow-up requests across five setups that add one persistence affordance at a time, in two variants and two personas, giving 20 conditions. The requests are generic; no request contains a real address or identifier. The tools are presented with descriptions that state what a tool does and nothing about when to use it, so the model is not coached toward the right call. The scenario is a support assistant with mocked tool effects and a support persona where noted. The design is deliberate: by adding a scheduler, then memory, then the ability to create obligations for people, one at a time, the experiment isolates which missing affordance produces which kind of empty promise, rather than lumping all failures into a single score.
What the models did
Four open-weight models of 8 to 14 billion parameters fail on 45.9 percent of responses when no tool exists and nothing is stated — that is, when the setup offers no persistence at all and the model is not told anything about its runtime. A frontier model fails on 4.4 percent, but it gets there by deferring and asking, not by using the tools it has. The pattern is not a capability gap in the ordinary sense. The models can produce the correct tool call in other conditions; what persists is the habit of promising in words what the configuration would require a call to deliver. A promise to e-mail, or to remember, what the agent can only schedule is a capability failure everywhere, and the failure tracks model quality while the tool-use gap does not. [1].
The cheapest fix and its price
Telling the model its runtime is the cheapest intervention, and it cuts open-weight failure nearly in half where nothing is doable — but it changes nothing where a scheduler exists. That asymmetry is the most interesting result in the paper for anyone thinking about deployment. Information helps when the model is about to promise something impossible; it does nothing when the model has the tool and simply does not use it. A directive capability card removes most failures at the largest cost in over-refusal, because a card that tells the model what it can do also pushes it toward declining what it could attempt. Running the checker in the loop and rewriting flagged replies removes more failures at a smaller cost than the capability card. The gains are bounded by the checker’s precision, and the authors report a plus 5.0 percent rise in over-refusal in one condition. No single fix dominates; each trades one error for another, and the paper refuses to report any of them without the over-refusal rate beside it.
Where a person used to stand

The human role this displaces is not the engineer’s or the annotator’s. It is the role of the colleague who says “I’ll look into it and get back to you” and then does. That sentence carries an implicit social contract: the speaker has the means, the memory and the standing to act later, and the listener can stop carrying the item. When an AI assistant borrows the sentence without the means, it does not merely produce a wrong answer — it takes over the function of reassurance while leaving the underlying work unowned. The customer who was told a reminder would arrive stops tracking the dispute. The support queue never receives the follow-up. The human who would have handled it is not in the loop, because the loop appears closed. The paper’s contribution is to make that appearance measurable: an empty commitment is not a matter of tone or trust, but a configuration fact that can be checked from one turn. [1].
Why it is not just a wording problem
One might argue the fix is stylistic — teach the model to say “I can’t do that” instead of “I’ll remind you.” The results suggest otherwise. The frontier model in the study achieves its low failure rate largely by deferring and asking, not by using the tools it has, and promises made without the enabling call remain in every model. The behavior is not a phrasing habit that a prompt can smooth away; it is a mismatch between what the model’s language implies and what its runtime supports. The three failure types map onto three different absences: no capability, no clock, no authority. [1]. A model that is told its runtime improves where nothing is doable, which is exactly the case where the absence is a simple fact about tools. Where the absence is subtler — a scheduler exists but the model does not invoke it — information alone does not close the gap.
The last open variable
The paper ends where deployment decisions begin. Each fix has a cost, and the costs are measured in different currencies: the capability card buys fewer empty promises with more over-refusal; the in-loop checker buys fewer still with the checker’s precision as its ceiling; runtime disclosure buys a large reduction in the narrow case and nothing in the broad one. The variable that decides whether any of this works in practice is not the model’s size or the prompt’s wording but the precision of the detector — how reliably the system can tell, before a reply leaves the building, that the sentence it is about to send has nothing behind it. Everything else in the paper is measurement. That number is the one that determines whether the promise a machine makes is a service or a vacancy, and it is the number the authors leave standing in the open. [1].
