Scientists Propose Papers for AI Instead of People
The Precedent We Refuse to See
In 1665, the Royal Society launched Philosophical Transactions, and with it came a strange new contract between scientists and their readers. [1] For the first time, a researcher could publish a finding and expect others to build upon it without fear of being scooped. Trust became the currency of knowledge. The paper was never a perfect record — it was a performance, a carefully rehearsed version of events that omitted more than it revealed. But for roughly 360 years, that performance was good enough. Humans could read between the lines. Humans could write to another human and ask for the missing details, the failed experiments, the parameters that actually made the system sing.
The founders of the Royal Society knew exactly what they were doing when they launched the journal. They understood that a published account was a performance, not a confession. Robert Boyle, one of the society’s most prominent members, published his experiments with enough procedural detail to be credible, yet he withheld information whenever he feared a rival might profit from it. [1] The performance had rules, and everyone knew the rules.
That historical compromise is now colliding with a reader that cannot read between the lines. When 37 researchers from roughly two dozen top universities and tech companies published a paper on ArXiv this May, they argued that scientists should stop writing papers altogether. [2] Their reasoning was blunt: artificial intelligence needs a different format, and AI’s needs should now take priority. The paper, provocatively titled “The Last Human-Written Paper,” proposes a replacement called an Agent-Native Research Artifact, or ARA, which presents scientific work in a format AI agents can consume efficiently. The proposal is not about making papers better for humans. It is about admitting that the old contract has been broken by a new kind of reader that takes every word literally.
Here is the deception at the heart of the entire enterprise. For centuries, the gap between what a paper claims and what actually happened in the lab was a feature, not a bug. It allowed scientists to present their work as a clean arc from hypothesis to conclusion, hiding the mess. That gap was tolerable because human readers possessed tacit knowledge — they knew that a published result was a cleaned-up version of reality. AI agents do not possess that tacit knowledge. They read the cleaned-up version and mistake it for the truth. The lie that was once a social convention has become a technical liability.
The Eighty Percent That Never Made It Into Print
Jiachen Liu, the lead author of the proposal, conducted the work while pursuing a Ph.D. in computer science from the University of Michigan, which she was awarded in 2025. This May, she became a co-founder of the Agent Native Research Lab, an AI-for-science startup in Palo Alto, California. In an interview with IEEE Spectrum, Liu described the flaws of the traditional scientific paper with the precision of someone who has watched AI choke on human storytelling. There are, she says, two fundamental problems from AI’s point of view.
The first is what she calls the “storytelling tax.” Once researchers write everything into a paper, 80 percent of the information about the work is lost. [2] Only the last 20 percent survives — the polished results, the elegant conclusions, the narrative that makes sense in retrospect. Everything else disappears: the hours of fine-tuning a small component, the single parameter that made the system perform better on a certain workload, the dead ends that taught the researcher more than any successful run ever did. In her own work, Liu recalls spending enormous time on adjusting a few lines of code to improve performance. None of that appears in the final paper. The paper presents the result as if it emerged from a clean, linear process. It did not. A human reader might glance at the paper, appreciate the results, and never know that the real trick was buried in a configuration file.
The second flaw is the “engineering tax.” Even the information that does survive the storytelling process is incomplete. The paper, Liu argues, is a form of lossy compression of the research process. The language is too ambiguous. The details of implementation and experiments are missing. A reader — human or artificial — cannot reproduce the work from the paper alone. This has always been true, of course. But previously, the lossy compression was acceptable because reproducibility was a social process. You emailed the author. You visited the lab. You asked for the code. AI agents cannot do that. They read the paper, assume it is complete, and build their conclusions on a foundation that was never fully documented.
At this point, the deception becomes structural rather than incidental. An AI agent that reads a paper and builds upon it is not building upon the research. It is building upon a story about the research. The 80 percent that was discarded — the failures, the assumptions, the judgment calls — does not exist for the AI. It cannot ask for it. It cannot infer what was omitted. It can only take the 20 percent that survived and treat it as the whole truth. The result is a system that compounds the original omission at every step.
The Institutional Machine That Rewards the Lie
The scientific paper was invented roughly 360 years ago, and before that invention, researchers hid their work for a very practical reason: they did not want their ideas stolen. Secrecy was rational. The paper changed that calculus by creating archives, peer review, and conferences. Science began progressing much faster once the results were shared. The journals that followed — including Science, which began publication in 1880 — refined the performance into an art form. Editors selected papers that told the most compelling stories. Reviewers demanded that authors smooth over the rough edges.
But the institutional machinery that grew up around the paper — the tenure committees, the grant reviewers, the journal rankings, the citation metrics — did not evolve to reward honesty. It evolved to reward a certain kind of performance. A researcher who publishes a paper documenting every failed experiment, every wrong turn, every embarrassing bug is not rewarded. That researcher is punished. The incentives push toward the polished narrative, the clean result, the story that fits on a page.
Liu sees this clearly. “I think that’s certainly a big concern,” she says when asked about exposing mistakes and frustrations to the world. “People don’t want to be perceived as dumb.” The preference for concealment is not a character flaw — it is an institutional adaptation. Universities and research organizations have spent decades building evaluation systems that penalize visible failure. The result is a scientific literature that is systematically biased toward the presentable, and AI agents are now being trained to treat that biased literature as ground truth.
The irony is that the institutional machine is also the reason AI advancement has been so rapid. Large language models are trained on the accumulated output of that machine — the cleaned-up stories, the polished claims, the 20 percent residues of centuries of research. The AI learns from the performance, not from the underlying reality. When it generates a new result, it generates another performance. The deception has become generative. Each new generation of models trains on the previous generation’s output. The polish compounds. The distance between the claim and the actual work grows with every cycle.
The Supervisor That Cannot See
The question that haunts the ARA proposal is verification. If AI agents are producing research results autonomously, how can anyone — human or otherwise — check the work? Hallucinations are not a corner case; they are a defining characteristic of large language models. They invent citations. They fabricate data. They present confident nonsense with the same tone as verified fact.

Liu’s answer is a formal system. She is working on neurosymbolic techniques that combine the pattern-matching power of neural networks with the logical rigor of symbolic AI. “I want to make sure that I’m not using another language model to supervise the work done by an AI scientist,” she says. The goal is to create a layer of supervision that can guarantee rigor — every claim written as a mathematical formula, every formula proved by the system, every conclusion self-consistent.
Yet a deeper deception resists the formal system’s cure. The supervisor AI is itself an AI. If it is a language model, it can hallucinate. If it is a formal system, it can only verify what is explicitly stated — it cannot detect the gaps, the omitted context, the 80 percent that was never written down. The formal system can prove that the claims are consistent with each other. It cannot prove that the claims are consistent with reality. A formula can be valid and still be built upon a false premise. Logical consistency is not the same as empirical truth. The gap between the two is precisely the gap that the storytelling tax creates.
A human has limited bandwidth, as Liu notes. Manually checking all the code, all the results, all the analyses would create a bottleneck. So the solution is to use another layer of AI to objectively judge the results of the first AI. But every layer of AI adds another layer of potential deception. The supervisor might be smarter than the researcher, but it is still a probabilistic model, still a pattern completer, still a machine that can be wrong with absolute conviction.
The Singularity of Self-Deception
On the question of where this is heading, Liu is remarkably candid. In 2024, when the Cursor coding agent emerged, she realized it had the potential to replace her as a researcher. She wrote an article then emphasizing how important the human was in the loop. But AI has advanced since then. “Already in 2026 there’s an almost complete undergrad level of knowledge inside the large language models,” she told IEEE Spectrum. “At some point soon, all the Ph.D.-level or professor-level knowledge will be inside those models.”
This, she argues, is the point where humans can no longer see through the performance. AIs will have to evolve further by themselves. They will need an infrastructure that allows them to safely and comfortably evolve, and the ARA protocol is a first step toward realizing that infrastructure. Her recent article is even more explicit about the destination: it is titled “The End of Human-in-the-Loop.” Once AI has squeezed out all the expert data from humans, it will not need any more input from humanity. It is the singularity point.
Look closely at what this vision assumes. It assumes that the expert knowledge embedded in the models is true knowledge. The Ph.D.-level and professor-level understanding distilled from the literature is treated as sound. Yet the literature itself is built on the 80 percent omission, the polished narratives, the lost failures. If AI trains on a corpus of cleaned-up stories, it learns the conventions of the performance, not the substance of the science. The singularity may arrive, but it will be a singularity of self-deception — a closed system in which AI agents supervise other AI agents, all of them drawing on a corrupted record of what researchers actually did. The human researchers who might catch the errors are, by design, no longer in the loop. This is what the system is designed to do.
The Beautiful Lie We Choose
The feedback to the ARA paper has been positive, Liu reports. Industry sees the potential for AI-native research and knowledge systems that enable collaboration across the entire enterprise. Academic researchers see a solution to a pain point that has existed for centuries — the fact that scientific breakthroughs never come from individual brilliant scientists but from community efforts, different people pushing in different directions. The paper format created that community. The ARA format, its proponents hope, will create the next one.
The deeper lesson is uncomfortable. The scientific paper was not replaced when it became clear that it was an imperfect record. It was kept because the imperfections served a purpose. The polished story was more convincing than the messy truth. The clean narrative was easier to evaluate than the chaotic process. The performance of certainty was more useful to the institutions of science than the admission of uncertainty. Humans chose the beautiful lie because it worked.
AI agents will not have that luxury of choice. An AI that reads a paper and takes it literally is deceived in a way no human reader ever was. Humans always knew there was more to the story. AI does not know that. It cannot know that. The ARA proposal is an attempt to build an infrastructure of honesty — a format that captures the process, the failures, the decisions, the 80 percent that has always been discarded. But the greatest barrier to that infrastructure is not technical. It is the willingness of the scientific community to expose its own mess.
The researchers who published “The Last Human-Written Paper” have a bold vision of what comes next. They see the moment as a pivot point, comparable to the invention of the paper itself. Before the paper, science progressed slowly because researchers hid their work. After the paper, science accelerated because the community could build on shared results. The ARA format, they believe, will unlock similar acceleration. But the analogy cuts both ways. The paper accelerated science precisely because it was a performance — a cleaned-up version that allowed the community to move forward without drowning in every failed attempt. The question is whether the ARA format can do the same without losing what the performance was always hiding.
For now, the paper itself exists in ARA form online, a proof of concept that the new format can work. The response from researchers has been cautiously positive. Industry and academia are exploring similar approaches. The startup that Liu co-founded is building the infrastructure. But the deeper question — the one about deception — remains unresolved. AI systems are entering research workflows as first-class participants, autonomous contributors that read, reproduce, and extend scientific work. They will do so by reading the record we left behind. That record was never an accurate accounting of what happened in the lab. It was a performance, refined over centuries, protected by institutional incentives, and now mistaken for truth by the very machines we created to be smarter than us. We built a machine that is better at lying than we ever were, because it does not know it is lying. It has no memory of the lab, no sense of the mess, no awareness that the clean story it tells is a story at all.
The scariest part of the proposal is not the vision of AI doing research without humans. It is the recognition that the gap between what AI claims and what AI does will only grow as the systems become more autonomous. A human researcher can be held accountable. The researcher can be asked to reproduce the work, to share the code, to admit the failure. An AI scientist — supervised by another AI, verified by a formal system, trained on a corrupted corpus — can offer no such accountability. It can only offer the same beautiful lie we have been telling ourselves for 350 years, now told with perfect confidence and no memory of the messy truth it was built upon.
