🌿freegardner

Synapse

AI in Academia The Gap Between Claims and Delivery

08 Oct 2026 · via News.mit.edu

AI in Academia The Gap Between Claims and Delivery
AI-generated image

AI in Academia The Gap Between Claims and Delivery

The Olympiad Medal That Proved Less Than It Seemed

Last year, a model reached gold-medal level at the International Mathematical Olympiad. Only a year later, models are producing new research results, including a recently proposed solution to one of the Millennium Prize Problems.

The trajectory appears unstoppable. But look closer at what actually happened, and a different picture emerges.

When an AI system produces a mathematical proof, it generates a sequence of symbols that happens to satisfy certain formal constraints. Whether that sequence constitutes a proof in the sense a mathematician would recognize — a chain of reasoning that illuminates why something is true — remains an open question. The model cannot tell you. It has no way of distinguishing between a solution it found through genuine insight and one that emerged from statistical pattern-matching against training data.

The appearance of mathematical understanding, in other words, is not the same as the substance.

Sasha Rakhlin, director of the MIT Statistics and Data Science Center, has synthesized recent discussions and readings about how AI is changing academia, particularly mathematics, statistics, machine learning, and engineering. [1] His analysis, published through MIT, identifies a central factor that determines the pace of AI progress in a discipline: the speed and reliability of verification.

Formalized proofs can be checked automatically. Programs can be run and tested. Where evaluation is fast and reliable, systems can generate candidates, learn from outcomes, and improve. Better models contribute to the next round of development, creating a compounding process that accelerates progress.

The process is elegant. It is also indifferent to whether the system understands what it is doing.

The Verification Trap

Here is where the deception deepens. The very feature that makes AI progress so rapid in mathematics — automatic verification — creates a blind spot. A proof checker confirms that each step follows from the previous one according to formal rules. It does not confirm that the proof is interesting, that it generalizes, or that it answers the question anyone actually wanted answered.

The gold-medal headlines obscured a distinction: obtaining a solution is not the same as understanding why it works, what generalizes from it, or what to ask next.

The compounding process accelerates regardless. Each verified result becomes training data for the next generation of models. The system learns to produce more results that pass verification. It does not learn to produce more results that matter.

This is not a flaw in the technology. It is a feature of how verification works. The same dynamic appears wherever automated checking is possible: software that passes test suites without being maintainable, statistical analyses that satisfy p-value thresholds without revealing anything true about the world, machine learning models that achieve benchmark scores without generalizing beyond the test set.

The Paper That Proves Nothing

Consider the academic paper. The reasoning was straightforward: if someone wrote a rigorous paper, they must understand the subject. The paper was evidence of intellectual work performed.

That inference is collapsing. A polished manuscript can now be generated with minimal human input. The prose will be fluent. The citations will be plausible. The analysis will follow the conventions of the field. What the paper cannot demonstrate is whether anyone involved understands the material.

Departments must reconsider what they reward. The temptation is to define valuable work as whatever AI cannot yet do. The approach is misguided. The boundary of AI capability shifts too quickly. A skill that seems safely human today may be automated tomorrow. Building an evaluation system around a moving target guarantees that the system will be obsolete before it is implemented.

Instead, contributions such as asking good questions, replicating results, synthesizing ideas across domains, reporting informative negative findings, and sharing datasets may deserve greater recognition. These activities reveal something about the researcher’s judgment. They cannot be faked as easily as a manuscript.

The deeper problem is attribution. When substantial parts of a research project are performed by AI, what exactly did the human contribute? Evaluation should establish what a researcher takes intellectual responsibility for. This expectation should guide hiring, promotion, and funding decisions. It should be made explicit for current and incoming doctoral students.

The old signal was at least correlated with the thing it purported to measure. The new reality severs that correlation: a paper can now be produced without the intellectual work the paper is supposed to represent.

The Muscle That Atrophies

AI in Academia The Gap Between Claims and Delivery (Image 1)
AI-generated image

Intellectual capabilities develop through exercise. This is not a metaphor. The ability to formulate a research question, to recognize when an approach is failing, to develop intuition about which interventions are likely to succeed — these capacities are built through repetition and failure.

Rakhlin identifies a specific risk: routine calculations, coding, failed approaches, and small discoveries have long served as the training ground for graduate students. [1] Delegating this work to AI removes the formative experiences through which expertise develops. The student who never struggles with a proof never learns what it feels like to be stuck, and therefore never learns how to get unstuck.

Some tasks are tedious without being educational. Formatting citations, debugging syntax errors, running standard statistical tests — these can be automated without loss. Other tasks are tedious precisely because they are educational. Working through a long calculation teaches patience and error-checking. Attempting an approach that fails teaches judgment about what to try next.

The problem is that these two categories look identical from the outside. Both involve time spent on tasks that a machine could perform faster. The student cannot know in advance which struggle will produce insight and which will merely consume hours. The advisor may not know either.

Students should learn to formulate problems, audit model outputs, reproduce results, and defend their choices. The fundamentals may become more valuable, not less, as the basis for using AI tools well. A researcher who does not understand the underlying mathematics cannot evaluate whether a model’s output is correct. A researcher who has never written code cannot debug the code a model produces.

The deception here is subtle. AI makes students more productive in the short term while potentially degrading their capabilities in the long term. The student who delegates everything graduates with a credential but without the skills the credential is supposed to certify. The gap between what the degree claims and what the graduate can do widens with each cohort.

The Knowledge That Never Gets Written Down

Rakhlin identifies what may be academia’s most important strategic asset: the accumulated knowledge and experience within its laboratories. [1] This asset is largely invisible.

Published papers present a selective record. They report what worked. They describe the experiment that produced the result. They do not report the experiments that failed, the directions that were abandoned, the reasons an approach did not pan out. A reader of the published literature sees a clean narrative: hypothesis, method, result, conclusion. The actual process was messier.

Scientists and engineers hold tacit knowledge about which interventions are likely to fail and why. This knowledge is acquired through experience. It is rarely written down. It is passed from advisor to student through conversation, through example, through the accumulated intuition of someone who has spent decades in a field.

Rakhlin suggests that this missing context may explain why AI models struggle in some domains, particularly the empirical sciences. An expert can anticipate consequences that are clear from experience but not from any published source. A model trained on the literature has access only to what the literature contains. It cannot know what was tried and failed if no one wrote it down.

The limitation may be temporary. If negative results and scientists’ interpretations could be captured systematically, models might become much better at scientific exploration. The knowledge that currently lives only in the minds of experienced researchers could be made available to the systems that need it.

But capturing that knowledge requires infrastructure that does not yet exist. It requires workflows that record hypotheses, interventions, outcomes, failures, and interpretations. It requires shared systems that connect these workflows across laboratories. It requires investment in compute, secure data systems, and expertise in adapting and post-training AI models.

Rakhlin imagines a future in which a research institution’s laboratories function more like a single, living, scientific organism. A neuroscience laboratory needs better methods for segmenting neurons in microscopy images. An AI agent recognizes a relevant advance from a computer vision group elsewhere in the institution. The agent connects the researchers, proposes benchmarks, helps them iterate. It surfaces unresolved questions to students and faculty. It makes lessons from one laboratory available to others.

The vision is compelling. It is also a description of what currently does not happen: the gap between the published record and the tacit knowledge is exactly the gap that AI systems cannot cross on their own.

The Independence That Cannot Be Bought

Rakhlin is explicit about the goal: universities should be able to pursue questions over long horizons, share results openly, and evaluate claims independently. These are not commercial priorities. Industry partnerships will be essential, but commercial interests will not cover the breadth of science or remain aligned with it over time.

Some degree of technological independence is necessary. The alternative is a research enterprise whose agenda is set by whoever controls the compute.

This is not a new concern. Universities have always depended on external funding. The difference is the scale and concentration of the dependency. When the tools of research are owned by a handful of corporations, the questions that can be asked are constrained by the tools that are available.

Rakhlin’s proposal for shared AI research infrastructure is an attempt to address this. If laboratories within an institution can share models, data, and workflows, they reduce their dependence on external providers. If they can trace the lineage of ideas through the system, they can recognize contributions more easily and collaborate more openly.

The tracing mechanism addresses a problem Rakhlin raised earlier: how to recognize and reward intellectual contributions when the old signals have broken down. If a system records who proposed an idea, who refined it, who tested it, and who interpreted the results, then credit can be assigned based on actual contribution rather than on authorship conventions that no longer reflect reality.

With agreed rules for consent and credit, researchers could collaborate with greater confidence that their work would be acknowledged. This could help answer the question of how to recognize and reward intellectual contributions in a world where AI performs much of the labor.

AI in Academia The Gap Between Claims and Delivery (Image 2)
AI-generated image

What the Machine Cannot Tell You About Itself

The recurring theme in Rakhlin’s analysis is a gap between appearance and reality. AI systems produce outputs that look like understanding. They generate proofs that pass verification. They write papers that follow conventions. They solve problems that once demonstrated expertise.

What they cannot do is explain the difference between producing a correct answer and understanding why it is correct. They cannot tell you which of their outputs are meaningful and which are noise. They cannot anticipate consequences that an expert would recognize from experience. They cannot know what they do not know.

The deception is not intentional. It is structural. The systems are optimized to produce outputs that satisfy measurable criteria. They succeed at this. The criteria, however, are not the same as the goals. A proof that passes a checker is not necessarily a proof that advances mathematics. A paper that meets publication standards is not necessarily a contribution to knowledge. A student who graduates with a credential is not necessarily educated.

AI systems can already synthesize and reason about information from more sources than any individual researcher could absorb, and their ability to make useful connections will continue to improve.

The answer depends on choices that institutions have not yet made. It depends on whether they build the infrastructure to capture tacit knowledge before it is lost. It depends on whether they develop evaluation systems that measure contribution rather than output. It depends on whether they preserve the formative struggles through which expertise develops.

None of these choices are forced by the technology. They are choices about what universities value and what they are willing to invest in. The technology will continue to improve regardless. Whether the improvement serves the mission of the university or undermines it depends on what the university decides to do.

The Next Necessary Question

The path forward that Rakhlin describes is not a solution. It is a direction. Build shared infrastructure. Capture negative results. Trace the lineage of ideas. Rethink evaluation. Preserve formative difficulty.

Each of these steps addresses a specific gap between what AI appears to do and what it actually does. Shared infrastructure addresses the gap between published knowledge and tacit knowledge. Capturing negative results addresses the gap between what models can learn from the literature and what experts know from experience. Tracing contributions addresses the gap between authorship conventions and actual intellectual work. Rethinking evaluation addresses the gap between paper production and research capability. Preserving formative difficulty addresses the gap between short-term productivity and long-term expertise.

The gaps will not close on their own. They will widen if left alone. The compounding process that makes AI so powerful in domains with fast verification will continue to accelerate. The domains where verification is slow, where judgment matters, where experience cannot be written down — these will fall further behind unless deliberate effort is made to bring them forward.

The question is not whether AI will change academia. It already has. The question is whether the change will be guided by an understanding of what the technology actually does or by the appearance of what it seems to do. The difference matters. The appearance is seductive. The reality is complicated.

Rakhlin’s essay is a map of that complication. It identifies where the gaps are, why they exist, and what might be done about them. It does not promise that the gaps can be closed. It does not claim that the path forward is easy. It asks whether universities will invest now to make the path possible.

That is the next necessary question. The answer is not yet written.


Sources

  1. News — Quote source (original article)

Mentioned organisations (context, not sources)

← back to the garden