LLMs Generate Scientific Laws But Cannot Select Best
Scientific law discovery has always been a human act of judgment as much as a mathematical one. From Kepler’s planetary laws to Hubble’s law, the compact equations that anchor physics, chemistry, and biology emerged from iterative cycles of generating hypotheses, testing them against empirical evidence, and refining them under scientific constraints. That cycle has long depended on a researcher’s ability to look at a candidate relationship and decide whether it is real, meaningful, or merely a numerical coincidence. A new benchmark suggests that when large language models take over the generation step, the judgment step does not come with them.
The paper introducing SciLaws-Bench, a curated collection of scientific task packages grounded in source literature, reports that current models can propose laws that fit data well while violating the physical constraints that make those laws scientifically valid. [1] The authors find that good predictive fit need not imply scientific validity, that recovering a published formula does not establish recovery of its new structural terms, and that candidate selection remains a bottleneck in scientific law discovery. The third finding is the one that matters most for anyone who assumed that better generation would automatically produce better science.
What the benchmark actually measures
SciLaws-Bench assembles 118 problems spanning six disciplines, drawing on 381 papers, 291 candidate laws, and roughly 8 million data points. [1] Each task package links a scientific problem to supporting data, published reference equations, and scientific-validity rubrics. The construction combined agent-assisted curation with human verification of critical variable mappings, formula implementations, data splits, and rubric items.
Two complementary evaluation settings structure the tests. SciLaws-Real fixes the existing scientific data and leaves the candidate law open, asking models to find laws that better explain held-out observations than published formulas while respecting each problem’s physical and other scientific constraints. SciLaws-Parallel fixes a synthesized generating equation and lets models query new data from a simulator. That hidden equation is a newly synthesized structural variant of a published formula, subject to the problem’s scientific constraints. Coefficients and residual noise are calibrated to the corresponding source data. Models start without observations, choose simulator queries, and infer the fixed hidden equation from returned responses, enabling evaluation of structural recovery against a target known by construction.
The design responds to a specific weakness in earlier evaluations. The challenge the benchmark sets is to evaluate discovery beyond familiar equations while keeping tasks grounded in scientific data and domain-specific constraints.
The gap between fitting and knowing
Fourteen LLMs spanning leading proprietary and open-weight families were evaluated under the same agent framework, with access to a Python sandbox and, in SciLaws-Parallel, a fixed-budget experimentation interface.

The results expose a failure mode that has nothing to do with computational power. On a nuclear-physics task, the model’s best-fitting formula introduced a nonexistent resonance peak near 47 eV. It therefore failed the benchmark’s validity checks. The model optimizes for fit, and fit alone will happily accommodate a phantom.
The paper observes that models use memory to reproduce known laws but struggle to discover novel structure. Reference-beating — producing a formula that outperforms the published baseline on held-out data — is more common on tasks with low recall, while full-law recovery varies substantially across models, from 0% for GPT-4o-mini to 42% for GPT-6. A closer analysis of trajectories reveals distinct discovery styles: some models emphasize breadth, screening many candidate laws in a few turns, while others emphasize depth, refining a few candidates over many turns. Neither style resolves the underlying problem.
The selection bottleneck
The sharpest finding concerns what happens after generation. Scoring the intermediate candidate laws within trajectories shows that weaker models propose laws as good as the leading models’ but submit worse ones. Best-of-N search exposes the same gap across trajectories, with self-selection capturing only a small fraction of the available gains. In other words, the models are not failing to find good laws. They are failing to recognize which of the laws they found is the good one.
This is a different kind of failure than the resonance peak, and it cuts deeper. A model that cannot generate a valid law is limited. A model that generates a valid law but cannot identify it as the best candidate has replaced the human researcher’s generating capacity without replacing the human researcher’s judgment. The candidate pool contains the answer. The model does not know it.
The authors state directly that self-evaluation and candidate selection remain a bottleneck in scientific law discovery, as agents do not consistently select the best laws available in their candidate pools. [1] The phrasing is careful and the finding is narrow, but its implication for the division of labor in scientific research is not. If the selection step is where the bottleneck sits, then adding more generation capacity — more models, more turns, more candidate laws — will not clear it. The bottleneck is not upstream.
Where the human role actually sits
It would be easy to read the benchmark as a scorecard for whether AI can do science, and to conclude that the answer is “not yet.” That reading misses what the paper actually demonstrates. The evaluation separates predictive fit from scientific validity, and the models fail on the second axis while often succeeding on the first. The separation is not a technical detail. It is the entire structure of scientific judgment.
A published law earns its place not because it fits a dataset but because it survives constraints: it remains well behaved in the relevant regime, it respects physical boundaries, it retains meaning in context. The SciLaws-Bench rubrics encode those constraints explicitly, translating scientific requirements into validation criteria. The models are being asked to satisfy criteria that a human researcher internalizes over years of training and that a research community enforces through peer review. The benchmark makes those criteria explicit precisely because the models do not carry them implicitly.

The curation process itself illustrates the point. Constructing the task packages required linking evidence across papers and datasets, standardizing heterogeneous data, implementing published formulas, and translating scientific constraints into explicit validation criteria. Domain experts and human verifiers checked critical variable mappings, formula implementations, split rationales, and rubric items against cited sources. The humans were not generating candidate laws. They were deciding what counted as a valid law, what data could be trusted, and what constraints applied. That is the work the models cannot yet do.
The open variable
The paper’s own framing points to what remains unresolved. The authors note that scientific law discovery requires more than goodness of fit or syntactic simplicity, and that proposed laws must respect physical constraints, remain well behaved in the relevant regime, and retain scientific meaning in context. The benchmark tests all three. The models pass some and fail others.
What the results leave open is whether the selection bottleneck is a property of current architectures, a consequence of training objectives that reward generation over discrimination, or a more fundamental limitation in how these systems represent scientific knowledge. The authors do not resolve this. They report that candidate selection remains a bottleneck and that self-selection captures only a small fraction of available gains. The last open variable — the one that determines whether AI-assisted law discovery becomes a genuine scientific tool or a generator of plausible-looking equations that no one can trust — is whether the judgment step can be built into the system or must remain outside it.
For now, the machine finds the law. The question of which law deserves to be called one still belongs to the person who knows what a resonance peak should look like. SciLaws-Bench
