🌿freegardner

Synapse

General Factor in Language Models Lacks Support

30 Sep 2026 · via Rss.arxiv

General Factor in Language Models Lacks Support
Image: Wikimedia Commons (Public Domain)

General Factor in Language Models Lacks Support

A single number has organized how the field talks about progress: the general factor, the one score that supposedly captures how smart a model is across the board. A factor analysis of 13,251 published evaluation scores drawn from 1,618 language models and 456 text-only benchmarks (Haznitrama et al.) puts that number under a microscope and finds it does less work than assumed — at the most generous estimate, a general factor accounts for 70.8 percent of variance in model performance, and far less in most of the solutions the authors tested. [1] The rest is domain-specific, idiosyncratic, and not identifiable in practice. That is the finding. What follows is what it means for the human judgment that the field has been quietly outsourcing to a construct that may not exist as advertised.

The Psychometric Borrowing

The paper, by Faiz Ghifari Haznitrama, takes what it calls a latent variable approach. [1] The premise is familiar from the human side: performance on any specific problem set is influenced by a domain-specific latent factor and a domain-agnostic one. The domain-agnostic one is the general factor, the g. In human intelligence research, g is the thing that survives after you factor out the specifics — the common thread running through vocabulary tests, spatial reasoning, pattern completion. The paper asks whether the same structure holds for language models, and it asks it at a scale no prior attempt had reached. The dataset is super-sparse, meaning most models were never evaluated on most benchmarks, so the authors triangulate across different data densifiers and imputation methods rather than trusting a single pipeline. [1]. That methodological caution is itself part of the argument: if the general factor were robust, you would not need to hedge this hard to find it.

What the Factor Analysis Actually Found

Three results come out of the analysis, and each one chips at a different assumption. First, the general factor’s share of variance is 70.8 percent at the most generous estimate and far less in most solutions — meaning the majority of what distinguishes one model’s performance from another is not captured by a single underlying ability. [1]. Second, content-similar benchmarks do not necessarily cluster together. Two evaluations that look like they test the same skill can load onto different factors, which undermines the practice of picking a handful of representative benchmarks and treating them as proxies for a capability. [1]. Third, the g factor is not dominated by any common theme, and the authors find a lack of evidence that it is well-proxied by standard “intelligence” benchmarks. [1]. The benchmarks the field treats as the canonical measure of general ability are not, in this analysis, the ones that best capture the general factor — if the general factor is even the right thing to be capturing.

The Construct That Cannot Be Targeted

Its conclusion is blunt: the findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. [1] That is a stronger claim than “general intelligence is hard to measure.” It is that the strategy of targeting a single conceptual ability lacks support, because the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice. In other words, you cannot build toward a target you cannot locate. The role this displaces is not the engineer’s — it is the evaluator’s, the person whose job was to look at a model’s scores and render a judgment about what it can do. That judgment has increasingly been delegated to the general factor: If g is only partially interpretable, the delegation was premature.

The Psychometric Precedent That Cuts Both Ways

General Factor in Language Models Lacks Support (Image 1)
AI-generated image

The borrowing from psychometrics is not incidental — it is the paper’s method and its implicit warning. The same ambiguity that psychometricians have long wrestled with now applies to models. A field that adopts the vocabulary of fluid intelligence inherits the vocabulary’s problems. And the problems are not academic: The factor analysis says the construct is stable enough to account for a majority of variance in the most favorable reading, and not much more than that in the rest.

Where the Human Judgment Was Supposed to Sit

The practical consequence is a vacuum. If no single ability can be targeted, then the work of deciding what a model is good for falls back to people who can hold multiple, partially conflicting signals at once. [1] — the domain expert who knows that a model acing a reasoning benchmark may still fail on the specific task in front of them, the auditor who knows that two similar-looking evaluations can diverge, the operator who knows that a high aggregate score says nothing about the failure mode that matters in their context. The paper does not say humans should make these calls. It says the technical shortcut for making them does not hold up. That is the displacement in reverse: the construct was supposed to make human judgment superfluous, and the evidence says it cannot.

The Sparsity

Problem as a Mirror The dataset’s super-sparse nature is worth sitting with, because it is not just a methodological nuisance. It reflects how the field actually evaluates: the authors’ need to triangulate across densifiers and imputation methods is a symptom of that fragmentation, not a fix for it. [1]. A general factor that only emerges under some imputation choices and shrinks under others is a factor whose existence depends on how you fill in the blanks. The evaluator filling in those blanks — deciding which missing scores to trust, which benchmarks are comparable, which models belong in the same analysis — is doing exactly the interpretive work the general factor was meant to automate. The paper’s method enacts the problem it describes.

The Cost of the Shortcut

There is a version of this story where the general factor is a useful fiction: not real, but predictive enough to guide decisions. The paper’s numbers make that version harder to defend. A construct that explains 70.8 percent of variance at its most generous and far less in most solutions is not a reliable guide when the decisions at stake are deployment, safety, or resource allocation. The gap between the generous estimate and the typical one is where the risk lives. An evaluator who treats the generous number as the real one will overestimate how much a single score tells them. An evaluator who treats the typical number as the real one will conclude, correctly, that they need to look at the specific abilities — and that looking is a human task, not a factor-analytic one.

The Idiosyncrasy That Resists Aggregation

The second finding — that content-similar benchmarks do not necessarily cluster together — deserves its own weight. It means the intuition that “these tests are basically the same” fails empirically. [1]. Two benchmarks that appear to probe the same skill can load on different factors, which implies that the skill they probe is not one thing. For anyone whose job is to certify a model’s fitness, this is the operative constraint: you cannot substitute one evaluation for another on the assumption that they measure the same underlying capacity. You have to run both, or accept that you are guessing. The construct was supposed to license the substitution. Without it, the evaluator’s workload does not shrink — it expands, because every benchmark now carries information that no other benchmark can replace.

General Factor in Language Models Lacks Support (Image 2)
AI-generated image

What the g Factor Is Not

The third finding is the quietest and possibly the most damaging: the g factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard “intelligence” benchmarks. [1]. This means the benchmarks the field most often cites as evidence of general capability are not the ones that best capture whatever general factor exists. The proxy relationship is broken. In practice, this invalidates a common move: citing performance on a canonical reasoning or knowledge benchmark as evidence of broad competence. The benchmark may be measuring something real, but it is not measuring the general factor, and the general factor is not measuring what the benchmark measures. Two constructs, loosely coupled, both treated as interchangeable. The judgment that was supposed to be replaced by the benchmark score is not replaced — it is obscured.

The Interpretability That Was Promised

The paper’s title says machine intelligence is only partially interpretable, and the factor analysis is the evidence for that claim. [1]. Partial interpretability is a specific condition: not opaque, not transparent, but somewhere in between, where some structure is visible and some is not. It is the visible part. The idiosyncratic first-order abilities are the invisible part. The problem is that the visible part has been treated as the whole. A field that reports a single capability score and calls it interpretability has mistaken a summary for an explanation. The paper does not offer a better summary. It offers a demonstration that the summary is incomplete, and that the incompleteness is not a temporary limitation of data or method — it is a feature of the construct.

The Human Role That Remains

If no single ability can be targeted, then the work of targeting — deciding what a model should be able to do, evaluating whether it does it, and acting on the answer — stays with people. [1]. Not because people are better at it, but because the technical substitute does not exist. This is the paper’s practical implication, and it is not a comfortable one. It means the evaluation burden does not decrease as models proliferate. It means the demand for domain expertise in assessment does not shrink. It means the judgment that was supposed to be automated is still being made, just less visibly, by whoever chooses which benchmarks to run and which scores to report. The factor analysis makes that invisible labor visible again, which is the first step toward doing it well.

The Next Step the Research Suggests

The paper’s own triangulation across densifiers and imputation methods points to the next move: rather than searching for a single general factor, map the first-order abilities directly, with methods that tolerate sparsity and do not assume a common thread. [1]. That is a harder research program than reporting a g score, and it is the one the evidence supports. For the people who use these evaluations, the corresponding move is to stop treating any single number as a summary of capability and start treating each benchmark as a partial, local measurement. It was a convenience. The paper shows what the convenience cost, and what remains when it is set aside: a set of specific, partially idiosyncratic abilities, and the human judgment required to navigate them. That judgment was never actually replaced. It was just hidden behind a number that did less than advertised.


Sources

1. arXiv — Paper

← back to the garden