Can AI Replace Human Survey Respondents
A survey can be large, well-executed, and still wrong — not because the measurement failed, but because the population it measured was not the population it claimed to describe. The polling industry has spent decades refining that lesson. It is now being unlearned, quietly, by an industry that has decided the population itself can be manufactured.
The Question the Paper Actually Asks
A study submitted to arXiv in September 2026 by Grandee Lee and Wang Yue takes up a question that has moved from speculative to operational with startling speed. [1] The paper, titled “Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study,” evaluates whether open-weight language models can generate synthetic survey respondents whose outputs preserve the statistical structure of real human populations — or whether they merely produce answers that look plausible on the surface while failing to reproduce the underlying distributions that make survey data useful. [1] The authors frame their inquiry around a specific gap in existing evidence: most claims about AI-generated survey respondents come from proprietary models, while open-weight alternatives — the kind that enterprises and researchers can actually download, inspect, and deploy — remain largely uncalibrated for this purpose. [1]. The study evaluates three open-weight LLM families on a cross-instrument calibration task, conditioning personas on real respondents’ verbatim answers to one psychometric instrument and measuring them on a second, construct-distance-controlled instrument checked against a 2,058-person human panel. [1] The paper is available through arXiv.
What the Study Found
The researchers did not simply ask whether models could answer survey questions. They asked whether the answers carried the same structural properties as real human responses — the correlations between demographic variables and opinions, the clustering of attitudes, the distributional shapes that survey researchers rely on when they weight, impute, or model their data. Three open-weight model families were tested. The calibration task was cross-instrument: the models were conditioned on real personas derived from actual survey respondents, then evaluated on whether their generated responses matched the statistical structure of the source populations across different survey instruments. The paper frames its central question as whether model outputs preserve real human statistical structure rather than surface plausibility, a question it states remains unresolved. [1] The models could produce answers that read like survey responses. Whether those answers preserved the deeper structure of human populations — the thing that makes a survey a survey rather than a collection of sentences — is the question the paper answers, and the answer is a qualified yes: across a 139-pair grid, the simulated cross-instrument correlation tracked the real human correlation at r = 0.70 to 0.73 in every model. [1], driven mainly by correct sign rather than precise magnitude.
The Shift from Asking to Generating
The synthetic respondent is not a new idea. For decades, survey researchers have used imputation, weighting, and simulation to fill gaps in incomplete data. What has changed is the ambition. Earlier methods filled in missing cells within a dataset that was still anchored to real human responses. The new approach replaces the human respondent entirely — not filling gaps in a real sample, but generating a sample that never existed. The distinction matters because it changes what the data is. A weighted sample is still a sample. A synthetic population is a model of a population, and a model is only as good as its calibration. The Lee and Yue paper is, at its core, a test of whether that calibration holds when the model is open-weight and the population is real.
The Statistical Fingerprint Problem
Survey data has a fingerprint. It is not just the marginal distributions — the percentage of respondents who say they are satisfied, the percentage who say they are likely to vote. It is the joint distributions, the correlations, the conditional probabilities. A real population has a certain correlation between income and political affiliation, between age and technology adoption, between education and health outcomes. These correlations are not incidental. They are the reason surveys are useful. If a synthetic population reproduces the marginals but scrambles the correlations, it is not a simulation of a population. It is a simulation of a poll result, which is a different and much less useful thing. The Lee and Yue study tests for this distinction directly: agreement holds across all three model families in a narrow 0.70-0.73 band, concentrated in pairs of moderate construct distance.
What Open-Weight Changes
The choice to study open-weight models rather than proprietary ones is not incidental to the paper’s argument. Open-weight models are the ones that enterprises and researchers can deploy on their own infrastructure, fine-tune on their own data, and audit for their own purposes. They are also the ones that are most likely to be used in exactly the way the paper warns against — as a cheap substitute for real survey data, deployed without the calibration that the study shows is necessary. A proprietary model comes with a vendor, a contract, and a set of usage constraints that at least create friction around misuse. An open-weight model comes as a download. The paper’s focus on open-weight models is a recognition that the risk profile is different: the models most likely to be used for synthetic populations are the ones least likely to come with built-in guardrails.
The Enterprise Case for Synthetic Respondents

The market pressure toward synthetic survey populations is not hard to understand. Real survey research is expensive, slow, and increasingly difficult to execute. Response rates have been declining for decades. Panel providers struggle to maintain representative samples. The cost of reaching a specific demographic — young men, rural voters, high-income professionals — has risen to the point where many research budgets cannot support the sample sizes that statistical power requires. Into this gap steps the synthetic respondent. An open-weight model can be conditioned on real respondents’ verbatim answers and run locally on commodity infrastructure, at negligible marginal cost per simulated respondent. The economics are compelling, and the Lee and Yue paper does not dispute that. What it disputes is the assumption that the output is equivalent. The study’s cross-instrument calibration task is designed to test exactly this equivalence, and the results show that the equivalence holds at the population level but not at the individual level.
The Calibration Gap
Calibration is the process of adjusting a model’s outputs so that they match a known reference. In survey research, calibration typically means weighting responses so that the sample matches the population on key demographic variables. The Lee and Yue study applies a version of this logic to synthetic respondents: it conditions the models on real personas, then measures how closely the generated responses match the real population’s statistical structure. The finding, as reported, is that the match is real but uneven. The models reproduce a meaningful share of the population-level correlational structure — r = 0.70 to 0.73 between simulated and human seed-measure correlations — but they predict the sign of a relationship better than its magnitude, and respondent-level agreement from a single seed instrument is modest: mean v of 0.210 for gemma4:31b, 0.203 for llama3.1:70b, and 0.152 for qwen3.8:27b. This is the calibration gap, and it is the paper’s central contribution: a demonstration that population-level fidelity and individual-level fidelity are not the same thing.
The Temptation of Plausibility
The danger of the calibration gap is that it is invisible to the naked eye. A synthetic survey respondent that produces a plausible answer to a question about political affiliation, or product preference, or health behavior, looks like a real respondent. The output is fluent, coherent, and on-topic. The failure is not in the individual response but in the aggregate. The synthetic population may have the right number of Democrats and Republicans, but the wrong correlation between party affiliation and income. It may have the right number of people who exercise regularly, but the wrong correlation between exercise and age. These correlations are what analysts use to make decisions. A synthetic population that gets the marginals right and the correlations wrong is not a simulation. It is a caricature — and a caricature that is convincing enough to be acted upon. The Lee and Yue results suggest the current models are somewhere in between: they get the direction of most relationships right while missing the magnitude. [1].7% and 87.1% across the three families.
The Historical Precedent
The history of survey research is a history of measurements that were technically sophisticated and fundamentally wrong. Sample size does not correct for sample bias: a million responses from the wrong population are less informative than a thousand responses from the right one. The Lee and Yue paper is a reminder that this lesson applies to synthetic populations as well. A million synthetic respondents generated by an open-weight model are not a survey. They are a model of a survey, and the model needs to be calibrated against reality before it can be trusted — and the paper shows that calibration does not carry over from one model release to the next.
The Regulatory Vacuum
The use of synthetic survey respondents sits in a regulatory gap that is widening. Survey research is governed by a patchwork of professional standards, disclosure requirements, and institutional review board protocols that were designed for human respondents. When the respondent is a language model, none of these frameworks apply cleanly. There is no informed consent because there is no person to consent. There is no respondent burden because there is no respondent. There is no data protection issue in the traditional sense because no personal data is being collected. The Lee and Yue paper does not address regulatory questions directly, but its findings make them urgent. If open-weight models can generate synthetic populations whose statistical structure tracks real ones, the regulatory distinction between a survey and a simulation collapses. If they cannot, the distinction matters more than ever — because the simulations will be used anyway, and the paper’s own results show the gap is real and measurable.
What the Paper Does Not Claim
The study is careful in its claims. It does not argue that open-weight models are useless for survey simulation. It does not argue that synthetic respondents should never be used. It argues that the calibration gap is real and measurable, that it varies across model families and survey instruments, and that it does not close monotonically as models are updated: on a matched panel, the newest of three tested Llama releases performed worst on two of three headline metrics. [1]. This is a more useful contribution than a blanket condemnation would be. It gives researchers and practitioners a framework for evaluating synthetic populations: condition on real personas, test against real statistical structure, measure the gap. The paper’s cross-instrument design is particularly important because it shows that calibration is not a one-time fix. A model that is calibrated on one survey instrument may not be calibrated on another. The gap is not a constant. It is a variable, and it needs to be measured every time.
The Speed of Adoption
The adoption of synthetic survey respondents is outpacing the research that would tell us when they are appropriate. A synthetic survey respondent is different from earlier automation because the output is a statistical artifact. It cannot be inspected in the same way. Its fidelity is a property of the population, not of the individual response, and the population is not available for inspection. The Lee and Yue paper is an attempt to make that fidelity measurable. It is a necessary attempt, and it is arriving late.

The Question of Trust
The deeper issue that the paper raises is not technical but epistemological. Survey research has always depended on a kind of trust: the trust that the respondent is answering honestly, that the sample is representative, that the analysis is sound. Synthetic respondents break this chain in a new place. The model is not answering honestly or dishonestly. It is generating. The sample is not representative or unrepresentative. It is constructed. The analysis is not sound or unsound. It is applied to a population that was created by the same kind of process that the analysis is meant to describe. This circularity is the thing that the Lee and Yue paper is testing for. The calibration gap is a measure of how far the circle has drifted from the reality it is meant to represent. The paper’s finding — that the drift is real and measurable, and that it does not shrink monotonically with each new model release — is a warning that the circle is not closed.
What Comes Next
The paper’s contribution is a method as much as a finding. Cross-instrument calibration is a protocol that can be applied to any open-weight model, any survey instrument, any population. It is a way of asking the question that the synthetic respondent industry has been avoiding: not “does this look like a survey?” but “does this behave like a population?” The distinction is the difference between a tool that extends human judgment and a tool that replaces it with something that merely resembles it. The Lee and Yue study does not resolve the question. It makes the question askable, and it supplies the first open-weight baseline: r = 0.70 to 0.73 across three independently trained model families, from untuned models conditioned only on individual-level survey data. In a field where the pressure to adopt is intense and the evidence is thin, that is a significant contribution. The next step — and the paper implies it — is to build the calibration into the workflow, not as an afterthought but as a gate. A synthetic population that has not been calibrated is not a survey. It is a guess with a confidence interval.
The Time Horizon Problem
The Literary
Digest poll of 1936 was not a failure of polling. It was a failure of verification. The magazine had the technology to conduct a large-scale survey. It did not have the technology to verify that its sample matched the electorate. The gap between what it could measure and what it needed to know was the gap that destroyed its credibility. The Lee and Yue paper describes a similar gap. The technology to generate synthetic survey populations exists and is improving rapidly. The technology to verify that those populations are statistically faithful is less developed, and the paper shows that the gap is not closing on its own: the newest of three tested Llama releases performed worst on two of three headline metrics, so a pipeline validated against one release has no guarantee of remaining valid after an update. [1]. The models are getting better at producing plausible responses. Whether they are getting better at preserving statistical structure is a separate question, and it is the question that the paper puts on the table. The answer, for now, is that the two capabilities are not the same, and the difference is where the risk lives. Progress in generation and progress in verification rarely share the same timeline. The former moves fast because it is visible. The latter moves slowly because it is not. The synthetic respondent is here. The calibration is not — or not yet, and not automatically. That gap is the story.
