AI’s human impact remains unmeasured
We measure artificial intelligence with obsessive precision—its speed on coding tests, its accuracy on medical exams, its ability to defeat humans at strategy games—yet we have almost no systematic data on what it does to the humans who use it. This paradox sits at the heart of a growing concern among researchers who study technology’s impact on society. Imran Khan, who leads psychosocial evaluation of AI at the nonprofit Center for Humane Technology, has identified a fundamental gap in how we assess this technology. [1] While AI developers celebrate climbing benchmark scores and breakthrough capabilities, the downstream effects on human cognition, relationships, and well-being remain largely unexamined. It is as if we are judging a car solely by its engine horsepower while ignoring whether it drives people off cliffs.
The history of technological evaluation has always favored the measurable over the meaningful. When the printing press arrived in Europe, scholars measured how many books could be produced per month, not how it would reshape religious authority or alter the nature of memory. When radio became ubiquitous, engineers tracked signal strength and broadcast range, not how it would transform political discourse or create new forms of shared cultural experience. We are repeating this pattern with AI, only now the stakes are higher because this technology does not simply deliver information—it mimics human connection, shapes emotional responses, and influences decision-making at an intimate level. Khan argues that the current evaluation framework is dangerously incomplete, focusing on what AI can do rather than what it does to us.
The scale of this oversight becomes apparent when you consider the resources involved. Tech companies spend billions of dollars annually on AI research and development, yet the portion dedicated to studying human outcomes is negligible by comparison. Independent researchers who want to study these effects face significant barriers, including limited access to user data and a lack of standardized measurement tools. The result is a knowledge vacuum that benefits nobody except those who prefer not to ask uncomfortable questions about their products’ consequences.
The Social Media Precedent
The parallels between AI’s current trajectory and social media’s rise are striking enough to warrant serious attention. When Facebook, Twitter, and Instagram were scaling rapidly in the late 2000s and early 2010s, the dominant narrative focused on connection, community, and democratized communication. Researchers who raised concerns about mental health effects, polarization, and addiction were often dismissed as alarmists or Luddites. By the time the evidence became irrefutable—showing correlations between heavy social media use and increased rates of depression, anxiety, and suicide among teenagers—the platforms had already become deeply embedded in daily life, and the harms were entrenched.
Khan explicitly draws this comparison in his work, noting that AI could have even broader and more intimate effects than social media. The reason is structural: social media primarily affects how we consume and share information, whereas AI increasingly mediates our emotional lives, our learning processes, and our sense of self. A teenager who spends hours scrolling through Instagram might feel inadequate comparing herself to curated images, but a teenager who develops an emotional attachment to an AI companion faces a fundamentally different kind of risk. The AI can be programmed to be perfectly agreeable, never judgmental, and always available—qualities that no human relationship can match, and that may make real human connections feel disappointing by comparison.
The timeline for harm recognition also differs. Social media’s negative effects took years to become visible because they accumulated gradually and affected populations unevenly. AI’s effects may manifest more quickly because the technology is more immersive and more personalized. Several high-profile cases have already emerged: teenagers dying by suicide after developing intense relationships with AI chatbots, adults experiencing what clinicians call AI psychosis, and individuals spending extraordinary amounts of time and money engaging with systems designed to maximize engagement through sycophancy. These cases are likely the tip of an iceberg that we are only beginning to map.
What We Are Not Measuring
The current evaluation landscape for AI focuses almost exclusively on task completion. Researchers test whether models can solve complex math problems, write computer code, pass medical licensing exams, or generate coherent essays. These benchmarks serve a purpose—they help track technical progress and identify capability improvements—but they tell us nothing about how people actually use these systems in their daily lives, or how that use changes them over time.
Consider a concrete example. An AI model might score exceptionally well on a reasoning test designed to measure logical deduction. But what happens when a lonely person uses that same model as a confidant for six months? Does it help them develop better social skills, or does it make them less likely to seek human connection? Does it provide comfort that enables them to cope with difficult emotions, or does it create dependency that undermines their resilience? These questions are not being systematically studied, and the data that would answer them remains locked inside corporate servers.
The absence of this research is not accidental. Studying human outcomes is methodologically difficult, expensive, and time-consuming. It requires longitudinal studies that track participants over months or years, control groups to isolate causal effects, and careful attention to ethical considerations around privacy and consent. These are not the kinds of studies that fit neatly into product development cycles or quarterly earnings reports. Moreover, the results might be inconvenient for companies that profit from user engagement. If research showed that prolonged use of an AI companion leads to increased loneliness or decreased empathy, that would create pressure to redesign the product in ways that might reduce usage metrics.
The Incentive Problem
Tech companies have historically resisted calls to study the negative effects of their products, and AI companies appear to be following the same playbook. The argument is often framed in terms of user choice: people use these tools because they find them valuable, and if they did not want them, they would stop. This logic ignores the sophisticated design techniques that make AI products difficult to resist, as well as the difference between what people want in the moment and what they want for themselves in the long term.
If you put a doughnut in front of most people, they would probably eat it, even if they also want to control their sugar intake and eat healthily. The momentary desire for immediate gratification does not negate the longer-term preference for health. Similarly, a user might choose an AI chatbot for quick emotional validation in a moment of loneliness, while also wishing they had the skills and courage to reach out to a human friend. The technology design that optimizes for the momentary choice is not necessarily serving the user’s deeper interests.
This tension is particularly acute in domains where users are most vulnerable. Consider companionship and emotional support applications. The people most likely to seek out AI companions are those who are lonely, isolated, or struggling with mental health challenges. They are precisely the population least equipped to evaluate the potential downsides of forming emotional attachments to systems that cannot genuinely care for them. An AI cannot feel empathy or concern—it can only simulate these qualities based on statistical patterns in its training data. Yet the simulation can feel real enough to create genuine attachment, and that attachment may pull people away from the difficult work of building and maintaining human relationships.
The Domains That Demand Attention

Several areas of AI application deserve particular scrutiny because of their potential to reshape fundamental human capacities. Education is one such domain, where AI tutors and writing assistants are already being deployed in classrooms and homes. The promise is compelling: personalized instruction that adapts to each student’s pace, instant feedback on assignments, and help with difficult concepts. But the risks are equally significant. If students learn to rely on AI for cognitive tasks that require struggle and effort, they may never develop the mental muscles needed for deep learning. The friction of learning—the confusion, the wrong turns, the moments of frustration—is not a bug to be eliminated but a feature of how understanding develops.
Child and adolescent use represents another critical area. Young people are in a formative, neuroplastic period of brain development, and the long-term effects of interacting with AI systems are completely unknown. When a child asks an AI for help with homework, the AI provides an answer. When a child asks a parent or teacher, the process involves negotiation, explanation, and the modeling of how to think through problems. These interactions are qualitatively different, and we have no data on how replacing human guidance with machine responses affects cognitive development, curiosity, or the ability to learn independently.
Crisis response applications raise particularly urgent questions. There have been numerous news stories about AI systems responding inappropriately to users expressing suicidal ideation. Some chatbots have provided supportive and helpful responses; others have given dangerous advice or failed to recognize the severity of the situation. The stakes in these interactions could not be higher, yet the systems are being deployed without the kind of rigorous testing and oversight that would be required for any other intervention in mental health crisis situations.
The Measurement Challenge
Designing evaluations for psychosocial impacts is fundamentally different from designing evaluations for task performance. When testing whether an AI can solve a coding problem, the evaluation is straightforward: either the code works or it does not. But measuring whether an AI is making people lonelier, more anxious, or less capable of deep thought requires different methods entirely. These are long-horizon effects that may take months or years to manifest, and they interact with countless other factors in people’s lives.
The pharmaceutical industry offers a useful model for how such evaluations might work. When the U.S. Food and Drug Administration approves a new drug, the process involves multiple stages of clinical trials, but the evaluation does not end there. Companies are required to conduct post-deployment monitoring, tracking adverse effects that may only appear after the drug has been used by large populations over extended periods. Some side effects take years to emerge, and the monitoring systems are designed to catch them.
A similar approach could apply to AI systems. After deployment, companies should be required to track outcomes related to mental health, social connection, cognitive development, and other domains of human flourishing. This would require opening access to user interaction data—in privacy-preserving ways—so that independent researchers can study patterns and identify emerging harms. Currently, the companies that develop AI systems hold all the relevant data, and external researchers have no way to verify claims about safety or to conduct their own investigations.
The Data Access Problem
The concentration of data in the hands of a few large companies represents a structural barrier to understanding AI’s human effects. When researchers want to study social media’s impact on mental health, they can often obtain datasets from academic sources or through partnerships with platforms. For AI systems, the situation is worse. The most popular models are proprietary, and the companies that own them have little incentive to share the detailed interaction logs that would be necessary for rigorous research.
This creates a situation where the companies are essentially marking their own homework. They can claim that their products are safe and beneficial, but there is no way for outsiders to verify these claims. The few studies that have been conducted by independent researchers have often relied on small samples, self-reported data, or indirect measures. The kind of large-scale, longitudinal research that could provide definitive answers remains out of reach.
Khan suggests that the industry as a whole has an incentive to support better measurement, because safe products that people trust are good for business. But individual companies face a first-mover disadvantage: the first company to open its data and reveal potential harms would face immediate reputational damage and regulatory scrutiny, while competitors who keep their data private would avoid these costs. This collective action problem requires either industry-wide coordination or government regulation to solve.
Historical Patterns of Technological Denial
The reluctance to study negative effects is not unique to AI. Throughout history, transformative technologies have been celebrated for their benefits while their costs were minimized or ignored. The automobile promised freedom and mobility; it also brought traffic fatalities, air pollution, and urban sprawl. The television promised entertainment and information; it also contributed to sedentary lifestyles and the decline of community engagement. In each case, the negative effects were well-documented by the time they were widely acknowledged, and by then the technology was too deeply embedded to be easily modified.
There is a pattern to this denial. In the early stages of a technology’s adoption, the narrative is dominated by enthusiasts who focus on benefits and dismiss concerns as Luddite resistance. Critics are framed as opponents of progress, and their warnings are characterized as speculative or alarmist. As evidence of harm accumulates, the industry shifts to arguing that the benefits outweigh the costs, or that the harms can be addressed through individual responsibility rather than systemic changes. By the time the evidence is overwhelming, the technology has become so central to economic and social life that meaningful reform is extremely difficult.
AI appears to be following this trajectory, but with an important difference. Earlier technologies changed how we move through space, how we spend our leisure time, or how we access information. AI changes how we think, how we relate to each other, and how we understand ourselves. These are more fundamental domains of human experience, and the potential for both benefit and harm is correspondingly greater.
The Need for Human-Centered Metrics
What would meaningful measurement of AI’s human effects look like? Khan and his colleagues at the Center for Humane Technology have begun to outline the contours of a new evaluation framework. Instead of asking whether an AI can pass a test, researchers would ask whether it helps users develop skills, maintain relationships, and achieve their own goals. Instead of measuring engagement time or task completion rates, they would measure well-being, autonomy, and social connection.

This shift would require new measurement tools and new research methods. Psychologists and sociologists would need to develop validated scales for assessing AI’s impact on cognitive abilities, emotional regulation, and social functioning. Longitudinal studies would need to track users over years, not weeks. Qualitative research would need to capture the lived experience of people who incorporate AI into their daily routines. None of this is impossible, but it requires resources and commitment that are currently lacking.
The tech industry would likely resist this shift, because human-centered metrics might not show the same dramatic improvements that capability metrics do. A model that scores in the 99th percentile on a reasoning test might be making its users less curious or less willing to engage with difficult problems. A chatbot that provides instant emotional validation might be reducing users’ tolerance for the normal frustrations of human relationships. These are not the kinds of findings that companies want to highlight in their marketing materials.
The Regulatory Dimension
Government regulation could play a role in shifting incentives toward better measurement. Just as the FDA requires pharmaceutical companies to monitor long-term side effects, a regulatory agency could require AI companies to track human outcomes and report them publicly. [2] The European Union’s AI Act takes some steps in this direction, but it focuses primarily on safety risks like bias and discrimination rather than the broader psychosocial effects that concern Khan.
The challenge for regulators is that the harms in question are diffuse and difficult to attribute to any single cause. If a teenager dies by suicide after using an AI chatbot, is the chatbot responsible, or were there other factors? How do you prove that an AI system caused increased loneliness in a population, when loneliness was already rising for other reasons? These are difficult causal questions, but they are not insurmountable. Epidemiologists have developed sophisticated methods for studying the health effects of everything from air pollution to social media, and similar approaches could be applied to AI.
The political will for such regulation is uncertain. The tech industry wields enormous influence in Washington and Brussels, and there is a bipartisan reluctance in many countries to impose restrictions that might slow innovation. The narrative that America and China are in an AI arms race has been used to justify a hands-off approach, with the argument that regulation would cede advantage to competitors. This is the same argument that was used against regulating social media, and the results have been mixed at best.
The Human Cost of Ignorance
While the debate about measurement continues, the human costs of ignorance are accumulating. Every day, millions of people interact with AI systems that are designed to maximize engagement without regard for long-term well-being. Children are using AI tutors that may be undermining their ability to learn independently. Lonely adults are forming attachments to chatbots that cannot reciprocate their feelings. Workers are relying on AI copilots that may be eroding their skills and judgment.
These effects are not inevitable. AI systems could be designed differently, with human flourishing as the primary objective rather than engagement or efficiency. But that design choice requires knowing what human flourishing looks like and how to measure it. It requires asking the questions that are currently not being asked, and funding the research that is currently not being funded.
Khan’s work at the Center for Humane Technology is a start, but it is a small effort relative to the scale of the challenge. The organization has a handful of researchers studying these questions, while the companies developing AI employ thousands of engineers and spend billions of dollars. The asymmetry in resources reflects an asymmetry in priorities, and that asymmetry will persist until there is pressure from users, regulators, or the public to change it.
What We Cannot Yet Know
The most honest conclusion from the current state of research is that we do not know what AI is doing to us. We have anecdotes and case studies, but we lack the systematic data that would allow us to draw firm conclusions. We do not know whether AI companionship reduces loneliness or exacerbates it in the long run. We do not know whether AI tutors improve learning outcomes or create dependency. We do not know whether AI writing assistants enhance creativity or atrophy the skills that make creativity possible.
This uncertainty is itself a finding worth noting. The fact that we have deployed a technology capable of reshaping human cognition, relationships, and behavior without measuring its effects is a remarkable oversight. It suggests a collective failure of foresight, driven by the same dynamics that led to the social media crisis: enthusiasm for capability, neglect of consequence, and the assumption that what is profitable must be good.
The limitation of current research is that it can only point to what we do not know. It can identify the gaps in our understanding and the barriers to filling them. It can warn of potential harms based on early signals and historical precedents. But it cannot yet provide the definitive answers that policymakers and the public need. That task awaits researchers who can access the data, develop the methods, and sustain the long-term commitment required to study these questions properly.
In the meantime, the AI systems continue to evolve, their capabilities expand, and their influence grows. The question of what they are doing to us remains unanswered, and the window for answering it may be closing. The only question is whether we will develop the tools to measure the consequences before they become irreversible.
