Essay
The Limits of Large Language Models in Predicting Social Science Experiment Outcomes: An Overestimated Promise
Dr. Jerry A. Smith · October 29, 2024 · 12 min read
The Overhyped Promise of AI: Why Large Language Models Fall Short in Predicting Human Behavior in Social Science

Prologue
“The Oracle’s Blind Spot”
Emily tapped her pen against the desk. The rhythmic clicking drowned out by the murmur of excited voices around her. On the projector screen, the bolded numbers flashed across the spreadsheet — 40%. It was the prediction everyone had been waiting for. GPT-4o analyzed demographic data, past test scores, and even socioeconomic indicators. It was supposed to be revolutionary. It was supposed to change everything.
“So, that’s it?” Kevin leaned back in his chair, arms crossed, but a grin creeping at the edge of his lips. “Forty percent improvement in test scores. This program is going to blow up the achievement gap.”
The room buzzed. Emily stared at the screen, the pit in her stomach tightening. She’d been working on the ground with these students for years. She’d seen their struggles, their resilience. But numbers? Numbers couldn’t tell the whole story.
“I mean, that’s… a lot,” Emily said, forcing a smile. “Is the model really that confident?”
“It’s data,” Kevin shrugged. “The model doesn’t get ‘confident.’ It just predicts. And we’ve got the data to back it up.”
The rest of the team nodded, and a few even scribbled notes. Emily’s hand hovered over her notepad, but she couldn’t write anything. She’d visited the classrooms herself and talked to the kids. The model didn’t know them like she did.
“Let’s move forward with implementation then,” someone said from the back. “If GPT’s that clear, we can’t ignore it.”
Days later, Emily sat in the school auditorium, her hands twisting in her lap as the first test results trickled in. She watched as the students filed out of the testing room, their faces expressionless, their pencils still gripped tight as if unsure of what had just happened. She had spent weeks convincing them this program would help.
Her phone buzzed in her pocket. Kevin had already sent the first round of results. She opened the email with a trembling finger. The numbers stared back at her — barely a change. Some students had even done worse.
She blinked hard, her breath catching. How could this be? The AI had been so confident, so precise. But looking at the scores now, it felt like a punch to the gut.
Kevin called her, his voice crackling through the line. “It’s… strange, right? The model predicted way higher outcomes. Maybe there’s an error in the data entry.”
Emily shook her head, even though he couldn’t see it. “It’s not an error. The model didn’t know.”
“Didn’t know what?”
“It didn’t know that Jayden’s been sleeping on his aunt’s couch for weeks because his mom’s apartment got condemned,” she whispered. “It didn’t know Maria’s been caring for her little brothers every night because her dad works two jobs now. And it sure as hell didn’t know that half of these kids didn’t even have internet access to use the study materials we sent them.”
Silence on the other end. She could hear Kevin breathe in, but he didn’t say anything.
“The data,” she continued, “isn’t the whole story.”
Relevant Questions to Consider:
- What does this story reveal about the limitations of relying solely on data and predictions in human-centered fields like education?
- How might models like GPT-4 struggle with understanding individuals' context and lived experiences, particularly those from underrepresented or marginalized groups?
- Can AI systems ever account for the unpredictable and deeply personal aspects of human behavior, or are they inherently limited to what’s visible in the data?
- What ethical implications arise when policymakers or researchers rely heavily on AI-generated predictions without considering the nuances they miss?
- How should we balance the power of AI with the human insight needed to interpret and apply its predictions in real-world contexts?
Abstract
Large Language Models (LLMs), such as GPT-4, have been widely celebrated as tools capable of revolutionizing various research domains, including social science. Proponents argue that these models can effectively predict treatment effects in complex experiments. However, this paper challenges the prevailing enthusiasm and suggests that LLMs’ predictive capabilities in social science are significantly overestimated. Through a critical evaluation of LLMs’ inherent limitations — such as their tendency to overestimate effects, their biases against minority groups, and their inability to grasp the nuances of human behavior — we argue that LLMs, far from being a panacea, pose significant risks when applied to sensitive and context-dependent fields. These models, while powerful in pattern recognition, remain poorly suited to the intricacies of human-centric research and may mislead researchers in the pursuit of ethical and scientifically sound conclusions.
Introduction
In recent years, the development of advanced Large Language Models (LLMs) like GPT-4 has transformed the landscape of machine learning and artificial intelligence. These models, trained on vast corpora of text, have shown an uncanny ability to generate human-like language, summarize complex documents, and, most intriguingly, predict outcomes in various fields, including social science (Brown et al., 2020). The allure of these models is evident: LLMs promise to handle large-scale data, recognize patterns in human behavior, and make precise predictions based on historical data. This has led many to believe that LLMs represent a groundbreaking tool for predicting social science experiment outcomes, promising more efficient, data-driven research.
However, beneath this optimism lies a set of assumptions that are increasingly proving problematic. Unlike the fields where machine learning has traditionally thrived, social science deals with human complexities, behaviors, and emotions that resist simple categorization. While impressive in their capacity to model structured data, LLMs are far less adept at grappling with the unpredictable and context-sensitive nature of human actions. The core issue lies in the technical limitations of LLMs and the misguided application of such tools to domains that require far more than data pattern recognition. This paper aims to dissect the fundamental flaws in using LLMs for social science prediction, critically assess their limitations, and highlight the risks of over-reliance on these models in ethical and scientific research.
Core Problem: Overestimated Predictive Capabilities
The promise that LLMs, such as GPT-4, can predict outcomes in social science experiments has garnered considerable attention, mainly due to these models' impressive results in other fields. Machine learning excels in structured, data-heavy environments where patterns emerge from large datasets and can be easily identified and acted upon (Floridi and Chiriatti, 2020). LLMs have demonstrated remarkable accuracy by leveraging their training on vast datasets in fields like natural language processing, speech recognition, and even medical diagnostics. However, in social science, this reliance on patterns becomes problematic. Social science experiments, which often focus on human behavior, cultural context, and interpersonal dynamics, defy the rigid structures that machine learning models rely on.
A fundamental problem with using LLMs in social science is the overestimation issue. By their design, LLMs generalize from data patterns. When applied to social science experiments, they often exaggerate the significance of treatment effects, predicting outcomes that are larger or more definitive than those observed in practice (Marcus, 2020). This is because LLMs do not understand the underlying mechanisms of human behavior; they only recognize correlations. The complexity of social science, which involves layers of human emotion, culture, and environment, is reduced to statistical noise in LLM processing. For example, in predicting the effect of a social intervention (such as an educational program or public policy change), an LLM may identify patterns that suggest a strong positive impact when, in reality, the effect is marginal or even non-existent.
This overestimation is not a minor flaw but a critical limitation that risks misleading researchers and policymakers. The predictive accuracy of LLMs is contingent on the assumption that past data can reliably forecast future outcomes. However, in social science, where contextual factors can dramatically alter results, this assumption breaks down (Green et al., 2010). As human behavior does not always adhere to predictable patterns, LLMs’ predictions often miss the subtle nuances that drive real-world outcomes. Therefore, while LLMs may appear to offer precise predictions on the surface, these predictions are usually overinflated, offering a false sense of certainty.
The Complexity of Social Science: A Mismatch for LLMs
Social science, by its nature, deals with the unpredictable, context-driven nature of human behavior. This presents a fundamental challenge for LLMs, who rely on the structured and statistical nature of the data on which they are trained. While LLMs are remarkably effective in environments where clear patterns can be extracted from data, social science experiments are often more fluid and contextual, involving variables that are difficult to quantify or model. For instance, cultural differences, emotional responses, and ethical considerations play crucial roles in shaping outcomes, none of which can be easily reduced to the patterns LLMs rely on (Floridi and Chiriatti, 2020). This inherent complexity presents a critical barrier to the effective use of LLMs in social science.
A significant flaw in applying LLMs to social science is the assumption that past data can adequately predict human behavior. Social science experiments often require understanding what people say and why they say it — motivation, context, and socio-cultural influences are paramount. Yet, LLMs are incapable of understanding these deeper contexts. They operate purely on statistical correlations, making predictions based on patterns rather than an understanding of human intention (Bender et al., 2021). As a result, they frequently miss vital variables essential to accurately predicting human behavior in real-world settings.
The inability of LLMs to grasp the complexities of human interaction is particularly problematic in experiments that involve interpersonal dynamics or psychological factors. For example, in studying the impact of social norms on behavior, an LLM might predict outcomes based solely on previously observed behavior patterns without understanding the cultural or ethical factors that underlie those norms. This creates a significant disconnect between the model’s predictions and the actual experiment outcomes, leading to oversimplifications that can distort the conclusions drawn from research (Green et al., 2010). In sum, the mismatch between the rigid, data-driven nature of LLMs and social science's fluid, unpredictable reality is a fundamental flaw that cannot be overlooked.
Bias Against Minority Groups: An Ethical Oversight
One of the most severe limitations of LLMs in social science research is their inherent bias, particularly against minority groups. LLMs are trained on large datasets that, by design, reflect the predominant cultures, narratives, and values of those who produce the data. As a result, minority groups — whether defined by race, gender, socioeconomic status, or other factors — are often underrepresented in the data used to train these models (Gebru et al., 2020). This underrepresentation leads to significant biases when LLMs attempt to predict outcomes for these groups, resulting in skewed or inaccurate predictions that may reinforce existing social inequalities.
This bias is not a trivial issue. Social science research often focuses on understanding and addressing societal disparities, particularly those that affect marginalized groups. However, when LLMs are used to predict outcomes in such studies, their inherent biases can obscure or distort the phenomena researchers are trying to address (Bender et al., 2021). For instance, if a study aims to predict the effectiveness of a policy designed to benefit an underserved community, an LLM trained on data that primarily reflects the experiences of more privileged groups may fail to account for the unique challenges faced by that community. As a result, the model may predict overly optimistic outcomes or, worse, dismissive of the policy’s potential impact on the group in question.
Moreover, using biased LLMs in social science research raises significant ethical concerns. By perpetuating biases present in the training data, LLMs may lead to conclusions that not only misrepresent the experiences of minority groups but also reinforce harmful stereotypes (Gebru et al., 2020). This can have severe implications for public policy and societal understanding. If, for example, an LLM predicts that a minority group is less likely to benefit from a social program based on biased historical data, policymakers may wrongly conclude that the program is ineffective for that group, thus perpetuating systemic inequality. The ethical risks associated with this misrepresentation cannot be overstated and highlight the need for greater scrutiny of LLMs’ role in social science research.
LLMs as Tools of Misuse: The Ethical Risk
As LLMs become more integrated into social science research, there is a growing risk of misuse. By their nature, LLMs are neutral tools, yet they can be easily misapplied in ways that lead to flawed or harmful conclusions. The combination of their tendency to overestimate effects and their inherent biases presents a severe risk when LLMs are used without sufficient oversight. In the hands of researchers who may need to understand the limitations of these models fully, LLMs can produce misleading results that have real-world consequences, particularly in fields like public policy, education, and healthcare (Gebru et al., 2020).
The risk of misuse is particularly acute in studies involving sensitive populations or ethically complex issues. For example, if an LLM is used to predict the outcomes of a survey of criminal justice reform, its biases could lead to predictions that disproportionately affect marginalized groups. This could, in turn, influence policy decisions to exacerbate existing disparities rather than mitigate them (Bender et al., 2021). Furthermore, because LLMs often produce exact and data-driven results, there is a risk that policymakers or other decision-makers will accept these results at face value without questioning the underlying assumptions or biases embedded in the model.
Ethical oversight is crucial to preventing such misuse. Researchers and policymakers must be aware of the limitations of LLMs and ensure that their use is carefully monitored, particularly in studies with significant social implications. Moreover, steps must be taken to mitigate the biases present in LLMs, either by improving the diversity of training data or by developing new methods for adjusting predictions to account for known biases. Such oversight is necessary for using LLMs in social science research to avoid perpetuating the problems these models are intended to solve, leading to flawed conclusions and unethical outcomes.
Conclusion: The Future of LLMs in Social Science — A Cautious Approach
The promise of LLMs in predicting social science outcomes needs to be more accurately estimated. While these models have shown remarkable abilities in fields where data is structured and predictable, their application to social science research is challenging. LLMs tend to overestimate treatment effects, struggle with bias against minority groups, and risk being misused in ways that can lead to unethical or flawed conclusions. The inherent complexity of human behavior, influenced by a wide range of socio-cultural and emotional factors, cannot be easily reduced to the patterns that LLMs are trained to recognize.
The future use of LLMs in social science must be approached with caution. While these models can be valuable tools for analyzing large datasets, their limitations must be acknowledged, and they should not be relied upon as definitive sources of prediction in complex, human-centered fields. Researchers must take steps to mitigate the biases inherent in these models, and policymakers should be wary of relying too heavily on LLM-generated predictions, particularly when it comes to sensitive or ethically charged issues. LLMs, while powerful, are far from being a panacea for social science research, and their role in this field should be carefully scrutinized to prevent unintended harm.
References
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big?. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., … & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
Floridi, L., & Chiriatti, M. (2020). GPT-3: Its nature, scope, limits, and consequences. Minds and Machines, 30(4), 681–694.
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2020). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.
Green, D. P., Ha, S. E., & Bullock, J. G. (2010). Enough already about “black box” experiments: Studying mediation is more difficult than most scholars suppose. The Annals of the American Academy of Political and Social Science, 628(1), 200–208.
Marcus, G. (2020). The next decade in AI: Four steps towards robust artificial intelligence. AI Magazine, 41(3), 53–63.