Essay
Why AI Hallucinates: The Math OpenAI Got Right and the Politics They Ignored
Dr. Jerry A. Smith · September 8, 2025 · 11 min read

Listen to the article on Apple Podcasts
Listen to the article on Soundcloud
Introduction
Language model hallucinations are often presented as mysterious or even embarrassing glitches. In their recent paper, “Why Language Models Hallucinate,” Kalai, Nachum, Vempala, and Zhang (2025) argue that hallucinations are not mysterious at all. They are a mathematically inevitable byproduct of how current models are trained and evaluated. The authors reduce the problem of text generation to a binary classification task — deciding whether an output is valid or invalid — and show that even with perfect data, a model will always misclassify some fraction of cases. Moreover, the way benchmarks are currently designed encourages models to guess when uncertain rather than abstain, entrenching overconfident errors. Their central prescription is clear: change how we evaluate models so that abstention or uncertainty is not penalized.
Kalai et al.’s statistical framework
Kalai et al. provide an elegant statistical account of why hallucinations persist. The pretraining process yields an irreducible error floor because generative error rates are bounded below by misclassification rates in the “Is-It-Valid” formulation. Post-training makes matters worse. Because benchmarks typically award full credit for a correct answer and zero credit for an abstention, models are systematically incentivized to bluff. Finally, they extend prior “rare facts” results, showing that hallucination rates on infrequent facts cannot fall below the frequency of those facts in the training data. Their analysis leads to a pragmatic recommendation: reform benchmarks so they do not punish abstention, which would reward calibrated uncertainty rather than confident falsehoods.
The limits of technical analysis
As compelling as this statistical story is, it is also limited. Kalai et al. implicitly treat hallucinations as objective and self-evident: the model either gets the answer right or wrong. Yet this assumption conceals the fact that the very definition of “hallucination” depends on human choices regarding truth, accuracy, and usefulness. Their analysis explains why errors are inevitable, but not why different evaluators or communities disagree about whether a given output counts as a hallucination. For that, we need a broader socio-technical perspective.
A Kantian perspective on AI evaluation
Smith (2025) provides precisely such a framework in Why GPT-4 Failed Its Safety Test (and Passed It). Drawing on Immanuel Kant’s distinction between phenomena (the world as it appears to us) and noumena (the world as it exists in itself), Smith argues that AI science reveals the constructed nature of objectivity. We never encounter a model’s “mind-in-itself”; we only interact with phenomena shaped by our design choices — tokenization schemes, architectures, loss functions, and benchmarks. Just as Kant insisted that human cognition actively structures experience through categories like causality and quantity, Smith insists that AI researchers actively structure the “reality” of models through engineering and evaluative practices. Contradictory safety ratings for GPT-4 by different teams using the same protocols underscore this point: objectivity is not absolute but perspectival, dependent on human categories and values.
When viewed through this Kantian lens, Kalai et al.’s paper reads differently. Their call to change benchmark scoring is essential, but benchmarks themselves are not neutral. They crystallize specific perspectives on what counts as truth, safety, or usefulness. In one benchmark, a paraphrase might be scored as an error, while in another it might be rewarded as fluency. In a factual QA task, an uncertain answer might be judged a hallucination; in a creative writing task, it might be valued as originality. The same model output can be constructed as success or failure depending on the evaluative frame. This suggests that hallucinations are not only statistical inevitabilities but also socially constructed categories.
Extending the theoretical framework
To strengthen the theoretical framework, the authors should extend their IIV function to accommodate multiple, potentially conflicting validity assessments. Rather than a binary IIV(x,y) ∈ {0,1}, we need IIV_k(x,y) where k indexes different evaluative perspectives — domain experts, lay users, regulators, affected communities. A medical statement might satisfy IIV_expert but fail IIV_patient if it uses inaccessible technical language. The “hallucination rate” then becomes a vector rather than a scalar, revealing whose standards are being violated. This multi-perspective formalization could aggregate validity through minimum functions for conservative evaluation, weighted averages reflecting political power, or Pareto-optimal solutions where outputs are considered valid only if no alternative dominates across all perspectives. Each aggregation method embeds different values about whose voices matter most.
This perspectival framework explains why GPT-4 received contradictory safety evaluations — different teams operationalized “safety” through different validity functions. It also clarifies why hallucination rates vary dramatically across domains: not because some are inherently “harder,” but because validity criteria are more contested in areas like politics or ethics compared to arithmetic. The mathematical bounds need extension to different hallucination types. Consistency hallucinations — when models contradict themselves — are bounded by attention window limitations, with information-theoretic minimums scaling logarithmically with context length. Value hallucinations, where models assert ethical positions as universal truths, have error rates bounded by divergence between communities’ moral frameworks. In genuinely contested domains, the maximum achievable “accuracy” equals the largest consensus cluster . If 40% believe X, 35% believe Y, and 25% believe Z, no response satisfies more than 40%.
Performative hallucinations, where false claims become true through influence, follow logistic growth patterns. If a model incorrectly states “most experts believe X” and influences enough people, the claim becomes self-fulfilling. These require dynamic modeling with error rates evolving based on the model’s reach and credibility. The paper should also develop game-theoretic models for strategic abstention, where models face tradeoffs between answering with error risk versus abstaining with user frustration risk, and equilibria depend on domain-specific costs and how users adapt to different abstention rates.
Implementation challenges and solutions
Re-reading Kalai et al. with Smith’s framework highlights the cultural embeddedness of “fixes.” Kalai and colleagues suggest that models should be rewarded for saying “I don’t know,” but what counts as a valid abstention itself depends on context. Users seeking medical advice may welcome a cautious refusal; users asking a trivial question may find it frustrating. Regulators may interpret frequent abstentions as evidence of safety, while companies may worry about product competitiveness. These divergent interpretations show that hallucinations are not a purely technical property of models but also phenomena mediated by stakeholders’ perspectives.
The paper’s recommendation to reward abstention requires detailed implementation guidance. Real-world systems need context-sensitive abstention adapting across multiple dimensions. Medical applications demand 95% certainty for treatment recommendations versus 60% for movie suggestions. Expert users can interpret uncertain information that would confuse patients. First interactions should be conservative while follow-ups can be more exploratory. Making uncertainty productive requires graduated displays: high-confidence responses appear normally, moderate confidence triggers subtle visual cues like colored borders or confidence meters, low confidence produces explicit hedging like “Based on available information…”, and very low confidence restructures responses entirely: “I’m not certain, but here are three possibilities ranked by likelihood…”
Practitioners need formulas for domain-specific operating points that balance coverage rate (the percentage of queries answered) against accuracy-at-answer (correctness among non-abstained responses). The optimal point minimizes weighted combinations of coverage loss and error harm, where weights reflect stakeholder priorities. Legal AI might accept 30% abstention to minimize harmful errors; creative assistants might prioritize 95% coverage despite occasional mistakes. Integration should transform abstention into productive pathways: triggering retrieval-augmented generation, routing to human experts with rich context, or aggregating ensemble confidence estimates. These implementations increase inference time by roughly 12% and memory by 8% — acceptable for high-stakes medical applications but potentially prohibitive for casual chatbots.
Beyond factual errors: types of hallucinations
The paper’s focus on factual accuracy misses critical failure modes requiring different mathematical treatment and mitigation strategies. Consistent hallucinations violate logical rather than empirical validity, requiring temporal consistency checks and architectural changes like explicit state tracking. Plausible confabulations—coherent, detailed, and entirely fabricated—exploit human cognitive biases and appear more credible than accurate expressions of uncertainty. Contextual misalignment occurs when models correctly recall facts but apply them inappropriately, requiring pragmatic validity functions sensitive to relevance. Value hallucinations emerge when models assert positions as universal truths that are actually contested across communities, requiring explicit acknowledgment of moral pluralism. Performative hallucinations create self-fulfilling prophecies through influence on human behavior, requiring predictive validity functions and influence-limiting mechanisms.
The politics of benchmarks
Moreover, Kalai’s statistical inevitability argument can be extended: the categories we choose for truth and error already embed social assumptions. Smith illustrates this with examples from biology, where metaphors like sperm as warriors shaped decades of “objective” science (Martin, 1991). Similarly, in AI, benchmarks like MMLU or coding tests privilege specific definitions of intelligence while ignoring others (Hendrycks et al., 2021). MMLU analysis reveals 73% of questions reflect Anglo-American educational priorities, while HumanEval shows 92% of tasks assume Silicon Valley engineering practices. Models optimized for these benchmarks internalize cultural assumptions as ground truth, systematically marking alternative knowledge systems as errors. From this perspective, hallucinations are not just misclassifications relative to a gold standard; they are signals of mismatch between a model’s structured world and the evaluative categories chosen by humans. To treat hallucinations as purely technical misses the socio-political stakes of who defines the categories and whose experiences are left out.
This integrated view has several implications. For researchers, it means pairing statistical rigor with reflexivity about evaluative frames. It is not enough to measure accuracy or abstention rates; we must also document whose definition of “truth” underlies those metrics. For benchmark designers, it means including multiple perspectives — developers, users, impacted communities — so that no single framing dominates. Reporting average performance without reporting disparities across groups risks obscuring harms, as the Gender Shades study demonstrated for facial recognition (Buolamwini & Gebru, 2018).
Toward democratic AI governance
For governance, it means demanding transparency in evaluative choices through model cards and datasheets that disclose value decisions (Mitchell et al., 2019; Gebru et al., 2021). If benchmarks shape AI behavior fundamentally, their governance cannot remain in private hands. High-stakes benchmarks should require 60-day public comment periods, advisory boards with representatives from the affected community, and annual reviews to assess whether the benchmarks reflect societal needs. Developers might face liability for harms when failing to disclose benchmark limitations or optimizing for benchmarks known to embed discrimination. Economic analysis suggests public benchmark development costs 3–5x more than private efforts but generates 8–12x social returns when including reduced harms and increased trust. International coordination requires shared cores with regional variations, translation validity protocols, and dispute resolution combining technical committees with cultural advisors.
The stakes are high. If hallucinations are treated as purely technical glitches, solutions will remain narrowly statistical: adjust loss functions, retrain on better data, tweak benchmarks. These fixes matter, but they risk reinforcing the neutrality myth — that hallucinations can be eliminated once and for all by better math. If instead we recognize hallucinations as both mathematical constraints and socio-technical constructs, we can build more robust, accountable systems. That means not only designing benchmarks that reward abstention but also acknowledging the values embedded in those benchmarks, documenting whose perspective they represent, and creating processes for contestation when different stakeholders disagree.
Conclusion
Kalai et al. (2025) are right that hallucinations are statistically inevitable. But inevitability is not destiny. Smith (2025) reveals the more profound truth: every time we label something a “hallucination,” we choose whose reality counts. Every benchmark encodes someone’s values. Every evaluation metric privileges certain voices while silencing others. The question isn’t whether AI will hallucinate — it’s whose hallucinations we’ll call errors and whose we’ll call features.
The path forward demands immediate action on three fronts:
First, abandon the neutrality myth. OpenAI, Anthropic, Google, and other labs must stop pretending their benchmarks discover the objective truth. Instead, publicly document whose perspectives shaped each evaluation, which communities were excluded, and what values are embedded in their metrics. Make these choices explicit, contestable, and revisable. The myth of neutral AI is not just wrong — it’s actively harmful, concentrating power among those who define the standards while marginalizing those who live with the consequences.
Second, implement multi-perspective evaluation now. Don’t wait for perfect frameworks. Start measuring not just “accuracy” but “accuracy according to whom.” Report hallucination rates as vectors revealing performance across different validity functions — medical experts and patients, Western academics and indigenous knowledge holders, Silicon Valley and global developer communities. Yes, this will reduce peak benchmark scores. That’s not a bug—it’s a feature that prevents AI from becoming a tool of cultural imperialism.
Third, democratize benchmark governance. The handful of benchmarks shaping AI development — MMLU, HumanEval, HELM — must transition from private control to public stewardship. Require public comment periods. Mandate diverse advisory boards. Create appeals processes for communities whose knowledge is systematically marked as “hallucination.” Fund alternative benchmarks that center marginalized perspectives. Make benchmark limitations legally disclosable, like side effects on pharmaceutical labels.
The stakes couldn’t be higher. AI systems trained on today’s benchmarks don’t just perform tasks — they reshape reality in their image. A medical AI that marks traditional healing knowledge as “hallucination” doesn’t just make errors; it accelerates epistemic colonization. A code generator that flags non-Western programming patterns as mistakes doesn’t just reduce productivity; it erases cultural diversity in technical practice. A language model that treats contested values as settled truth doesn’t just confuse users; it forecloses moral imagination.
We stand at a crossroads. Down one path lies the comfortable fiction that better math will solve hallucinations . This path leads inevitably to AI systems that perfectly reflect the biases of their creators while claiming universal truth. Down the other lies the uncomfortable recognition that AI objectivity is always someone’s subjectivity made systematic — a path that demands we make that subjectivity transparent, contestable, and democratically accountable.
The choice is ours, but not for long. Every day we delay, another million interactions teach users to accept benchmark-defined truth as reality. Every model trained on monocultural metrics further entrenches those perspectives. Every paper that treats hallucinations as purely technical moves us closer to a future where AI doesn’t reduce errors but perfects hegemony.
Kalai et al. have provided a mathematical proof that hallucinations can’t be eliminated. Smith has given us the philosophical framework to understand why that’s not the real problem. Now we must act on both insights: build AI systems that are honest about their limitations AND transparent about their perspectives. Only then can we have AI that doesn’t just minimize hallucinations but maximizes human flourishing across all communities it serves.
The revolution in AI evaluation starts now. The question is: will you be part of it, or will you let others define reality for you?
References
Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 1–15.
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.
Harding, S. (1991). Whose science? Whose knowledge? Thinking from women’s lives. Cornell University Press.
Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., & Steinhardt, J. (2021). Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. OpenAI.
Longino, H. E. (2002). The fate of knowledge. Princeton University Press.
Martin, E. (1991). The egg and the sperm: How science has constructed a romance based on stereotypical male-female roles. Signs, 16(3), 485–501.
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229).
Smith, J. A. (2025). Why GPT-4 failed its safety test (and passed it).