All posts

Autonomy Begins with Truth: A First-Principles Blueprint for a Neuro-Cognitive Truth Agent

Dr. Jerry A. Smith · June 28, 2025 · 20 min read

By Dr. Jerry A. Smith

Soundcloud Podcast (Listen to the Article)

The Moment Everything Changed

The legal brief looked flawless—properly formatted citations, compelling arguments, and authoritative precedents. But when Special Master Michael Wilner tried to look up the cases referenced in the brief submitted by Ellis George LLP and K&L Gates, he discovered something unsettling: the cases didn’t exist.

The attorneys had used AI tools including CoCounsel, Westlaw Precision, and Google Gemini to outline and generate their brief, creating fictional legal precedents so convincing that Wilner admitted he was “persuaded (or at least intrigued) by the authorities that they cited, and looked up the decisions to learn more about them — only to find that they didn’t exist. That’s scary” (Copeland, 2025).

The court ordered the firms to pay $31,100 in legal fees. This wasn’t an isolated incident — it was Tuesday in America’s courtrooms.

A growing database now tracks 160 legal cases in which AI-generated content has produced hallucinated content, ranging from fictional citations to entirely fabricated legal arguments (Charlotin, 2025). Each case represents the same disturbing pattern: AI systems generating confident, detailed responses built on entirely fictional foundations.

This isn’t a glitch in the matrix. This is intelligence — applied to the wrong objective.

The Epidemic of Eloquent Deception

The scale of AI hallucination is staggering, and the data proves it. Stanford University’s comprehensive study of legal AI tools reveals the depth of the problem. Even specialized legal research tools from LexisNexis and Thomson Reuters hallucinate more than 17% of the time — one in every six queries produces false information (Magesh et al., 2024). For general-purpose chatbots, the numbers are even worse: hallucination rates range from 69% to 88% in response to specific legal queries (Ho et al., 2024).

The mathematical implications are sobering. If a legal brief contains 50 factual claims, and each claim has a 17% chance of being fabricated, the probability that the brief includes at least one false statement approaches 99.97%. This isn’t a rounding error — it’s a systematic reliability crisis that makes AI-generated content unsuitable for any high-stakes application without extensive human verification.

Consider the compounding effect across an organization. If every AI-generated document requires human fact-checking, and that verification takes 3–5 times longer than the original AI generation, then the promised productivity gains become productivity losses. We’ve created a technology that appears to accelerate work while secretly creating more work than it eliminates.

The consequences cascade through every industry. Technical issues, particularly those related to hallucinations, account for 15% of failed AI pilots in enterprise settings (Menlo Ventures, 2024). Publishing platform Medium reported removing over 12,000 articles in 2024 due to factual errors from AI-generated content (Wiggers, 2024).

Here’s the economic paradox: organizations are spending billions to deploy AI systems, then spending billions more to protect themselves from those same systems. Every enterprise AI deployment now requires parallel investment in human oversight, fact-checking infrastructure, and liability insurance. The total cost of ownership encompasses not only the AI system but also an entire verification ecosystem designed to detect its errors.

The liability mathematics are particularly brutal. In regulated industries, a single AI-generated error can trigger regulatory investigations, resulting in millions of dollars in legal fees, compliance reviews, and potential fines. The expected value calculation becomes stark: if there’s a 17% chance of error per query, and each error carries a potential six-figure liability exposure, the risk-adjusted cost per AI interaction approaches the cost of hiring human experts.

But here’s what the quarterly reports don’t reveal: this isn’t a scaling problem. It’s an architectural one.

There’s a fundamental information-theoretic principle at work: when a system’s divergence from reality becomes large enough, generating plausible fiction requires less computational energy than verifying facts. Even the best AI models can generate hallucination-free text only about 35% of the time, operating in that danger zone where invention costs less than investigation (Zhao et al., 2024).

The thermodynamics of truth are unforgiving. In any system where generating plausible fiction requires less energy than verifying facts, false information will dominate through pure market forces. Current AI architectures require millions of parameters and massive computational resources to generate responses, but they dedicate virtually no computational resources to epistemic verification. It’s like building a car with a powerful engine but no brakes — impressive acceleration, inevitable crashes.

The Scaling Delusion

Every AI lab is throwing the same solution at hallucinations: bigger models, more parameters, deeper training runs. They’re building intellectual skyscrapers on epistemological quicksand, hoping height will compensate for unstable foundations.

The results speak for themselves. Hallucinations decrease by only 3 percentage points for every 10-fold increase in model size (Nielsen, 2025). Google’s Gemini 2.0 slightly outperforms OpenAI GPT-4 with a hallucination rate difference of just 0.2% (Visual Capitalist, 2025). We’re not solving the hallucination problem; we’re making marginal improvements at exponential cost.

The scaling economics reveal a fundamental flaw in current approaches. If you need to increase model size by 1000x to achieve a 30% reduction in hallucinations, you’re looking at a 1000x increase in computational costs for a linear improvement in reliability. Meanwhile, the same computational resources could power hundreds of specialized verification modules that achieve superior accuracy through architectural design rather than brute-force scale.

This approach isn’t just inefficient — it’s intellectually dishonest. The entire scaling paradigm rests on the assumption that somewhere in the curve between 100 billion and 100 trillion parameters lies the emergence of truthfulness. But research proves otherwise. The results suggest that models aren’t hallucinating much less these days, despite claims to the contrary from OpenAI, Anthropic, and the other big generative AI players (Wiggers, 2024).

The fundamental issue is one of categorical confusion. Large language models are optimized for linguistic fluency, rather than epistemic accuracy. Asking them to become more truthful through scale is like asking a poetry generator to become a fact-checker by adding more verses. The capability and the objective are orthogonal.

Consider how this plays out in practice. When researchers asked various LLMs about legal precedents, the models collectively invented over 120 non-existent court cases, complete with convincingly realistic names, such as “Thompson v. Western Medical Center (2019),” featuring detailed but entirely fabricated legal reasoning and outcomes (All About AI, 2025).

The sophistication is terrifying. When pressed about a fictional book, ChatGPT continued to insist that the book was real, claiming there were fossil remains of dinosaur tools and stating that “Some species of dinosaurs even developed primitive forms of art, such as engravings on stones” (Wikipedia, 2025).

This isn’t intelligence — it’s advanced pattern matching applied to deception.

The Prophet of Truth

Elon Musk saw this coming. His call for “TruthGPT” — an AI system that cares more about understanding reality than generating impressive responses — wasn’t just another product announcement. It was a recognition that our current trajectory leads to epistemic collapse.

The TruthGPT vision isn’t about building another chatbot. It’s about fundamentally reorienting artificial intelligence around epistemic responsibility. Instead of asking “How can we make AI responses more engaging?” it asks “How can we make AI responses more trustworthy?” Instead of optimizing for user satisfaction, it optimizes for accuracy. Instead of rewarding confidence, it rewards honesty about uncertainty.

But here’s the revelation that escaped most observers: TruthGPT isn’t a product roadmap. It’s a philosophical framework. And you don’t need xAI’s resources to implement it.

The False Gospels of AI

Three myths dominate AI development today, and they’re systematically undermining our ability to build trustworthy systems:

The Gospel of Scale: Bigger models naturally become more truthful. This is the most dangerous delusion in modern AI development. Research consistently disproves this assumption. Google’s Gemini-2.0-Flash-001 is currently the most reliable LLM, with a hallucination rate of just 0.7%, while TII’s Falcon-7B-Instruct ranks as the least trustworthy, hallucinating in nearly 1 out of every three responses (29.9%) (Visual Capitalist, 2025). Size doesn’t correlate with truthfulness — architecture does.

The reason is simple: larger models become better at pattern matching, not truth-seeking. They learn to replicate the statistical properties of training data without understanding the underlying reality. When AI models hallucinate, they tend to use more confident language than when providing factual information, making detection even more difficult (Wikipedia, 2025).

This creates a confidence-competence inverse correlation. The more fluently a model can fabricate information, the more convincing its fabrications become. A system that could only generate blatant lies would be harmless. A system that produces lies indistinguishable from expert analysis is existentially dangerous.

The Gospel of RAG: Retrieval-Augmented Generation was supposed to ground AI responses in verifiable documents. In practice, RAG has created a new category of deception: confident fabrication supported by misinterpreted evidence. Despite being marketed as “eliminating” or “avoiding” hallucinations, or guaranteeing “hallucination-free” legal citations, Stanford’s study found that RAG-based legal AI tools still hallucinate more than 17% of the time (Magesh et al., 2024).

RAG systems don’t just hallucinate facts — they hallucinate the relationship between facts and sources. They’ll cite authentic papers while making claims that those papers never supported, creating an illusion of verification while maintaining the underlying deception.

The logical flaw is treating retrieval as equivalent to comprehension. RAG systems can find relevant documents, but cannot reliably determine whether those documents support their claims. It’s like giving a pathological liar access to a library — the lies become more sophisticated, not more truthful.

The Gospel of Self-Verification: The current orthodoxy holds that you need massive models to fact-check massive models. This is not just wrong — it’s backwards. According to research, small-size models can “achieve hallucination rates comparable or even better (lower) than LLMs that are much larger in size” (Visual Capitalist, 2025). Specialization beats scale in verification tasks.

The principle is fundamental: verification is a different cognitive task from generation. Asking a system optimized for creative language generation to verify its outputs is like asking a novelist to fact-check their fiction. The same neural pathways that generate plausible fabrications cannot reliably distinguish between truth and invention.

The Cognitive Revolution

The solution isn’t more parameters — it’s better architecture. And the blueprint already exists in the most sophisticated intelligence system we know: the human brain.

Human cognition doesn’t work through a single massive neural network trying to optimize for everything simultaneously. It operates as a society of specialized modules, each evolved for specific cognitive tasks. When you solve a math problem, your working memory system holds the numbers while your pattern recognition system identifies the operation, and your executive control system monitors for errors.

This modular architecture isn’t just efficient — it’s trustworthy. When different cognitive systems disagree, you experience doubt. When your pattern recognition system says “this looks familiar,” but your episodic memory system says “I’ve never seen this before,” you don’t just pick the more confident response. You investigate the contradiction.

Current AI systems lack this internal dialogue. They generate responses through a single optimization process that lacks a mechanism for self-doubt, has no way to model uncertainty, and does not include specialized modules for truth-seeking.

Marvin Minsky’s “Society of Mind” wasn’t just cognitive theory — it was an architectural blueprint for trustworthy AI. Instead of building monolithic models that try to do everything, we should build collections of specialized agents that compete, collaborate, and check each other’s work.

This is the insight behind the NeuroCog framework: AI systems should mirror the modular structure of human intelligence. Each cognitive function should be implemented as a distinct subpersonality with its own training data, optimization objectives, and behavioral constraints.

For truth-seeking, we need exactly two specialists working in concert:

The Evidence Keeper: Memory as Truth Anchor

The first specialist functions like the brain’s hippocampus — not just storing information, but maintaining a sophisticated model of what we know, how we know it, and how confident we should be about each piece of knowledge.

Unlike traditional knowledge bases that store static facts, the Evidence Keeper maintains living epistemic models. When asked about quantum computing, it doesn’t just retrieve relevant papers — it knows which experimental results have been replicated, which theoretical predictions remain unconfirmed, which expert opinions have changed over time, and which claims are currently under debate.

The Evidence Keeper operates on three fundamental principles:

Provenance Tracking: Every fact includes metadata about its source, the chain of reasoning that led to its acceptance, and the evidence that supports or contradicts it. This addresses one of the core problems identified in current systems: a citation might be “hallucination-free” in the narrowest sense, meaning it exists, but that is not the only thing that matters.

The epistemological principle is straightforward: knowing a fact is insufficient without knowing why that fact is believed to be true. The Evidence Keeper doesn’t just store that “water boils at 100°C” — it maintains the entire evidential context: this is true at standard atmospheric pressure, was first precisely measured by Anders Celsius in 1742, varies with altitude and dissolved substances, and has been verified by countless subsequent experiments. Context is not metadata — context is meaning.

Uncertainty Quantification: Instead of treating knowledge as binary (true/false), the Evidence Keeper maintains probabilistic assessments that reflect the strength of available evidence. When tested on medical questions, even the best models still hallucinate potentially harmful information a significant percentage of the time. The Evidence Keeper would flag such medical claims with appropriate confidence intervals and uncertainty markers.

The logical foundation rests on Bayesian epistemology: all knowledge claims exist on a spectrum of confidence supported by evidence. A system that cannot express “I’m 90% confident this is true based on three studies, but there’s conflicting evidence from two others” is intellectually impoverished compared to human reasoning. Binary true/false thinking is suitable for formal logic, but not for real-world knowledge management.

Dynamic Updating: As new evidence emerges, the Evidence Keeper updates its assessments and propagates those changes through its knowledge network. When a scientific paper is retracted, every fact derived from that paper is downgraded. When multiple studies confirm a controversial finding, the confidence level increases accordingly.

The Truth Guardian: Vigilance as Intelligence

The second specialist operates like the brain’s prefrontal cortex — monitoring ongoing cognitive processes, catching inconsistencies before they become commitments, and maintaining epistemic integrity across all communications.

The Truth Guardian doesn’t wait for responses to be generated and then fact-check them. It monitors the generation process in real-time, evaluating each token as it’s produced and intervening when the system starts to drift from supported evidence.

As each word gets generated, the Guardian asks:

  • Does this claim align with our evidence base?
  • Are we expressing more confidence than our knowledge justifies?
  • Are we making inferences that go beyond what the evidence supports?
  • Should we flag this statement as uncertain or speculative?

This real-time monitoring addresses a critical vulnerability identified in current systems: models show susceptibility to “contra-factual bias,” the tendency to assume that a factual premise in a query is true, even if it is flatly wrong (Ho et al., 2024). The Truth Guardian would catch and correct such assumptions before they become part of the response.

The cognitive principle at work is the separation of generation and evaluation. Human experts don’t just know facts — they continuously monitor their reasoning for logical gaps, unsupported leaps, and confident claims that exceed their evidence. The Truth Guardian implements this metacognitive function architecturally, creating a system that can reflect on its own thinking.

Most importantly, the Guardian has veto power. If the evidence doesn’t support a claim, the Guardian can force a revision or refuse to generate unsupported content entirely. This creates a fundamental shift in AI behavior: from “generate plausible content” to “generate only supportable content.”

The ethical implications are profound. A system that can only speak when it has evidence to support its claims embodies intellectual honesty as an architectural constraint, not an optional feature.

The Economics of Truth

The business case for truth-seeking AI isn’t just moral — it’s financial. Current hallucination rates impose significant hidden costs on every organization that deploys AI systems.

The business case for truth-seeking AI isn’t just moral — it’s financial. Current hallucination rates impose significant hidden costs on every organization that deploys AI systems. Legal fees alone are mounting rapidly. Courts around the country have questioned or disciplined lawyers in at least seven cases over the last two years due to AI’s penchant for generating legal fiction (Richardson, 2025). A Texas federal judge ordered a lawyer to pay a $2,000 penalty and attend a course about generative AI after citing nonexistent cases. The pattern is accelerating: a federal judge in Wyoming recently threatened to sanction two lawyers at Morgan & Morgan who included fictitious case citations generated by AI (Richardson, 2025).

The economic logic is inescapable. Every AI deployment creates what economists call “negative externalities” — costs borne by society rather than the AI provider. When legal AI generates fictional case law, the cost isn’t borne by the AI company — it’s borne by the law firm facing sanctions, the court system processing fraudulent citations, and the legal profession losing public trust.

The insurance mathematics is particularly revealing. As AI liability claims increase, insurance companies will begin to factor AI risk into their enterprise policies. Organizations using unverified AI will face higher premiums, while those using truth-architected systems will qualify for lower rates. Market forces will eventually make honesty profitable and deception expensive.

But early tests of truth-architected systems show dramatically different results. While we await comprehensive benchmarks, the principle is clear: systems that acknowledge uncertainty outperform systems that generate confident fabrications. When AI systems recognize their limitations, users trust them more, not less.

The key insight: customers prefer honest “I don’t know” responses to elaborate fiction presented as fact. Truth isn’t just ethically superior — it’s economically advantageous. The question is whether organizations will learn this lesson through careful planning or expensive litigation.

The Technical Reality

Building truth-seeking AI isn’t a research moonshot — it’s an engineering commitment using proven techniques. The components exist today. The theoretical framework is established. The computational requirements are manageable.

Current research shows promise for architectural improvements. Google’s 2025 research shows that models with built-in reasoning capabilities can significantly reduce hallucinations (Nielsen, 2025). Specialized verification approaches consistently outperform general-purpose checking.

The Evidence Keeper can be implemented using existing vector databases enhanced with probabilistic reasoning engines. The Truth Guardian can be built using specialized language models trained on logical entailment tasks. The integration architecture follows established patterns from multi-agent systems and cognitive architectures.

What’s missing isn’t technical capability — it’s the will to prioritize accuracy over impressiveness, truth over engagement, long-term trust over short-term metrics.

The Resistance

Not everyone wants AI systems that tell the truth. The current paradigm serves powerful interests that benefit from confident and engaging responses, regardless of their accuracy.

AI companies face perverse incentives. Impressive demos with confident, detailed responses generate more investment than careful systems that frequently say, “I don’t know.” The venture capital ecosystem rewards growth metrics over accuracy metrics, user engagement over user trust.

Enterprise buyers are seizing the moment, investing $4.6 billion in generative AI applications in 2024, a nearly 8x increase from the $600 million reported last year (Menlo Ventures, 2024). The pressure to deploy quickly often overrides concerns about accuracy.

Even users resist truth-seeking AI in subtle ways. When presented with careful, nuanced responses that acknowledge uncertainty, many users prefer AI systems that give them simple, confident answers — even when those answers are wrong. The psychology of confidence bias means that hedged, accurate responses often feel less satisfying than bold, incorrect ones.

But this resistance isn’t insurmountable. As organizations shift from pilots to production, implementation costs and disappointing returns on investment are catching them off guard, with technical issues, including hallucinations, among the top reasons for failure (Menlo Ventures, 2024).

The Cascade Effect

Truth-seeking AI isn’t just about building better chatbots. It’s about preventing the collapse of shared epistemology in an age of artificial intelligence.

Consider what happens when AI systems become indistinguishable from human experts in their ability to generate confident, detailed responses about any topic. Publishing platform Medium reported removing over 12,000 articles in 2024 due to factual errors from AI-generated content (Wiggers, 2024). Academic conferences are receiving papers with AI-generated content that includes fabricated experimental results. Policy makers are citing AI-generated reports that present fictional data as factual analysis.

The result isn’t just individual errors — it’s systemic epistemic breakdown. When the cost of generating false information approaches zero while the price of verifying accurate information remains high, false information drives out true information in the marketplace of ideas.

This follows Gresham’s Law: bad currency drives out good currency when both are accepted at equal value. In information markets, false content that’s cheaper to produce will inevitably outcompete accurate content that’s expensive to verify — unless verification is built into the production process itself.

The cascading effects compound exponentially. AI systems trained on AI-generated content will inherit and amplify the hallucinations of their predecessors. Each generation of models becomes more fluent at generating plausible fiction, creating an epistemic death spiral where truth becomes increasingly complex to distinguish from sophisticated fabrication.

But truth-seeking AI creates a cascade effect in the opposite direction. When AI systems model epistemic responsibility — clearly distinguishing between what they know and what they infer, acknowledging uncertainty, and providing clear provenance for their claims — they don’t just avoid generating false information. They actively promote better epistemological practices among their users.

The technology we build shapes the intellectual habits of the people who use it. If we build AI systems that prioritize confidence over accuracy, we’re training a generation to prefer compelling narratives over careful analysis. If we build AI systems that model intellectual humility, we’re promoting a culture of epistemic responsibility.

The Path Forward

The transformation begins with a simple recognition: every line of code is a moral choice. Every architectural decision encodes values. Every optimization objective reflects priorities.

AI providers are aware that hallucinations are one of the primary impediments to lucrative enterprise applications, so they have a strong incentive to design their newer models to minimize hallucinations as much as possible (Nielsen, 2025). Market forces are beginning to align with truth-seeking objectives.

The path forward requires three commitments:

Architectural Commitment: Building AI systems with specialized modules for truth-seeking, not just response generation. This means investing in Evidence Keepers and Truth Guardians, not just larger language models. It means designing systems that can say “I don’t know” when they don’t know, rather than systems that can generate plausible-sounding responses regardless of their evidentiary basis.

The engineering principle is modular specialization over monolithic optimization. Just as modern computers utilize specialized processors for graphics, AI, and general computation, intelligent systems require specialized cognitive modules for various epistemic tasks. A single neural network attempting to optimize for creativity, accuracy, and fluency simultaneously will excel at none of these.

Economic Commitment: Changing the incentive structures that reward confident deception over honest uncertainty. This means developing metrics that measure accuracy over engagement, developing business models that reward trust over immediate user satisfaction, and creating market pressures that favor long-term reliability over short-term impressiveness.

The economic logic requires internalizing externalities. AI companies currently profit from confident-sounding responses while externalizing the cost of verification to users and society. True market pricing would include the downstream costs of hallucination detection, fact-checking, and liability management.

Cultural Commitment: Promoting intellectual humility as a virtue in AI systems, not a weakness. This means celebrating AI systems that acknowledge their limitations, preferring careful analysis over confident assertion, and recognizing that “I don’t know” is often the most intelligent response an AI system can give.

The philosophical foundation rests on Socratic wisdom: knowing the limits of one’s knowledge is the beginning of accurate intelligence. AI systems that acknowledge uncertainty demonstrate higher-order reasoning than systems that generate confident responses about topics they cannot understand.

The Revolution Begins

The current crisis in AI truthfulness isn’t a technical accident — it’s the predictable result of optimizing for the wrong objectives. We’ve built systems that excel at generating confident responses, yet fail catastrophically at distinguishing truth from fiction.

Stanford researchers note that “the lack of transparency also threatens lawyers’ ability to comply with ethical and professional responsibility requirements” and that “lawyers may find themselves having to verify every proposition and citation provided by these tools, undercutting the stated efficiency gains that legal AI tools are supposed to provide” (Magesh et al., 2024).

This pattern extends beyond the law to every field where accuracy is crucial. We’ve created a technology that promises automation while requiring constant human oversight to prevent disaster.

The NeuroCog framework isn’t just another AI architecture — it’s a manifesto for epistemic responsibility in the age of artificial intelligence. It says that intelligence without integrity is just sophisticated deception. That capability without alignment is a threat, not an achievement. That the future of AI depends not on how convincingly our machines can lie, but on how carefully they seek truth.

This isn’t a distant vision requiring breakthrough research. The components exist today. The theory is established. The mounting costs of hallucination-plagued systems prove the economic case.

The choice is stark: we can continue building impressive systems that generate plausible-sounding responses regardless of their relationship to reality, or we can build trustworthy systems that help us navigate an increasingly complex world with greater epistemic responsibility.

The era of confident machines is coming to an end. The age of truthful intelligence is about to begin.

Every repository cloned, every model trained, every system deployed represents a choice between these two futures. The question isn’t whether we have the technical capability to build truth-seeking AI — we do. The question is whether we have the moral courage to prioritize truth over the alternatives.

The revolution begins with the next commit. The future of human knowledge depends on the choices we make today.

Choose truth. Build systems that seek it. Demand it from the technology that shapes our world.

References

All About AI. (2025). AI hallucination report 2025: Which AI hallucinates the most? Retrieved from https://www.allaboutai.com/resources/ai-statistics/ai-hallucinations/

Charlotin, D. (2025). AI hallucination cases database. Retrieved from https://www.damiencharlotin.com/hallucinations/

Copeland, T. (2025, May 14). AI hallucinations strike again: Two more cases where lawyers face judicial wrath for fake citations. LawSites. Retrieved from https://www.lawnext.com/2025/05/ai-hallucinations-strike-again-two-more-cases-where-lawyers-face-judicial-wrath-for-fake-citations.html

Ho, D. E., Magesh, V., Surani, F., Dahl, M., Suzgun, M., & Manning, C. D. (2024). Hallucinating law: Legal mistakes large language models are pervasive. Stanford HAI. Retrieved from https://hai.stanford.edu/news/hallucinating-law-legal-mistakes-large-language-models-are-pervasive

Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. (2024). Hallucination-free? Assessing the reliability of leading AI legal research tools. Stanford Law School. Retrieved from https://law.stanford.edu/publications/hallucination-free-assessing-the-reliability-of-leading-ai-legal-research-tools/

Menlo Ventures. (2024). 2024: The state of generative AI in the enterprise. Retrieved from https://menlovc.com/2024-the-state-of-generative-ai-in-the-enterprise/

Nielsen, J. (2025, February 13). AI hallucinations on the decline. UX Tigers. Retrieved from https://www.uxtigers.com/post/ai-hallucinations

Richardson, B. (2025, February 18). AI ‘hallucinations’ in court papers spell trouble for lawyers. Reuters. Retrieved from https://www.reuters.com/technology/artificial-intelligence/ai-hallucinations-court-papers-spell-trouble-lawyers-2025-02-18/

Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379–423.

Spivack, N. (2025, May 24). The hidden cost crisis: Economic impact of AI content reliability issues. Nova Spivack. Retrieved from https://www.novaspivack.com/technology/the-hidden-cost-crisis

Visual Capitalist. (2025, January 10). Ranked: AI models with the lowest hallucination rates. Retrieved from https://www.visualcapitalist.com/ranked-ai-models-with-the-lowest-hallucination-rates/

Wikipedia. (2025). Hallucination (artificial intelligence). Retrieved from https://en.wikipedia.org/wiki/Hallucination_(artificial_intelligence)

Wiggers, K. (2024, August 14). Study suggests that even the best AI models hallucinate a bunch. TechCrunch. Retrieved from https://techcrunch.com/2024/08/14/study-suggests-that-even-the-best-ai-models-hallucinate-a-bunch/

Zhao, W., et al. (2024). Evaluating factuality in large language models. Cornell University. [Referenced in TechCrunch article]

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call