Essay
When All Your AI Agents Are Wrong Together
Dr. Jerry A. Smith · November 24, 2025 · 9 min read

Listen to the article on Apple Podcasts
Listen to the article on Soundcloud
Recent work on massively decomposed agentic processes — most notably the MAKER system used to complete over one million sequential reasoning steps without error — demonstrates that long-horizon reliability in LLM-based agents is achievable (Meyerson et al., 2025). However, MAKER’s success relies on extreme specialization of the problem domain, a carefully engineered deterministic state representation, and heavy parallel sampling with probabilistic voting.
The approach is powerful, but it exposes several structural weaknesses that limit its applicability in real-world environments where ambiguity, uncertainty, partial observability, and correlated errors are common.
This article proposes a new architecture, TAC-HAVA-K, which combines three key innovations:
- TAC (Thesis, Antithesis, Consolidator): adversarial reasoning between independent agents
- HAVA (Hierarchical Active Verification): belief states, simulators, and multi-channel verification
- K-fold parallelism: diversity on both the Worker and Devil’s Advocate sides
This combined system resolves MAKER’s three core flaws while preserving its ability to reach million-step reliability. It supports messy, partially observed settings, provides structural error correction through adversarial reasoning, and reduces cost through hierarchical control.
Why Parallel Voting Works: The Mathematics of Logarithmic Efficiency
Before examining what breaks, we need to understand why MAKER’s voting approach is so powerful when it works (Meyerson et al., 2025).
At each micro-step, let p represent the probability that a single LLM sample gives the correct action, with q = 1 − p representing the error probability. MAKER draws samples until one candidate action is ahead of every other by at least k votes — the “first-to-ahead-by-k” rule.
The key theoretical result is elegant: the probability of selecting the wrong action decreases exponentially in k, while the number of samples needed grows only linearly with k. This means reliability scales exponentially while cost scales only logarithmically.
Picture this as a biased random walk. Let X represent the vote lead for the correct action after n samples. When a correct sample arrives, X increases by 1 with probability p; when wrong, X decreases by 1 with probability q. We stop when X reaches +k (correct wins) or −k (incorrect wins).
For a biased random walk with p > q, the probability of error is approximately (q/p)^k — exponential in −k.
Consider a concrete example with step-level success rate p = 0.6. The ratio q/p equals roughly 0.67. At k = 10, the error probability drops to about 1.7%. At k = 20, it falls to 0.03%. At k = 30, we reach approximately 0.002%.
For million-step reliability — where we want the probability of zero errors across 1,000,000 steps to exceed 90% — we need per-step error rates around 10⁻⁷. This is achievable with modest accuracy (p = 0.6–0.7) and k values in the 5–12 range.
This is why MAKER scales: reliability needs grow exponentially, while costs grow only logarithmically. The Towers of Hanoi problem, recently introduced as a benchmark for investigating LLM reasoning limitations (Shojaee et al., 2025), provides an ideal testbed because the problem scales naturally to enormous numbers of required steps — solving with twenty disks requires just over one million steps.
The Three Core Flaws in the MAKER Approach
MAKER provides an unprecedented demonstration of long-horizon stability (Meyerson et al., 2025), but its own architecture reveals critical limitations that constrain generalization.

Flaw 1: Reliance on Uncorrelated Error Distributions
The mathematics only works when errors are mostly independent across samples. MAKER assumes that LLM sampling noise behaves like independent random trials — essential for voting to wash out mistakes.
However, real LLMs often fail in correlated modes. When a prompt induces a systematic misunderstanding, when a reasoning pattern collapses into a wrong schema, when a particular error is frequent in the model’s internal distribution, or when a domain is inherently ambiguous — in all these cases, most samples make the same wrong mistake.
If errors are correlated, voting doesn’t fix them. Increasing k doesn’t fix them. Samples converge to the wrong answer faster.
In random-walk terms, the bias p > q simply disappears. The “attractor” mistakes — those extremely stable errors representing repeated conceptual flaws — dominate the distribution. If all samples share the same wrong mental model, voting cannot help.
Flaw 2: Requirement for Fully Explicit, Unambiguous State
MAKER requires that the precise environment state at every step is represented in canonical, serialized form. The state must be fully known and unambiguous. The following action must be describable entirely from that state. Correctness must be definable and checkable without ambiguity.
This limits MAKER to deterministic, fully observable, crisp state machines.
In real tasks, information may be incomplete, observations may be noisy, some states may be latent and not directly observable, multiple interpretations may be plausible, and actions may have uncertain or delayed consequences.
Voting requires that every sample sees the same complete state with no ambiguity and no hidden variables. Under partial observability, parallel samples may yield different, yet equally plausible, interpretations. Instead of converging on a single best answer, the sample distribution becomes multimodal, inconsistent, and unstable.
MAKER’s approach becomes brittle or inapplicable when the world cannot be cleanly serialized.
Flaw 3: High Per-Step Cost and Scaling Limitations
MAKER incurs cost because every step must be sampled multiple times. Voting requires multiple model calls, sometimes dozens. Verification relies on strict filtering, heuristics, and formatting constraints.
For a million steps, even small per-step overhead becomes enormous. While parallelism reduces wall-clock time, the token-level and call-level costs remain high. In many real-world tasks, performing heavy sampling per atomic step is impractical.
The TAC-HAVA-K Solution
Our architecture marries adversarial debate, hierarchical verification, and parallel epistemic diversity into a robust agentic system designed to withstand the three flaws above.

The TAC Triad: Thesis, Antithesis, Consolidator
The foundation is structured adversarial reasoning through three distinct roles.
Workers (Thesis Agents) produce proposed actions, explanations, plans, supporting evidence, assumptions, and confidence levels. Critically, Workers cannot communicate with each other or with Devil’s Advocates — they may only respond to Consolidator questions. This isolation preserves independence of thought.
Devil’s Advocates (Antithesis Agents) serve as systematic epistemic adversaries. They find flaws, identify logical failures, challenge assumptions, surface misinterpretations, highlight edge cases, and suggest alternative views. Like Workers, they cannot speak to Workers or influence each other directly — they interact only with the Consolidator.
The Consolidator (Synthesis Agent) sees everything. It can ask unlimited questions, interrogate Workers and Devil’s Advocates repeatedly, and request justification, clarifications, or corrections. However, it does not create new arguments — it adjudicates, synthesizes, and stabilizes. The Consolidator serves as the epistemic governor of the system.
K-Fold Parallelism: Best-of-K Selection
A single Worker and single Devil’s Advocate provide basic adversarial robustness, but diversity remains limited. By introducing K Workers and K Devil’s Advocates, we expand the diversity of reasoning while controlling costs.
K Workers provide independent solution proposals, diverse reasoning styles, different heuristics and biases, and alternative pathways toward the same goal. This directly combats correlated reasoning patterns.
K Devil’s Advocates provide independent critiques, varied adversarial perspectives, different attack strategies, and broad coverage of potential error modes. This radically suppresses undetected systematic failures.
After interrogation, the Consolidator selects the strongest Worker proposal paired with the most compelling Devil’s critiques — the most consistent and plausible combination. This selection mechanism is not a vote but a qualitative assessment based on reasoning depth, consistency, and resilience under questioning.
HAVA Integration: Belief States, Simulation, and Verification
To handle MAKER’s second and third flaws, we add the HAVA layer.
The Belief-State Engine solves the explicit state requirement. Instead of requiring perfect state knowledge, the system maintains a belief-state that represents distributions over possible world states, uncertainty models, and latent variables inferred from observations. Workers and Advocates reason about belief, not a perfect state. This generalizes the architecture to partial observability, ambiguous domains, and noisy data.
The World Model provides a learned or rule-based simulator that predicts outcomes of proposed actions. The Consolidator compares Worker reasoning, Advocate critiques, world model predictions, and belief-state constraints. This enables consequence-checking, reality anchoring, and detection of inconsistencies.
The Verifier checks logical invariants, safety rules, format correctness, and domain-specific constraints through symbolic and constraint-based methods. This ensures that even subtle errors are caught — using verification channels that don’t share LLM biases.
Checkpoint and Local Repair address the cost problem. Instead of requiring perfect forward progress, the system periodically checkpoints the belief-state, detects inconsistencies later in the steps, triggers rollback and regeneration, and applies stronger verification only when needed. Local repair prevents single-step mistakes from compounding over long horizons.
The Full Pipeline
At each iteration, the system executes a structured loop:
- Workers W₁ through Wₖ propose actions
- Devil’s Advocates D₁ through Dₖ critique all Worker outputs
- The Consolidator interrogates Workers and Devils through questioning
- The World Model simulates consequences of proposed actions
- The Verifier checks constraints
- The Consolidator synthesizes the final action
- The Belief-State updates based on observations
- Checkpoint and rollback logic maintains long-run integrity

This loop operates at both micro-level (atomic steps) and macro-level (hundreds of micro-steps grouped). Using macro control reduces cost drastically — heavy verification applies at macro boundaries, not every micro-step.
How TAC-HAVA-K Addresses the Three Flaws
Each component of the architecture targets a specific weakness in the MAKER approach. The adversarial structure breaks error correlation. The belief-state engine handles ambiguity. The hierarchical design controls cost. Here’s how they work together.
Solving Correlated Errors
Adversarial agents catch what sampling cannot. The Devil’s Advocate role is specifically designed to surface systematic failures that would survive voting. Multiple independent Workers and Devils further reduce correlation — different prompts, different reasoning styles, different attack strategies. Simulators and constraint checkers add verification channels that don’t share model biases at all.
Undetected errors become rare not through statistical independence, but through structural diversity of verification.
Handling Partial Observability
Belief-state reasoning combined with adversarial cross-examination and world model simulation allows operation in ambiguous problems, noisy environments, and real-world tasks where explicit state cannot be serialized. The system no longer asks “is this action correct for this ground-truth state?” but rather “is this action consistent with beliefs and constraints?”
MAKER cannot operate in these conditions. TAC-HAVA-K can.
Achieving Cost Efficiency
Debate is reserved for macro decisions. Light checks handle micro steps. Heavy verification triggers only on uncertainty or risk. Rollback reduces the need for perfect atomic accuracy.
The net effect: million-step operation at manageable cost through hierarchical control and selective verification.
Conclusion
MAKER showed the world that million-step error-free execution is possible. Its elegant mathematics — logarithmic cost scaling against exponential reliability growth — provided the theoretical foundation. But its core assumptions limit direct application to simple, fully-known state machines where errors are independent.
The TAC-HAVA-K architecture preserves MAKER’s reliability while solving its three central flaws. It achieves robustness to correlated errors through adversarial multi-agent reasoning, operates in partial observability using belief states and simulators, and reduces cost through hierarchical reasoning and selective verification.
By combining Workers, Devil’s Advocates, a powerful Consolidator, a belief-state engine, a world model, and constraint verification — plus the best-of-K diversity mechanism — this architecture offers a path toward robust, scalable, general-purpose long-horizon agents.
The future of reliable AI agents lies not in perfecting individual model calls, but in architecting systems where errors are detected, challenged, and repaired before they compound.
This article is part of ongoing research into agent reliability architectures. For technical implementation details or collaboration inquiries, please reach out.
References
Meyerson, E., Paolo, G., Dailey, R., Shahrzad, H., Francon, O., Hayes, C. F., Qiu, X., Hodjat, B., & Miikkulainen, R. (2025). Solving a million-step LLM task with zero errors. arXiv preprint arXiv:2511.09030. https://arxiv.org/abs/2511.09030
Shojaee, P., Mirzadeh, I., Alizadeh, K., Horton, M., Bengio, S., & Farajtabar, M. (2025). The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. arXiv preprint arXiv:2506.06941. https://arxiv.org/abs/2506.06941