All posts

The Exponential Stack: Why Your AI Still Feels Like a Chatbot

Dr. Jerry A. Smith · April 21, 2026 · 9 min read

Building Minds | Edition 20

A board meeting, a CTO, and the three rungs of agent topology — with receipts.

By Dr. Jerry A. Smith | April 2026


A board meeting. A CTO. A demo.

The VP of Engineering has done exactly what she was asked. An AI assistant in the sales stack. An AI summarizer in the legal stack. An AI drafter in the product stack. Three single agents, each well-prompted, each with access to the right tools. Each one, individually, works.

Then the CEO asks a question nobody rehearsed.

What has the AI learned about our company over the last six months?

The room goes quiet.

The honest answer is nothing. Three agents, six months of use, no retained institutional knowledge. The AI has been typing. It has not been thinking.

This is what happens when you buy intelligence and forget to buy anatomy.

The last two years of enterprise AI have been a category error. Organizations bought ChatGPT-shaped tools and expected organizational outcomes. They benchmarked models and ignored topology. They confused a brain with a company. The capability curve that matters in 2026 is not about model size. It is about agent topology — and topology grows capability exponentially, not linearly. Three rungs are now visible. The delta between them is measurable in Anthropic's own papers, in private-equity deal rooms, and in the gap between the firms winning with AI and the ones who still think they are.


Rung 1 — The Single Agent

One context window. One reasoning stream. One point of view. The ChatGPT your team opens every morning. The custom GPT with a persona. The Claude chat with your PDFs attached.

Rung 1 is not nothing. For bounded tasks — write this email, summarize this document, answer this question — it is remarkable. But the architecture has a ceiling, and the ceiling is easy to state.

A single agent's knowledge acquisition is serial. It asks one question at a time. It cannot hold two branches of a reasoning tree open at once. As context grows, it rots: recent work on context degradation shows up to a 40% performance loss due to bloated windows. Practitioners measure single generalist agents burning an extra $0.52 per message in inference waste once their permanent overhead passes 45,000 tokens, before the user types.

A single specialist — a sniper agent built for one job — outperforms that same generalist by 5–10× on its slice.

That is the Rung 1 ceiling in a single line: one generalist brain, no matter how good the model, will always be overkill for any given task and insufficient for the whole.

Rung 2 — The Agent with Subagents

One orchestrating mind. Many disposable hands. The parent splits the question, dispatches workers in parallel, and integrates returns.

This is the architecture Anthropic ships for their Research product, and the numbers they published in June 2025 are the cleanest signal in the public record:

A multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on Anthropic's internal research eval.

Ninety-point-two percent. Same model family. Same company. The only variable was topology.

Anthropic went further and decomposed the performance: three factors explain 95% of the variance on the BrowseComp benchmark, and token usage alone explains 80% of it. In plain English, what the subagent architecture buys is parallel token capacity against an exponentially branching search tree. N subagents cover N branches of the tree at once. The tree is exponential; coverage grows far faster than agent count.

This is the first non-trivial multiplier. Depth and breadth grow together because the parent can ask in many directions simultaneously without losing the thread.

Rung 3 — The Orchestrated Team

Peer agents with role, private memory, and persistent identity. Coordinated by protocol rather than by a single parent. The orchestrator is the chief of staff. The sub-agents are employees. The orchestrator never does execution work directly.

Rung 3 is where knowledge acquisition becomes compositional. Specialists build private expertise. Handoffs happen at structural interfaces, not conversational ones. Outputs emerge at the seams — analyses no single agent produced in isolation.

The canonical external receipt comes from MetaGPT. Role-based team (PM, architect, engineer, QA): 85.9% on HumanEval pass@1, against GPT-4 single-agent at 67.0% — a 28% relative lift (arXiv:2308.00352). The specialists were not bigger models. They were the same model class, in different roles, with structured handoffs.

The internal receipt is sharper. The patent-pending 5-3-1 architecture, deployed against private-equity Commercial Information Memoranda, yielded 8–12 findings per CIM through sequential single-agent analysis. The five-extractor / three-analyst / one-synthesizer topology, working on the exact same documents, produced 25–35 findings at the same quality bar. That is a 200% increase. Anchor drift collapsed from 20% to 2%. Analyst agreement rose from 65% to 92%.

Same documents. Same models. Different anatomy.


Why the curve is exponential, not additive

Three compounding effects are now well-documented. None is speculation.

Recommended by LinkedIn

[Empowering Our People: OneAI at Beyond ONE Empowering Our People: OneAI at Beyond ONE Beyond ONE

1 year ago](https://www.linkedin.com/pulse/empowering-our-people-oneai-beyond-one-beyond-1-cqgmf) [The Model That Refuses to Write: Jev, System One, and the Coming Decision Layer in Enterprise AIThe Model That Refuses to Write: Jev, System One, and… Padmini Soni

1 week ago](https://www.linkedin.com/pulse/model-refuses-write-jev-system-one-coming-decision-layer-padmini-soni-lgflc) [The Real AI Revolution: How Every Department is Being Reinvented from the Inside outThe Real AI Revolution: How Every Department is Being… Robin G.

11 months ago](https://www.linkedin.com/pulse/real-ai-revolution-how-every-department-being-from-inside-granberg-v1vcf)

Parallel coverage of an exponential tree. Search and reasoning trees branch exponentially. N parallel agents cover N branches at every step, so coverage grows as roughly b^d with parallelism, where it was linear in time with a single agent. Anthropic's own finding — research time cut by up to 90% on complex queries when subagents ran in parallel — is this effect measured? The 5-3-1 architecture makes the combinatorics explicit: five orthogonal extraction dimensions yield 2⁵ − 1 = 31 interaction patterns, whereas sequential analysis sees only five. Literal exponential-versus-linear, on the same input.

Specialization gain. A focused agent with the right tools and the right context outperforms a generalist on its slice. Stack specialists and the gains compound across the workflow. MetaGPT is the canonical external result. The internal version, if you have seen it, is the difference between a sales-enabled generalist and a sales-enabled team — a pricer, a researcher, a drafter. You do not need a better model. You need a different org chart.

Emergent integration at the seams. The most valuable outputs live at the interfaces between agents, not inside any one of them. On Rung 1, these interfaces do not exist. On Rung 2, they are shallow — subagents return, the parent synthesizes, done. On Rung 3, they are structural: the 5-3-1 design forces analysts to see only the anchors produced by extractors, never the original source. That constraint is what produces a synthesis no single agent had in its context window. Controlled studies of coordinated multi-agent systems show revision rates falling 74% and coherence rising 47% from this effect alone.


The honest ceiling

This is where most multi-agent pieces go soft. So here is the other side of the chart.

The same Google/MIT topology paper that shows centralized teams outperforming single-agents by +80.8% on parallelizable financial reasoning shows every multi-agent variant degrading performance by 39–70% on sequential planning tasks. Topology is not a free lunch. The wrong topology for the wrong task shape is worse than a single agent.

Uncoordinated parallel agents amplify errors 17.2×. Centralized coordination pulls that down to 4.4×. The orchestrator is not a performance knob. It is a safety mechanism.

Berkeley's MAST paper catalogued 14 distinct failure modes across 1,600+ multi-agent traces (arXiv:2503.13657) and concluded, bluntly, that performance gains over single-agent baselines are often minimal — with most failures rooted in design flaws, not model limitations.

And the sharpest natural experiment in the entire field: Anthropic itself ships single-agent architecture for Claude Code and multi-agent architecture for Research. Same company. Same model family. Opposite architectural choices. The Claude Code team reached the same conclusion Cognition published in their "Don't Build Multi-Agents" post — for coding tasks where subtasks share deep context, "running multiple agents in collaboration only results in fragile systems." The Research team, working against parallelizable search trees, reached the opposite conclusion and got +90.2%.

There is no dogma here. There is a question: what shape is your task?

Parallelizable, coverage-heavy, high-value, bounded context → climb to Rung 2, then Rung 3. Sequential, state-dependent, deeply coupled → stay on Rung 1 with better tools.

Anything else is theater.


What a CTO actually pays for

If you are buying AI, you are buying a topology, not a model. The frontier models are now roughly equivalent on capability benchmarks, and that gap is narrowing. The durable advantage lives in the cognitive architecture around the models — the memory systems, the coordination mechanisms, the governance layers.

Three questions to bring to your next vendor meeting.

Which rung is this? If the answer is Rung 1 — a model behind a prompt — you are buying a chatbot. Pay chatbot prices.

Does the architecture match the task shape? Parallelizable research, coverage-heavy analysis, portfolio-scale extraction: multi-agent wins in the +80% to +200% range. Sequential coding, state-heavy drafting, iterative negotiation: single-agent with good tools wins, and forcing multi-agent will amplify errors 17×.

Where is the institutional memory? If your AI is not retaining what it learned six months in, you are still on Rung 1, no matter what the demo looked like. Rung 3 stores knowledge in persistent, role-scoped memory across a team. That is the difference between a typing tool and a learning organization.

The cost premium is real. Multi-agent systems burn roughly 15× the tokens of a chat interaction. That number rules out the lowest-value tasks. But it unlocks work that cannot be done at Rung 1 at any token price. You cannot buy organization-scale cognition by paying more for a chatbot. You buy it by paying for anatomy.


The portable line

Models are organs. Topology is anatomy. You cannot replace anatomy with a bigger organ.

You can stack the most powerful model on the planet inside a Rung 1 harness, and the company will still feel, at the end of the fiscal year, that AI changed very little. It will have typed a great deal. It will not have learned anything.

Three rungs are visible now. The deltas between them are measured. And the natural experiment — Anthropic shipping single-agent for Code and multi-agent for Research — has settled the question of whether topology is preference or fit. It is fit.

The firms that compound an AI advantage over the next two years will not be the ones that choose the right model. That choice is converging. They will be the ones who choose the right topology for the right problem, and who build the memory, the coordination, and the governance to make Rung 3 structurally stable.

This is architectural work. It is quieter than a product launch. It shows up as a line on the income statement 18 months later, when the firm on Rung 3 describes its customers, its contracts, and its own internal operations in a way no one on Rung 1 can.

Ask the vendor which rung they are on.

Ask yourself the same question.


If your firm is working through which rung fits which task — or what it takes to architect the climb from Rung 1 to Rung 3 — reach out. I'm Dr. Jerry Smith at Verity Vantage Group, and that's what we do.


Building Minds is a newsletter about cognitive architecture for the age of agents. The premise across every edition is the same: durable AI advantage is not being built in the models — it is being built in the structures around them. The memory systems, the coordination protocols, the governance layers, the temporal intelligence, the topological anatomy. Each edition investigates one axis of that architecture with a single question in view: what does it actually take to turn a model into an organization?

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call