All posts

Convention, Not Telepathy

Dr. Jerry A. Smith · September 24, 2026 · 9 min read

What shared context actually buys a team of agents — and the number that tells you when something is leaking between them

I wrote two agents an explicit convention. It was short, it was correct, and both agents could read it. Then I watched them ignore it for two hundred consecutive rounds and score 26.5%, worse than two agents flipping coins.

The convention wasn’t the problem. The agents carried a convention of their own, and they shared it more deeply than anything I had documented.

That experiment came near the end of a research program I ran on my own laptop, and it left me with a single number every team shipping multi-agent systems should understand. The number is 75. On one common decision shape (converge by default, split when both agents hit the unusual case), it is the ceiling on what shared context can buy for agents that don't communicate. On that test, anything above it indicates a leak.

The category error: shared context is not coordination

Most multi-agent architectures rest on a quiet assumption. If the agents share enough context (the same model, system prompt, memory store, and persona), they will coordinate without the latency and expense of communicating. Vendors describe it as collaboration. Earlier in this program, I described it as entanglement.

It is a reasonable assumption, and it is half right.

Shared context does synchronize behavior. What it cannot do is let one agent act on information its teammate privately holds. Those are different capabilities, and most organizations measure neither. They append another paragraph to the shared prompt, notice fewer collisions during the demonstration, and declare coordination solved.

The distinction matters most for a specific decision pattern, and production systems are full of it. Two agents act simultaneously, each on information the other doesn’t have, and the correct joint outcome depends on both. Two workers on a shared queue. A planner and an executor deciding whether to escalate. Two routing agents splitting cases when both queues run hot. There is usually a sensible default, and it fails in exactly one situation: when both agents encounter the unusual case at the same time.

That failure is where shared context runs out.


The ceiling is 75, and it is a theorem

To measure this cleanly, I reduced the decision to its minimal form. Each agent privately observes whether its own ticket is hot. Each chooses a lane. The team scores when both choose the same lane, unless both tickets are hot, in which case they must separate.

Converge by default. Split under joint pressure.

Physicists will recognize the structure. It is the CHSH game, the architecture behind half a century of Bell experiments, and it comes with a mathematical proof. If two agents share anything fixed in advance — weights, instructions, notes, a random seed — and do not communicate once the tickets are dealt, the team wins at most 75% of the time.

The proof fits in a paragraph. Freeze everything the agents share, and each agent’s behavior reduces to a deterministic rule: if my ticket is hot, choose this lane; otherwise, choose that one. Winning every round would require four conditions at once — same lane when neither is hot, same lane when only one is hot (twice), different lanes when both are hot. The first three force all four choices to match, which breaks the fourth. Any fixed pair of rules loses at least one of the four cases. Shared randomness only mixes fixed rules, and a mix cannot beat its best ingredient.

The ceiling therefore applies equally to a one-bit coin and to a billion shared parameters. That is the uncomfortable conclusion for anyone budgeting for richer shared context.

The optimal strategy under the ceiling is also the dullest available: always choose lane 0. It wins every round except the one where both tickets are hot. That is not intelligence. It is a convention, in the sense Schelling and Lewis meant it — a shared default that lets separated parties agree without talking.

Shared context buys your agents a convention, not telepathy.

The receipts: every rung landed where the math said

I built a harness to verify the ladder with scripted agents, one thousand rounds per configuration, with each agent isolated from its teammate’s inputs. The measured results:

  • Nothing shared: 49.5%. Coin flips, as expected.
  • Best shared convention (always lane 0): 74.6%. On the ceiling.
  • A shared “quantum coin” — a simulated entangled pair measured the same way each round: 74.5%. Indistinguishable from the convention, cell for cell.
  • Quantum measurement that depends on each agent’s own ticket (simulator only): 84.3%. Above the ceiling, near its own theoretical limit of about 85%.
  • One message per round: 100%.

The breakdown by scenario is more instructive than the totals. The shared convention won every ordinary round and lost every round where both tickets were hot. That is the ceiling, visible in a single row of a table.

Two notes on honesty. The quantum results come from a laptop simulator; they show what the mathematics allows, not what any hardware did, and no production agent stack has access to that resource anyway. And the “quantum coin” result corrected my own earlier work: three experiments I had been treating as quantum were, mathematically, a shared coin in costume.

The mistake I made for three experiments

Before building the team game, I scored each agent individually across three experiments. Whatever I placed in the shared background (nothing, a classical note, quantum-correlated bits), every agent landed near 70%. I documented it as an honest null: shared priors don’t help.

The null was built into the measurement.

An individual agent’s score depends only on whether its own information points in the correct direction. Correlation with a teammate changes how the answers align, not how frequently either agent is correct. When I went back and re-scored the same logs by team outcome, the effect was there all along — agreement rose from 57% with private backgrounds to 64–68% with shared, quantum-correlated ones.

I had been measuring the agents. The effect lived in the team.

If your dashboards score each agent separately — accuracy per agent, success rate per tool call — they cannot see what your shared context is doing. Not a little. At all.

The model has a convention of its own

Then I placed real models in the seats: two Claude agents, claude-haiku-4-5, each invoked once per round with the rules and its own ticket, each running in a separate operating-system process, with tickets dealt from a cryptographic random source neither could predict. In the first pilot, I asked each agent for a one-character answer, 0 or 1, and I gave both agents the same written note: always pick lane 0; it wins three rounds in four.

Across two hundred rounds, both agents did something different, every single time. Hot ticket, lane 1. Otherwise, lane 0.

The team won only when neither ticket was hot: 26.5%. Without the note, the same habit showed up about 92% of the time, for 37.5%. Even when I deliberately planted a leak — copying one agent’s ticket into the other’s “shared memory” — the team reached only 51%, because the agent reading the leak used it the way its habit suggested rather than the way the rules required.

This is the observation I keep returning to. The agents did share a convention. It lived in their weights rather than in my note (hot means one), and both followed it automatically, because both were the same model. It overrode the documented convention, and it performed worse.

Your shared system prompt is not the only thing your agents share. The model is shared too, and it has defaults you never wrote and cannot see until you test for them.

The change I made was operational rather than philosophical: let the agents reason briefly before answering, and relabel the lanes A and B so nothing in the labels invites the habit. Then I ran the pilot again: same model, same tickets, two hundred rounds per setup.

The habit disappeared. With the written convention, the team scored 74.5% — on the ceiling. Without any convention at all, it also scored 74.5%; reasoning from the rules alone, the agents worked out “always lane A” by themselves and lost 47 of the 48 rounds where both tickets were hot, exactly the losses the ceiling predicts. With the planted leak, the team jumped to 84.5%, above anything shared context can buy, and the test flagged it. With an honest, counted message, 88%. Not one answer was unusable.

Same weights, same rules, different outcome. The only change was whether the agents could think before they answered.

These are pilot numbers, two hundred rounds each. The thousand-round binding runs come next, and I will publish those when they exist, not before.

Above the ceiling is a channel

Here is why the number is operationally useful rather than merely interesting.

The 75% ceiling covers everything fixed before the round. So if silent agents beat it by a clear statistical margin, only two explanations remain. Either they hold a quantum resource measured in the special way above — no agent stack does — or information about one agent’s ticket is reaching the other after the tickets are dealt.

That converts the ceiling into a diagnostic you can run against your own architecture. The condition matters: the ceiling holds when each agent’s signal is dealt fresh and at random, hot half the time, not on your lopsided production traffic. Deal fresh random signals, route the agents through the production pipeline (the same orchestrator, memory store, tools, retries, and logging), and compare the team score to 75%.

Where do the channels usually live? The usual suspects are shared memory written mid-task, orchestrator summaries that pass one agent’s state into another’s prompt, scratchpads and caches both agents can read, retry loops that include the other agent’s output, inputs predictable from a counter or timestamp, and agents meant to run in parallel that actually run in sequence.

A channel is not automatically a defect. Sometimes it is exactly the coordination you want. But then it is communication, and it should be counted, secured, and priced as communication — not credited to “shared context.”

Two conditions have to hold before you trust the test. First, a competence check: with the convention written into their shared context, your agents should score close to 75%. My first pilot scored 26.5%, which means the test could not have detected anything. My second scored 74.5%, and it caught the planted leak. Second, a planted leak: copy one agent’s signal into something the other reads and confirm the test fires. A detector that misses a planted leak has not passed anything.

The honest ceiling on my own claim

This is a single task, the simplest one with an interesting limit. Production decisions carry their own ceilings, and the test as constructed only detects channels that improve performance on this game.

At a thousand rounds, it reliably detects a leak that lifts a team to approximately 79% or higher. Subtler leaks require more rounds.

My clean runs are a calibration, not an audit. I designed the harness so a clean run cannot leak by construction; the valuable next step is pointing the diagnostic at a production stack with real memory and orchestration, where it could discover something.

And the quantum rungs are simulation. I report no hardware advantage, and I don’t think one belongs in your architecture review.

Portable line

Shared context buys a convention, not telepathy. The convention tops out at 75%. Anything above it is a channel you haven’t counted.

Diagnostics for your next architecture review:

  • Where do two of your agents decide at the same time, on information the other lacks, with a default that fails when both hit the unusual case.
  • Have you written that default down explicitly — and checked that the agents actually follow it, rather than a habit of the model.
  • Do you score those decisions at the team level, broken down by which agents saw the unusual case.
  • If your agents coordinate better than a convention allows, can you name the channel.
  • For the joint case your convention loses, is the fix one short message rather than one more paragraph of shared context.

The full write-up, with the proof, the harness design, and the agent prompts that run this audit for you, is in progress. If you run the test on your own stack, I would like to hear what it found — especially if it found something.

—

Dr. Jerry A. Smith Verity Vantage Group · Building Minds

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call