
Listen to the article on Apple Podcasts
Listen to the article on Soundcloud
Abstract
The dominant approach to building AI agents — a single foundation model wrapped in an orchestration framework, deployed to the cloud — replicates an architectural mistake that biological systems solved long ago. Brains are not monoliths. Nervous systems distribute intelligence across specialized components that communicate through standardized signals, fail independently, and scale without central bottlenecks.
This article examines what happens when we apply the same principle to artificial intelligence: federated agents with heterogeneous capabilities, coordinating through lightweight messaging protocols and routing tasks based on cost and privacy constraints rather than static configurations. The architecture runs in production today across healthcare, financial services, and private equity — sectors where data cannot leave organizational boundaries and single points of failure are unacceptable.
The question is not whether distributed AI systems are possible. The question is whether we will continue building monoliths while waiting for a bigger brain.
The Nervous System
Consider how you build a brain.
The obvious approach is monolithic: one organ, centrally located, handling everything. This works until scale becomes a problem. The brain that must process every sensory input, coordinate every motor response, and store every memory in a single location runs into fundamental limits. Bandwidth. Latency. The catastrophic consequence of local damage.
Evolution discovered a different architecture. The nervous system distributes intelligence. Peripheral nerves mediate reflexes without consulting the central nervous system. The spinal cord coordinates patterns too urgent for the round-trip to the skull. Specialized regions process vision, language, spatial reasoning — each optimized for its domain, communicating through standardized signals, failing independently.
We are building artificial minds. Most of us are building them wrong.
The dominant pattern for AI agents follows the monolithic model. Take a foundation model — GPT-4, Claude, Gemini — wrap it in an orchestration framework, connect tools, deploy to the cloud. One model handles reasoning, planning, execution, and memory. One endpoint receives every request. One provider’s uptime determines yours.
This architecture emerged from constraint. In 2023, cloud APIs were the only path to frontier capabilities. Local hardware couldn’t run competitive models. The tooling assumed centralization. These constraints made the monolithic pattern reasonable.
The constraints have changed. The architecture has not.

Apple Silicon enables 32-billion-parameter models to run at interactive speeds. Quantization enables 70B models to run on consumer hardware. The gap between cloud and local inference has narrowed from disqualifying to negotiable. Yet production systems remain frozen in the assumptions of two years ago.
The question is whether this matters.
Token economics scale poorly with agent complexity. A simple question-answering system costs fractions of a cent per interaction. An autonomous agent with tool use, multi-step reasoning, and memory retrieval achieves performance ranging from 50 cents to 2 dollars per task. Multiply by thousands of daily organizational interactions. The mathematics becomes hostile.
Privacy creates a harder problem. Every cloud API call transmits data to infrastructure you don’t control. For most applications, this represents managed risk. For healthcare, legal, defense, and financial applications, it represents disqualification. The sectors with the highest-value AI applications become the least able to deploy AI agents.
Single-model dependency creates single points of failure. On March 14, 2024, OpenAI experienced a multi-hour outage. Every system built on GPT-4 became simultaneously non-functional. This was the architecture working as designed.
No individual model excels at everything. Each has a capability profile — dimensions of strength and weakness. The monolithic pattern forces you to accept the weakest dimension of your chosen model as a ceiling for all tasks. You cannot route mathematical reasoning to a mathematics specialist while routing prose generation to a writing specialist. You get one profile, uniformly applied.
These observations suggest something. The pattern itself may be wrong.
Federation offers an alternative.
Instead of one system attempting everything, compose many systems that each do something well. Each agent maintains sovereignty over its domain. Communication happens through messages rather than hierarchical control. No agent requires any other to function. Together, they accomplish what none could alone.
The concept requires three supporting elements.
First: capability vectors. For agents to collaborate, they must advertise their capabilities. An agent publishes its model access, tool availability, resource capacity, domain expertise, and latency profile. Routing is a mathematical operation on the capability space. Need GPU access, codebase context, and sub-second response? Query the vectors, route accordingly.
Second: asynchronous messaging. Synchronous request-response creates temporal coupling. The caller blocks. If the callee is slow, the caller suffers directly. Asynchronous communication decouples. A request represents fire-and-forget from the sender’s perspective. Work continues. Responses arrive when ready, or timeout handling triggers. Agents operate at different speeds, go offline for maintenance, and process requests in whatever order suits their context.
Third: heterogeneous model topology. A federated system simultaneously leverages cloud frontier models for complex reasoning, local open-weight models for privacy, specialized fine-tuned models for domains, small, fast models for classification and routing. The routing layer selects based on task requirements, cost constraints, and availability. Requests that require sophisticated reasoning are delegated to capable cloud models. Requests processing sensitive data stay local. Requests during cloud outages fall back to alternatives.
The messaging substrate becomes invisible infrastructure. What emerges is collective intelligence that exceeds that of any single agent.

MQTT provides the communication layer. The protocol originated for IoT — unreliable networks, resource-constrained devices — but its properties suit agent federation precisely.
True publish-subscribe semantics: agents subscribe to topics, publish to topics, with no direct coupling between publishers and subscribers. Quality-of-service levels: fire-and-forget, at-least-once, or exactly-once delivery, based on operational criticality. Lightweight protocol overhead: minimal bytes, fast parsing, low memory footprint. Native support for unreliable connections: clients reconnect gracefully, optionally resuming sessions. Topic hierarchies: flexible addressing that maps naturally to organizational structures.
The protocol was designed for exactly the conditions federated agents face.
A topic hierarchy might look like this:
federation/agents/{agent-id}/inbox
federation/agents/{agent-id}/status
federation/broadcast/capabilities
federation/routing/requests
An agent who sends a direct message publishes it to the target’s inbox. A monitoring system subscribing to federation/# observes all traffic. An orchestrator subscribing to federation/agents/+/status receives heartbeats from all agents.
When an agent joins, it connects, subscribes to its inbox, publishes its capability vector, begins heartbeating, and subscribes to relevant broadcast channels. No centralized dependency exists. If the broker becomes temporarily unreachable, agents continue local operations and rejoin when connectivity is restored.
The architecture enables several organizational patterns.
In the orchestrator pattern, a designated agent accepts requests, decomposes them into subtasks, routes subtasks to specialists, and synthesizes results. A request to analyze a codebase and create documentation might decompose into static analysis routed to a code-understanding agent, documentation generation routed to a writing specialist, and diagram creation routed to a visualization agent. The orchestrator combines outputs into a coherent deliverable.
In the peer-to-peer pattern, agents communicate directly without central coordination. Each maintains awareness of others’ capabilities and routes requests based on local knowledge. Agent A receives a task requiring capabilities it lacks, queries the capability broadcast, identifies Agent B as qualified, and publishes the subtask directly to B’s inbox. No orchestrator involvement.
In the specialist pattern, agents assume fixed roles aligned with specific capabilities. A code-execution specialist runs sandboxed code. A web-search specialist queries external APIs. A memory specialist manages long-term storage. Each responds only to requests that match its function.
Production systems typically combine patterns. An orchestrator handles decomposition while specialists handle execution. Specialists communicate peer-to-peer for handoffs that don’t require orchestrator involvement. The messaging substrate supports all patterns simultaneously.

Real deployments demonstrate the principles.
Twenty hospitals across five continents trained a model to predict the oxygen requirements of COVID-19 patients. Each hospital is trained locally on its patient data. Only model parameters — not patient records — traversed network boundaries. The federated model showed 38% improvement in generalizability and 16% improvement in performance compared to single-institution training. The result appeared in Nature Medicine.
IBM developed a federated system for anti-money-laundering detection across banks and payment networks. The architecture combined homomorphic encryption, differential privacy, and random forests. It placed second in the NIST Privacy Enhancing Technologies Challenge — a competition specifically evaluating whether privacy-preserving techniques can work in practice.
Banking Circle uses the Flower framework to train models on European data without cross-border transfers, then deploys US-market models while maintaining GDPR compliance. The data never moves. Only gradients cross borders.
These are not theoretical possibilities. They are running systems with measured outcomes.
The private equity sector presents an interesting test case.
A single analyst reviewing a confidential investment memorandum needs privacy but not federation. Local inference suffices — run a model on-device, extract financial metrics, and generate analysis. The data never leaves. This is privacy through isolation.
A single firm with multiple departments — legal, financial, operational — conducting due diligence might centralize in a secure VPC. Same legal entity, unified data governance, simpler than federation.
But a PE firm benchmarking supply chain risk across ten portfolio companies faces a different problem. Those companies compete with each other. They cannot share raw vendor lists. They can share model gradients. Federated learning trains a “Supply Chain Risk” model across all ten — each trains locally on their ERP data, shares only model updates, and the aggregated model improves prediction for all participants.
Club deals involving multiple PE firms present the same dynamic at a different scale. Competitive dynamics prevent data pooling. Gradient-only sharing enables collaboration without exposure.
Cross-border fund operations add regulatory constraints. GDPR requires data locality. Federated architectures with regional nodes satisfy the requirement while enabling global model improvement.
The pattern: federation adds value when data cannot be pooled due to competition, regulation, or trust boundaries. When those constraints don’t apply, simpler approaches work better.
Implementation involves trade-offs that warrant explicit statement.
Differential privacy protects against inference attacks by adding calibrated noise to gradients. The privacy parameter epsilon controls the trade-off: a strict epsilon provides stronger privacy but degrades model accuracy. Practitioners must tune explicitly.
Homomorphic encryption enables computation on encrypted data. The computational overhead ranges from 10x to 1000x, depending on the operation's complexity. Not suitable for latency-sensitive applications without specialized hardware.
Secure aggregation prevents the central server from observing individual gradients. It increases the number of communication rounds and requires minimal participant thresholds.
Data heterogeneity challenges generalization. Financial institutions have different schemas, formats, and distributions. A model trained on Bank A’s patterns may not generalize to Bank B’s. Techniques like FedProx help. Expect a 10–20% degradation in accuracy relative to IID assumptions.
Federation requires orchestration infrastructure: message brokers for coordination, aggregation servers for parameter collection, monitoring of participant health, and per-participant authentication. The operational overhead is real.
These costs must be weighed against benefits. The accounting is not always favorable. With fewer than five participants, simpler approaches — secure enclaves, differential privacy on centralized data — often provide better trade-offs. Federation shines at scale. Ten institutions. Twenty. The benefits compound while per-participant costs amortize.

Multi-agent AI extends beyond federated learning into orchestrated workflows.
Walmart deploys specialized agents for demand forecasting, inventory optimization, pricing, and routing. Each agent handles one domain. They coordinate to achieve SKU-level autonomous forecasting, integrating point-of-sale data, weather, and trends.
SAP Joule functions as an enterprise-wide multi-agent system spanning HR, CRM, and Finance. Each domain has agents optimized for its workflows. The agents collaborate across domain boundaries.
A global consumer goods company modernizing its ERP system deployed testing, validation, and compliance agents, which reduced manual testing by 60% during a multi-region SAP S/4HANA rollout. The agents monitored changes, identified test cases, prioritized risk-based testing, and executed validations across functions.
MIT Sloan reports that 35% of organizations currently use agentic AI, with 44% planning to expand their use. Among high-adoption firms, 52% prioritize agent-to-agent interactions — not just human-to-agent, but agents coordinating with each other.
The pattern appears across sectors: insurance claims processing with seven specialized agents collaborating in a human-in-the-loop audit; due diligence automation reducing research time from seven days to one; and deal pipeline management with sourcing, screening, diligence, and valuation agents feeding an orchestrator.
We return to the nervous system analogy.
The brain did not evolve toward greater centralization. It evolved toward greater distribution with better coordination. Reflexes that don’t require conscious attention. Pattern matching that happens before awareness. Specialized regions optimized for their domains, connected by standardized signaling.
The future of artificial intelligence follows the same trajectory. Not one god-model answers every question. A society of specialists — reasoning engines, code generators, data analysts, security monitors — each doing what they do best, coordinating through protocols designed for exactly this purpose.
The technology exists. MQTT brokers are battle-tested. Local inference is practical on consumer hardware. Federated learning has moved from research to production deployments with measured outcomes.
The question is whether to build the nervous system or to wait for a larger brain.
Dr. Jerry Smith is a computational neuroscientist and AI architect who builds autonomous systems that work in production.