Building Minds
The Productivity Gap Will Not Be Solved One Task at a Time
Dr. Jerry A. Smith · July 27, 2026 · 18 min read

Leadership’s next job is to turn AI from a collection of assistants into a human-agent operating system
Building Minds
By Dr. Jerry A. Smith | Building Minds · Verity Vantage Group
A CEO walks into an operating review with a reasonable question.
The company has thousands of employees using generative AI. Analysts draft faster. Developers generate code. Sales teams prepare account briefs in minutes. Legal teams summarize contracts. Customer-service representatives receive suggested responses before they finish reading the ticket.
The internal surveys look excellent. Employees report saving time. Usage is climbing. Every function has a success story.
So the CEO asks: Where did the value go?
Did margin improve? Did the order-to-cash cycle shorten? Did the company serve more customers without adding cost? Did risk decline? Did a product reach market sooner? Did working capital move? Did the business avoid hiring it had already planned?
The room has plenty of productivity anecdotes. It has no common answer.
That silence is the AI productivity gap.
It is tempting to blame the model, the data, employee resistance, or a shortage of use cases. Sometimes those are real problems. But the deeper issue is architectural and managerial. Most organizations have deployed AI at the level of the task while continuing to operate the enterprise at the level of the old workflow.
They have made pieces of work faster without redesigning how the pieces become an outcome.
That is why the next phase of enterprise AI cannot be another wave of copilots. It must be the deliberate construction of agent-enabled human-agent ecosystems: persistent operating systems in which humans and specialized agents share goals, memory, state, controls, feedback, and responsibility for an end-to-end business program.
The shift is larger than moving from one agent to many agents. It is a shift from asking AI to complete tasks to organizing humans and agents around programs—customer retention, software delivery, claims resolution, revenue-cycle performance, supply continuity, regulatory readiness, product launch, portfolio operations.
A task ends when an output is produced.
A program persists until an outcome is achieved, measured, and sustained.
That difference is where enterprise value begins.
The evidence is not contradictory
The research on AI productivity is increasingly clear on one point: generative AI can produce substantial gains on well-bounded work.
In a large customer-support deployment studied by Erik Brynjolfsson, Danielle Li, and Lindsey Raymond, AI assistance increased issues resolved per hour by 14 percent overall and by 34 percent among novice and lower-skilled workers. In a randomized experiment by Shakked Noy and Whitney Zhang, professionals completed writing tasks 40 percent faster, while independent quality ratings improved by 18 percent. In the BCG field experiment led by Fabrizio Dell’Acqua and colleagues, consultants working inside the model’s capability frontier completed more tasks and worked faster.
Those are real effects. They should not be dismissed simply because enterprise income statements have not moved at the same speed.
But the counterevidence matters just as much.
Outside the model’s capability frontier, the BCG participants were 19 percentage points less likely to produce a correct answer. METR found that experienced open-source developers using early-2025 AI tools took 19 percent longer on real issues even though they believed AI had made them faster. A six-month field experiment across 66 firms and 7,137 knowledge workers found that active AI users spent roughly two fewer hours a week on email, yet researchers detected no corresponding change in the quantity or composition of work. Danish data covering approximately 25,000 workers and 7,000 workplaces found reported time savings but no measurable effect on earnings or recorded hours after two years.
The right conclusion is not that AI works or that AI does not work.
The right conclusion is that task productivity is a technical effect; enterprise ROI is an organizational accomplishment.
A model can improve a task without changing a workflow. A workflow can improve without releasing usable capacity. Capacity can be released without being captured. Operating value can be captured without being credibly attributed. At each boundary, leadership has to design the conversion.
The evidence becomes coherent once we separate four levels that enterprises routinely collapse:
- Model capability — Can the model reason, generate, retrieve, classify, plan, or act?
- Task performance — Does it make a bounded activity faster, cheaper, or better?
- Workflow performance — Does the end-to-end system improve after reviews, queues, handoffs, exceptions, and approvals?
- Enterprise outcome — Does the improvement become revenue, margin, cash, throughput, quality, customer value, risk reduction, or strategic capability?
Most AI dashboards stop at level two.
The board is asking about level four.
Between them sits the organization.
Why the task model runs out of road
The first generation of enterprise generative AI was built around a simple interaction: a person has a task, sends a prompt, receives an answer, and decides what to do next.
That pattern is useful. It is also structurally limited.
The employee remains the integration layer. The employee notices the need, gathers the context, chooses the tool, writes the prompt, checks the answer, moves it into the system of record, requests approval, follows up on the exception, and remembers what happened the next time.
AI accelerates the visible task. The human still carries the workflow.
This is how organizations end up with hundreds of local wins and almost no aggregate change. The assistant drafts the contract faster, but legal review remains queued for four days. The sales agent writes better outreach, but account selection and pricing approval still constrain conversion. The coding agent generates more software, but architecture, security review, testing, and release capacity become the new bottlenecks. The analyst produces the report sooner, but nobody changes the decision cadence the report was meant to support.
The gain does not vanish. It is absorbed.
It is absorbed by review, rework, waiting, handoffs, ambiguity, exception handling, and unchanged management routines. Sometimes it becomes a better employee experience. Sometimes it becomes slack. Sometimes it simply moves the bottleneck downstream.
The task model also resets too often. A chatbot knows what is in the current conversation. A business program must know what was promised last quarter, which exception was approved, how the customer reacted, why a policy changed, which hypothesis failed, what the finance team accepted, and what remains unresolved.
Without persistent state and institutional memory, every interaction begins with a talented stranger.
That is not an operating system. It is a series of encounters.
The unit of transformation is the program
A task has an instruction and an output.
A workflow connects tasks into a repeatable process.
A program coordinates multiple workflows, decisions, resources, people, and time horizons against a durable outcome.
This distinction matters because most valuable enterprise outcomes are program-shaped.
Customer retention is not a task. It includes signal monitoring, account history, relationship judgment, product usage, service recovery, pricing, contract timing, executive intervention, and follow-through.
Software delivery is not a task. It includes product intent, architecture, implementation, testing, security, release, observation, incident learning, and technical-debt management.
Revenue-cycle performance is not a task. It includes intake, coding, documentation, submission, denial prediction, appeals, payer rules, cash application, and process correction.
These programs do not need a single super-agent pretending to be the entire company. They need a coordinated ecology of specialized intelligence.
An orchestrated retention program might include:
- a signal agent watching usage, support, sentiment, payment, and renewal indicators;
- a research agent assembling account history and relevant commitments;
- a risk agent estimating churn and identifying contradictory evidence;
- a commercial agent testing renewal and expansion scenarios;
- a policy agent checking concessions and approval boundaries;
- an orchestration agent managing dependencies, deadlines, and handoffs;
- a shared memory layer preserving decisions, outcomes, corrections, and unresolved commitments;
- a human relationship owner deciding how to act when trust, negotiation, reputation, or strategic judgment dominates.
The agents do not merely hand a summary to the human and disappear. They remain connected to the program. They communicate changes in state. They request work from one another. They challenge evidence. They escalate exceptions. They update shared memory. They observe the effect of the intervention. They learn which signals mattered.
The program continues after the email is sent.
That is the difference between automating a task and operating a capability.
What a human-agent ecosystem actually is
The phrase human-agent ecosystem can easily become another inflated label. Adding five bots to a diagram does not create an ecosystem. Neither does allowing agents to talk to one another without boundaries.
A production human-agent ecosystem has at least seven parts.
1. A durable outcome
The system is organized around a business result, not around the use of AI. The outcome might be lower cost per resolved claim, faster product release with no increase in escaped defects, higher renewal retention, reduced inventory exposure, or shorter time from order to cash.
The objective must be measurable, economically material, and owned by someone with authority to change the work.
2. A human authority structure
Humans remain accountable for purpose, tradeoffs, risk, trust, and the legitimacy of action. Their role shifts from manually carrying out every step to designing and governing the conditions under which work occurs.
A serious program names at least:
- a business outcome owner;
- a workflow or program owner;
- a technical owner;
- a risk and control owner;
- a finance owner.
This is not committee-building. These owners resolve different classes of ambiguity. The business owner decides what matters. The workflow owner changes standard work. The technical owner assures system performance. The risk owner defines non-negotiable boundaries. Finance decides when value is real enough to book.
3. Specialized agents with bounded roles
Agents should have explicit jobs, context, tools, permissions, and success conditions. A specialist tasked with monitoring exceptions should not quietly acquire authority to approve them. A research agent should not become the final judge of policy. A commercial agent should not rewrite risk thresholds because doing so improves conversion.
Role boundaries make agent behavior inspectable. They also make failure containable.
4. An orchestration and communication layer
Intercommunication is necessary, but all-to-all conversation is not the goal. Unstructured agent chatter can amplify errors, duplicate work, and obscure accountability.
Agents need communication contracts: what message types exist, what evidence must accompany a claim, which agent can request or assign work, how conflicts are reconciled, when state changes, and what conditions trigger human escalation.
The orchestrator maintains the work graph. It knows dependencies, deadlines, unresolved questions, current state, and which actor—human or agent—has the next decision. It selects the topology that fits the work: a single agent for tightly coupled sequential reasoning, parallel specialists for decomposable investigation, or a staged team for work that needs independent checks and synthesis.
The orchestrator is not just a productivity feature. It is a control plane.
5. Shared and role-scoped memory
A program cannot compound if its intelligence disappears at the end of each session.
The memory architecture must preserve several kinds of state:
- semantic memory: facts, policies, products, customers, and operating knowledge;
- episodic memory: what happened, when, and with what outcome;
- procedural memory: approved methods, controls, and standard work;
- prospective memory: commitments, deadlines, dependencies, and future actions;
- decision memory: what was decided, by whom, on what evidence, and under which assumptions.
Not every agent should see every memory. Access follows role, need, sensitivity, and purpose. The goal is not a giant context window. It is the right institutional context reaching the right actor at the right moment.
6. Deterministic and human judgment gates
Agents are probabilistic workers. Enterprise systems still need deterministic controls.
Validation rules, reconciliations, permissions, schema checks, test suites, policy checks, thresholds, and audit logs should surround agent behavior. Human judgment should be concentrated where ambiguity, consequence, trust, ethics, or strategic tradeoffs are highest.
The mature design is not “human in every loop.” That simply recreates the bottleneck. Nor is it “human out of the loop.” That confuses autonomy with accountability.
The goal is human judgment at the right loops.
7. Operational and economic telemetry
The ecosystem must observe more than agent accuracy and token cost.
It should measure:
- task quality and critical-error rate;
- end-to-end cycle time and throughput;
- queue, review, rework, and exception burden;
- adoption and workflow penetration;
- overrides and escalations;
- model, data, integration, support, and governance cost;
- customer, quality, risk, revenue, margin, and cash outcomes;
- realized value against the baseline and counterfactual.
Without this layer, an organization can operate a technically impressive agent system while remaining economically blind.
The architecture looks less like a chatbot and more like an operating system:
Recommended by LinkedIn
[
Claude Cowork: Your AI Colleague Is Here
Sanil N.
8 months ago](https://www.linkedin.com/pulse/claude-cowork-your-ai-colleague-here-sanil-nadkarnee-8xxhf)
[
Why AI Time Savings Rarely Become Business Value
Kieran Gilmurray
2 months ago](https://www.linkedin.com/pulse/why-ai-time-savings-rarely-become-business-value-kieran-gilmurray-pwyye)
[
The real cost of speed over clarity
Yota Trom
2 months ago](https://www.linkedin.com/pulse/real-cost-speed-over-clarity-yota-trom-w6tie)
Business outcome + economic baseline
│
Human program authority
intent · tradeoffs · accountability
│
Orchestration / control plane
work graph · routing · state · escalation
┌───────┼────────┐
Specialist Specialist Specialist
agent agent agent
└───────┼────────┘
│
Shared memory + authoritative systems
│
Tests · policy · permissions · audit
│
Workflow results + financial telemetry
└──── feedback to program
The feedback path matters as much as the execution path. Without it, the ecosystem cannot tell whether its actions improved the program or simply generated more activity.
Interconnected does not mean uncontrolled
The attraction of multi-agent systems is easy to understand. Specialized agents can search in parallel, challenge one another, preserve role-specific context, and combine distinct forms of analysis. For parallelizable, coverage-heavy work, that topology can materially outperform a single generalist.
But more agents do not automatically mean more intelligence.
Tightly coupled, sequential work can degrade when split across too many actors. Handoffs lose context. Coordination consumes time and tokens. One agent’s unsupported assumption becomes another agent’s input. Errors can compound while every participant appears locally reasonable.
Leadership should therefore resist a new form of technological theater: the animated diagram filled with agents sending messages to one another.
The question is not, How many agents can we connect?
The question is, Which topology best preserves truth, contains error, and advances the program outcome?
Sometimes the answer is one capable agent with strong tools and memory. Sometimes it is an orchestrator with parallel specialists. Sometimes it is a staged architecture in which independent agents produce evidence, analyst agents reconcile it, and a synthesizer creates a unified recommendation. Sometimes the correct answer is to keep the decision human.
The system should earn complexity. It should not begin there.
Leadership has to redesign the enterprise around the new division of labor
The productivity gap persists because organizations often treat AI as a technology rollout. The business names a use case, IT integrates a tool, users receive training, and a dashboard reports adoption.
That process can deploy software. It cannot by itself redesign the production system.
Closing the gap requires a set of leadership decisions that cannot be delegated to the AI team.
1. Choose a value program, not a bag of use cases
Start with a material outcome and the system that produces it. Map the customer promise, the value stream, the binding constraint, and the decisions that govern flow.
Do not begin with “Where can we use an agent?” Begin with “Which enterprise outcome is constrained, and what would have to change for that constraint to move?”
A faster activity outside the constraint may be useful. It is not the foundation of an ROI case.
2. Establish the baseline before the build
If the baseline arrives after the demo, it is not a baseline. It is a story.
Measure current volume, labor, cycle time, queue time, quality, rework, risk, service level, and full cost. Decide how the counterfactual will be estimated: holdout, staggered rollout, matched comparison, trend adjustment, or an explicit attribution haircut.
Finance should help design this before deployment, not audit it after the team has declared victory.
3. Map decisions, not just process steps
Traditional process maps show activities and handoffs. Human-agent systems also require a map of authority.
Which decisions can an agent make? Which can it recommend? Which require human approval? Which conditions suspend autonomy? Who owns an exception? What evidence must accompany a recommendation? What happens when agents disagree? Who can change a threshold?
These choices are the real operating model.
4. Design the human roles at the same time as the agent roles
If leaders define agent jobs but leave human jobs unchanged, people become hidden middleware. They spend their days reviewing everything, repairing context, moving outputs between systems, and accepting accountability for decisions they did not meaningfully shape.
Human roles should move toward:
- setting intent and priorities;
- making high-consequence tradeoffs;
- designing constraints and escalation policy;
- handling ambiguous and relationship-sensitive exceptions;
- validating evidence and accepting risk;
- teaching the system through correction;
- redesigning the program as conditions change.
This is not simply “upskilling.” It is a new division of labor.
5. Build communication and memory as shared infrastructure
An enterprise does not need hundreds of disconnected agent memories and improvised handoffs. It needs a governed layer for identity, state, events, commitments, provenance, access, and communication.
Protocols such as MCP and A2A can help connect tools and agents, but protocol compatibility is only the transport layer. The harder questions are semantic and organizational: what does a status mean, which source is authoritative, how is a claim supported, who may act on it, and what must be retained?
Interoperability without shared meaning creates faster confusion.
6. Turn governance into executable architecture
Governance should not be a document that agents are expected to remember. It should appear in permissions, tool boundaries, schemas, tests, thresholds, evidence requirements, logs, rate limits, approval gates, and kill switches.
Policy becomes more reliable when it is expressed as a property of the system.
This does not eliminate governance committees or human review. It gives them something observable to govern.
7. Decide what happens to released capacity
Time saved is not money saved until the operating plan changes.
Leadership must choose an economic destination before declaring ROI:
- remove overtime, contractor spend, rework, licenses, or future hiring;
- absorb additional demand with the same cost base;
- redeploy capacity to a measured backlog or revenue-producing activity;
- improve conversion, retention, price, or service commitment;
- reduce loss, risk, inventory, downtime, or delayed cash;
- create a new product, service, or cognitive asset.
If leadership chooses none of these, the honest result is capacity created, not financial value realized.
That capacity may still be worthwhile. Better employee experience and strategic option value matter. But they should not be mislabeled as booked savings.
8. Change incentives and management cadence
An agent-enabled program cannot survive if employees are rewarded for preserving the old process, managers are measured on local utilization, and functions optimize their own queues.
Metrics must move from individual activity to program outcomes. Operating reviews should examine workflow flow, agent behavior, exception patterns, human review burden, and economic capture together. Teams need permission to stop weak automations, alter decision rights, and redesign work when the bottleneck moves.
The program should have a learning cadence:
- weekly operational review of failures, queues, and exceptions;
- monthly review of adoption, quality, cost, and operating outcomes;
- quarterly finance reconciliation and redesign decision;
- explicit scale, revise, contain, or stop gates.
9. Treat production as the midpoint
A system entering production proves that it can run in the real environment.
It does not prove that people use it correctly, that the workflow changed, that capacity was captured, that economics remain favorable, or that the result persists.
Before production, leadership designs the value hypothesis, baseline, workflow, topology, controls, and ownership.
After production, leadership must drive adoption, reliability, behavior change, economic capture, finance reconciliation, and reuse.
Production is not the finish line. It is where value realization becomes testable.
A 90-day path from assistant to program
Leaders do not need to redesign the whole enterprise at once. They need one bounded program where the economics are visible and the authority to change work is real.
Days 1–30: Define the program
Choose one material outcome. Map the current workflow, decisions, queues, exceptions, and binding constraint. Establish the baseline. Name the business, program, technical, risk, and finance owners. Write the benefits-capture contract: what will improve, how it will be measured, what released capacity will become, and what would cause the initiative to stop.
Days 31–60: Design the ecosystem
Define the human and agent roles. Select the simplest topology that fits the work. Specify communication contracts, memory boundaries, authoritative systems, permissions, deterministic controls, human judgment gates, and escalation paths. Build an evaluation set from representative work, including difficult exceptions and outside-the-frontier cases.
Days 61–90: Operate and prove
Run the new system on real work with controlled exposure. Measure end-to-end performance, not only agent output. Track review, rework, abandonment, escalation, cost, quality, and risk. Change standard work. Execute the capacity decision. Reconcile results with finance. Then scale, redesign, contain, or stop.
The objective of the first 90 days is not maximum autonomy.
It is credible evidence that a human-agent program can improve a business outcome under real operating conditions.
The future enterprise is not human or agent
The usual debate asks which jobs AI will replace and which jobs humans will keep. That framing is too static.
The more important question is what kind of organization becomes possible when human judgment and machine agency are designed as one production system.
Humans bring purpose, accountability, moral and strategic judgment, social trust, contextual interpretation, and the authority to make tradeoffs on behalf of the institution. Agents bring persistent monitoring, rapid retrieval, parallel search, simulation, routine execution, cross-system coordination, and tireless attention to state.
Neither side is enough.
Humans alone become the coordination bottleneck as complexity grows. Agents alone lack legitimate authority, grounded purpose, and reliable judgment under novel consequences. The advantage comes from architecture: putting each form of intelligence where it is strongest, then building interfaces that allow the whole to learn.
This will change the shape of management.
Managers will spend less time distributing and chasing tasks and more time defining outcomes, designing decision systems, resolving exceptions, allocating trust, and improving the ecology of work. Domain experts will become teachers and evaluators of institutional intelligence. Finance will move upstream into system design. Risk leaders will express controls as operating architecture. Technology leaders will manage not only applications and data, but populations of agents with identity, permissions, memory, cost, provenance, and performance histories.
The organization itself becomes programmable—not because code replaces leadership, but because leadership decisions are increasingly expressed in the topology, memory, permissions, metrics, and escalation logic of the human-agent system.
That is a much larger idea than task automation.
The leadership mandate
The productivity gap will not close because the next model is faster. It will not close because prompt libraries improve. It will not close because every employee receives an assistant.
Those developments may raise local productivity. They do not perform the organizational conversion.
Leadership has to do that.
Leaders must choose the outcomes worth reorganizing around. They must define how humans and agents divide authority. They must connect specialized agents through governed communication rather than uncontrolled conversation. They must build institutional memory so the system learns across time. They must redesign workflows, incentives, reviews, and exception paths. They must decide what released capacity becomes. They must require finance-verifiable proof.
The next enterprise AI frontier is therefore not the isolated agent. It is the orchestrated, interconnected, intercommunicating human-agent ecosystem working against a persistent program.
Not a bot that completes a task.
A system that senses, reasons, coordinates, acts, measures, remembers, and improves—while humans retain responsibility for purpose, consequence, and trust.
A company does not become agentic when it buys more AI seats.
It becomes agentic when humans and agents can pursue a shared outcome through a governed operating system—and when that system changes what the enterprise can reliably achieve.
Copilots can make employees faster.
Human-agent ecosystems can make the company different.
Dr. Jerry A. Smith