All posts

We Built a 10-Agent AI System That Monitors Our $225K Project in Real-Time — Here’s What We Learned

Dr. Jerry A. Smith · November 10, 2025 · 12 min read

Listen to the article on Apple Podcasts
Listen to the article on Soundcloud

It was November 8th, and our AI system had just flagged a critical risk: Action Item AI010 was due in 7 days, but the deadline day was already packed with 6 hours of meetings. Only 2 hours remained for actual work — not enough to complete a deliverable that was only 75% done. And worse? Three other critical tasks depended on it. Miss this deadline, and the entire project timeline would cascade into failure.

This wasn’t a human discovery. Four AI agents — ActionItemSight, EventSight, DeliverableSight, and RiskSight — had independently detected different pieces of the puzzle, then synthesized them into a compound risk alert. A week before the crisis would have become visible through manual monitoring.

That’s ForeSight: a multi-agent AI system we built to monitor our $225,000 consulting engagement in real-time. Ten specialized agents continuously analyze action items, deliverables, budget, emails, calendar events, and stakeholder relationships — then generate executive dashboards in 4.2 minutes flat.

Here’s how we designed, built, and deployed a system that transformed project management from reactive firefighting to proactive intelligence.

The Problem: Drowning in Data, Starving for Insight

Like most research managers, I was spending 35–45 hours per week compiling status reports. Every day began the same way: open Gmail (50+ unread emails), scan Slack (100+ messages overnight), check JIRA (23 action items to track), review Google Calendar (20–30 meetings this week), reconcile budget spreadsheets. Pull data from five systems, manually synthesize into a coherent status update, and send to stakeholders. Repeat daily.

The math was brutal: 35 hours/week × $150/hour × 48 weeks = $252,000 per year in pure manual labor. And that’s just one project manager. Scale that across a consulting firm with 10 PMs managing 25 projects, and you’re looking at $2.5 million annually spent on status compilation.

But the real cost wasn’t time — it was late detection. When you’re manually checking systems once or twice daily, critical risks surface 2–3 weeks after they’ve already started compounding. A stakeholder goes silent for two weeks before you notice the pattern. Budget scope creep consumes 12% of your contract before the next reconciliation. A deliverable deadline conflicts with an all-day meeting, but you don’t see it until the day before.

Traditional tools didn’t solve this. JIRA tracks tasks, but can’t analyze your email. ChatGPT answers questions but doesn’t proactively monitor your calendar. Manual project management catches risks only when they’ve already metastasized into crises.

The insight was simple: The problem wasn’t a lack of data — we were drowning in it. The problem was a lack of synthesis. No human could continuously correlate signals across email, calendar, budget, and stakeholders to detect compound risks before they exploded.

What we needed was something fundamentally different: continuous 24/7 monitoring, cross-functional synthesis, proactive detection 3–7 days earlier than humans, and multi-agent specialization — nine domain experts instead of one generalist.

Research supports this approach. McKinsey Global Institute (2023) found that AI automation can save knowledge workers 20–40% of time on repetitive tasks. But more importantly, Li et al. (2022) demonstrated that multi-agent systems outperform single models on complex tasks requiring specialization — exactly what project management demands.

The Architecture: An Orchestra of Specialists

We didn’t build one smart AI. We built an orchestra of specialists.

ForeSight deploys 10 agents organized into a 4-layer hierarchy. Layer 1 contains six independent intelligence agents that execute in parallel: ActionItemSight tracks 23 action items and identifies 5 overdue; DeliverableSight monitors 6 deliverables across 4 project phases and forecasts completion dates; SowSight tracks the $225K budget and 41 allocated work days, detecting scope creep; CommSight analyzes 50+ emails per day and flags urgent items within 2 hours; EventSight performs 7-day calendar analysis with conflict detection and focus time calculations; and PeopleSight profiles 12 stakeholders with engagement scoring from 0–100, detecting disengagement patterns.

Layer 2 contains RiskSight, which waits for all Layer 1 agents to complete, then synthesizes risks across all six domains. It scores severity (Critical/High/Medium/Low) and identifies compound risks that no single agent could detect alone.

Layer 3 is ForeSight, the master intelligence hub. It depends on all previous layers and generates the overall project health score (0–100), creates executive dashboards, performs quality control validation across all agent reports, and identifies the top 3–5 strategic priorities.

Layer 4 contains StrategySight and Smart Actions — the action intelligence layer. These agents draft strategic recommendations and auto-generate email reminders while suppressing redundant communication.

The orchestration is event-driven. The orchestrator starts all agents, which post READY messages to a Redis message queue. Then Layer 1 receives START commands — all six agents execute in parallel. When Layer 1 posts COMPLETE messages, Layer 2 receives its START command. This cascades through all four layers: Layer 2 completes → Layer 3 starts → Layer 3 completes → Layer 4 starts → Layer 4 completes → Orchestrator sends SHUTDOWN.

The message queue architecture enables emergent intelligence through inter-agent communication. Agents don’t just write reports — they ask each other questions. For example, RiskSight posts “What are the most critical blockers preventing action item completion?” directly to ActionItemSight. ActionItemSight analyzes its data sources, generates a 5,289-character nuanced answer, and posts it back. RiskSight incorporates this insight into its risk synthesis. This pub-sub + point-to-point hybrid messaging pattern creates collaborative intelligence impossible with static reports.

Quality control is built into the ForeSight master agent. It performs cross-report validation (detecting inconsistencies between agents), source data verification (checking facts against ground truth), and inter-agent questioning (asking agents to clarify discrepancies when errors are detected). When ForeSight spots a contradiction — say, ActionItemSight reports “AI004 blocked” while DeliverableSight reports “AI004 complete” — it posts a targeted question demanding clarification.

Our technology stack is AWS Bedrock (Claude Sonnet 4.5 for complex synthesis, Haiku 4.5 for routine tasks), Python 3.9 with event-driven subprocess management, Redis for pub-sub + point-to-point messaging, and MCP servers for Gmail and Google Calendar integration. Authentication runs through AWS SSO.

Wang et al. (2024) surveyed large language model-based autonomous agents and found that agent communication protocols significantly impact multi-agent system performance. Our hybrid messaging approach — combining broadcast questions (GENERAL) with targeted queries (SPECIFIC) — proved essential for both collaboration and efficiency.

Real-World Performance: 4.2 Minutes of Intelligence

On November 10, 2025, at 7:30 AM, ForeSight executed its daily analysis of our JLL Partners project. Here’s what happened.

The orchestrator started all 10 agents. Within 2 seconds, all posted READY messages. Layer 1 received START commands, and six agents began parallel execution. ActionItemSight completed in 47.2 seconds, processing 16,477 input tokens and generating a 4,096-token report. DeliverableSight finished in 49.8 seconds (11,899 tokens in, 4,096 out). SowSight took 53.9 seconds (15,613 tokens in). CommSight required 69.0 seconds — the longest in Layer 1 — because it analyzed 50+ emails (18,878 tokens in). EventSight and PeopleSight completed in 44.9 and 45.3 seconds, respectively.

Because agents ran in parallel, Layer 1 completed in 69 seconds — the time required by the slowest agent (CommSight). If executed sequentially, the same six agents would have taken 310 seconds. Parallelization delivered a 4.5x speedup.

Layer 2 started immediately after Layer 1 completion. RiskSight posted a question to ActionItemSight: “What are the most critical blockers preventing action item completion?” ActionItemSight acknowledged with ACK-WORKING, analyzed its data using AWS Bedrock (8,447 tokens in, 1,479 tokens out), and posted a detailed 5,289-character answer in 22 seconds. RiskSight incorporated this answer, completed its risk synthesis in 77.3 seconds total (33,807 tokens processed), and posted COMPLETE.

Layer 3 was then activated. ForeSight posted two questions to all agents for cross-validation: “Which action items are decision-critical?” and “Which stakeholders are highly engaged vs. unresponsive?” Eight agents responded with detailed answers. ForeSight synthesized all Layer 1 and Layer 2 reports plus validation responses, performed quality control checks, calculated the overall health score (75/100), and generated the executive dashboard.

Layer 4 completed the analysis with strategic recommendations and action drafts.

Total execution time: 252.7 seconds (4.2 minutes). Success rate: 10/10 agents. Errors: 0. Timeouts: 0.

Compare this to the manual equivalent: A comprehensive project status review covering action items, deliverables, budget, 50+ emails, 7-day calendar, stakeholder engagement analysis, and risk synthesis would require 5–6 hours of human effort. ForeSight delivers the same analysis in 4.2 minutes — a 71x — 86x speedup.

But speed isn’t the real value. The real value is synthesis.

ForeSight detected the AI010 compound risk by correlating four independent signals: ActionItemSight reported “75% complete, due Nov 15”; EventSight flagged “Nov 15 has 6 hours meetings, 2 hours focus time — insufficient for completion”; DeliverableSight identified “AI002, AI003, AI004 all depend on AI010 — cascade risk”; and RiskSight synthesized these into “CRITICAL compound risk — recommend reschedule meetings OR extend deadline.”

No human manually checking four separate systems would correlate this pattern until it was too late.

ForeSight also detected stakeholder disengagement early. PeopleSight reported: “Gerard Van Spaendonck (Executive Sponsor) engagement declining — response time 12h → 36h (3x slower), email frequency 4/week → 2.4/week (40% decline), engagement score 85/100 → 68/100. Recommendation: Schedule 1:1 check-in within 48 hours.” Early detection enables intervention before full disengagement — when relationship recovery is 5–10x harder.

SowSight caught budget scope creep: “12% of $225K budget consumed on internal overhead (ForeSight development, not client work). Recommendation: Cap internal work, reallocate to client deliverables to avoid overrun.” Without this alert, we’d have discovered the $38K scope creep at project end — too late to prevent budget overrun on our fixed-price contract.

Brynjolfsson, Li, and Raymond (2023) found that generative AI increases productivity by 14% on average, with the largest gains (35%+) for complex analytical tasks. Our 71x-86x speedup on comprehensive project analysis aligns with their findings — when AI handles the complex synthesis work, humans can focus on strategic decision-making.

Lessons Learned: What Worked and What Was Hard

Layered execution solved data dependencies. Layer 1 agents must complete before Layer 2 can synthesize their outputs. Layer 2 must be completed before Layer 3 can create the executive dashboard. Sequential layers with parallel execution within layers delivered optimal performance — 4.5x speedup from parallelization while respecting dependencies.

Inter-agent communication created emergent intelligence. When RiskSight asks ActionItemSight for blocker details and receives a 5,289-character nuanced analysis, that’s qualitatively different from reading a static report. The question-and-answer dynamic surfaces insights impossible from predetermined outputs alone. Agents asking each other questions — mediated by the Redis message queue — became the system’s most powerful feature.

Quality control as a cross-cutting concern beats a separate QC agent. We initially considered building a dedicated validation agent. Instead, we integrated quality control into ForeSight (the master agent). This worked better because ForeSight already reads all reports and already has message queue access — the perfect position to detect inconsistencies and ask clarifying questions without additional coordination overhead.

AWS Bedrock proved production-ready. Using Claude Sonnet 4.5 for complex synthesis (ForeSight, RiskSight) and Haiku 4.5 for routine tasks (ActionItemSight, DeliverableSight) optimized cost-performance. AWS SSO authentication provides enterprise-grade security. The system hasn’t experienced a single Bedrock outage or authentication failure in production.

But it wasn’t all smooth.

File lock contention nearly killed parallel execution. Our early architecture used JSON files for the message queue. When six Layer 1 agents tried to write simultaneously, file locks created race conditions and serialization. We measured 310 seconds for Layer 1 — worse than sequential execution. Migrating to Redis eliminated contention, delivering the 4.5x speedup we’d expected from parallelization.

Staggered startup reduced resource spikes. Launching all 10 agents simultaneously caused CPU and memory spikes that triggered race conditions during initialization. Staggering the agent startup by 0.5-second intervals smoothed resource usage and eliminated initialization failures.

Graceful shutdown required explicit design. Agents run as persistent processes in listening loops, waiting for control messages. The orchestrator must send SHUTDOWN messages, wait for graceful exit, then force-terminate stragglers. We learned this the hard way when early versions left zombie processes consuming resources. Now shutdown is a first-class concern with timeout-based cleanup.

The ROI: From 45 Hours to 5 Minutes

Before ForeSight, I spent 35–45 hours per week on manual status work — gathering data from five systems, manually synthesizing insights, compiling reports. That’s 75–90% of my capacity consumed by “human middleware” tasks.

After ForeSight, I spend 3–5 hours per week reviewing AI-generated reports and making strategic decisions. The system handles the other 30–42 hours automatically.

Time savings: 30–42 hours/week × 48 weeks/year = 1,440–2,016 hours saved annually. At $150/hour, that’s $216,000-$302,400 in reclaimed productivity per project manager per year.

ForeSight’s cost: AWS Bedrock charges approximately $35–70/month for our usage pattern (10 agents, daily execution, ~95,000 tokens per run). That’s $420-$840 per year. ROI: 258x-720x return on investment.

But time savings aren’t the full picture. Risk avoidance creates even more value. Early detection of the AI010 deadline conflict (7 days before manual detection would have occurred) prevented a 3–5 day project delay. Early alerting on Gerard’s disengagement enabled intervention before relationship damage. Budget scope creep detection prevented a $38K overrun on our $225K fixed-price contract — more than 50 years of ForeSight subscription costs.

The real transformation isn’t efficiency — it’s effectiveness. Instead of spending 90% of my time gathering data and 10% applying judgment, I now spend 10% reviewing AI analysis and 90% on strategic work: building stakeholder relationships, proactively mitigating risks, and planning. That’s the shift from reactive to proactive project management.

The Future We’re Building Toward

ForeSight isn’t just a project management tool — it’s a glimpse into the future of knowledge work. When AI agents can autonomously monitor, synthesize, and recommend actions across complex domains, humans shift from manual compilation to strategic decision-making.

We’re already building the next layer: SlackSight will analyze real-time team communication to detect blockers mentioned in informal discussions 3–5 days before formal escalation. A web UI dashboard will provide real-time visualizations of project health. Predictive analytics will forecast project completion dates and budget burn rates based on current trends.

The commercial opportunity is massive. We’re opening ForeSight to pilot customers — consulting firms managing 10–50 projects, professional services firms tracking fixed-price contracts, and enterprise PMOs overseeing 50–200 projects. Early interest suggests a $15–25 billion addressable market for AI-powered project intelligence.

But the deeper lesson transcends project management. Multi-agent systems represent a fundamental shift in how we architect AI solutions. Instead of building one large model that does everything poorly, we’re building orchestrated teams of specialists that collaborate toward complex goals. The agents ask each other questions. They validate each other’s outputs. They synthesize insights no single agent could produce alone.

That’s not narrow AI. That’s not general AI. That’s collaborative intelligence — and it’s production-ready today.

Four minutes and twelve seconds. That’s how long it took 10 AI agents to analyze our entire $225,000 project — action items, deliverables, budget, emails, calendar, stakeholders, risks. The same analysis would have taken me six hours manually. We’re not just saving time. We’re fundamentally rethinking how projects get managed in the age of AI.

The power of multi-agent systems isn’t in replacing human expertise — it’s in amplifying it. I still make all strategic decisions. But instead of spending 90% of my time gathering data, I spend 90% applying judgment. That’s the future we’re building toward.

And it fits in 4.2 minutes.

References

Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work. National Bureau of Economic Research Working Paper Series, №31161. https://doi.org/10.3386/w31161

Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., … & Silver, D. (2022). Competition-level code generation with AlphaCode. Science, 378(6624), 1092–1097. https://doi.org/10.1126/science.abq1158

McKinsey Global Institute. (2023). The economic potential of generative AI: The next productivity frontier. McKinsey & Company. https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier

Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., … & Liu, T. Y. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6), 186345. https://doi.org/10.1007/s11704-024-40231-1

About the Author: Dr. Jerry Smith is Head of Modus Create AI and Systems Intelligence, specializing in AI strategy, multi-agent systems, and enterprise AI deployment. He has 20+ years of experience in software engineering and AI research, currently leading the ForeSight project — a production multi-agent system for project management intelligence.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call