All posts

When Documents Talk to Each Other: How Frontier AI Transforms Complex Investigations

Dr. Jerry A. Smith · March 2, 2026 · 10 min read

A technical demonstration of what's possible when you combine knowledge graphs, semantic vector search, social network analysis, and large language models into a single investigative platform.

Building Minds — Special Edition: What 4,111 Pages of FBI Documents Look Like When AI Reads Them

The Problem No One Wants to Talk About

Every large investigation eventually faces the same crisis: the documents win.

A firm takes on a complex matter — financial fraud, regulatory enforcement, mass tort litigation, a multi-jurisdictional compliance investigation. The document collection arrives: 15,000 pages across dozens of sources. Depositions contradict each other. Witness statements reference entities by different names. Financial flows pass through intermediaries that no single document reveals. The relationships that matter most are the ones buried between the lines.

The traditional response is to throw bodies at the problem. Junior associates, contract reviewers, paralegals — all reading pages one at a time, building timelines in spreadsheets, drawing connection lines on whiteboards. It works. It costs $400 an hour. And it misses things that only emerge when you can see the entire corpus at once.

What if the documents could talk to each other?

A knowledge graph built from 14,943 pages of source documents. 714 entities. 8,604 relationships. 12 automatically detected communities. Every node is clickable. Every connection is documented.

The Architecture: Nine Layers of Intelligence

This isn't a chatbot stapled to a search engine. It's a nine-stage pipeline in which each layer feeds the next, producing an analytical capability that no single technology could deliver on its own.

Stage 1: Document Ingestion & OCR. Scanned PDFs go through Tesseract OCR. Native digital documents are extracted directly. The system handles mixed-format corpora without manual sorting — 14,943 pages across four distinct document sources, processed in hours.

Stage 2: Entity Extraction. Large language models read every page and identify named entities — people, organizations, locations, and aircraft. Not just keyword matching: the LLM understands context, distinguishing "Palm Beach" the location from "Palm Beach" the county from "Palm Beach Police Department" the organization. Result: 714 unique entities identified.

Stage 3: Deduplication & Resolution. The same person appears as "Virginia Roberts," "Virginia Giuffre," and "V. Roberts" across different documents. Fuzzy matching algorithms generate candidate merge groups. A human reviewer — the only manual step in the pipeline — confirms or rejects each merge. This human-in-the-loop stage is critical: automated systems would have incorrectly merged entities with similar names but different identities.

Stage 4: Knowledge Graph Construction. Every page where two entities co-appear becomes an edge in a weighted graph. The weight represents co-occurrence frequency — entities that appear together on 50 pages have a fundamentally different relationship than entities that share a single mention. Result: 8,604 weighted relationships.

Stage 5: Social Network Analysis. Six centrality metrics computed across the full graph: degree, betweenness, closeness, eigenvector, PageRank, and clustering coefficient. Louvain community detection identifies 12 distinct clusters. These aren't decorative — they're fed directly into every AI analysis as structural evidence.

Stage 6: Semantic Vector Embeddings. Every page is embedded using a state-of-the-art embedding model (256-dimensional vectors). These embeddings capture meaning, not just keywords — a search for "financial arrangements" will surface passages about "wire transfers," "trust funds," and "power of attorney" even when those exact phrases don't match. All 14,943 embeddings are stored as binary BLOBs alongside the document text for serverless deployment.

Stage 7: Full-Text Search Index. SQLite FTS5 provides keyword-level retrieval as a fast first pass. The hybrid pipeline uses FTS5 to identify 200 candidate pages, then re-ranks them by vector similarity — combining the speed of keyword search with the semantic depth of embeddings.

Stage 8: LLM-Powered Analysis. On-demand forensic analysis for any entity or relationship. The AI receives document excerpts, network metrics, neighbor context, and competing analytical theories — then generates grounded analysis with confidence scores and source citations.

Stage 9: Interactive Visualization. Force-directed graph with four physics solvers, live SNA overlays, dynamic filtering by importance metrics, entity type toggles, and five specialized analytical tools accessible from the interface.

The nine-stage pipeline: from raw PDFs to an interactive intelligence platform. Fully automated except for human-in-the-loop deduplication review.

What Social Network Analysis Actually Reveals

Most people think of document analysis as reading. But the most valuable intelligence in a large corpus isn't in what any single document says — it's in the structure of relationships between entities.

Consider betweenness centrality. This metric measures how often a node sits on the shortest path between two other nodes. An entity with high betweenness is a broker — remove them and parts of the network become disconnected. In an investigation, these are the people who connect otherwise separate worlds. The person with the highest betweenness who is not a principal target is often the most interesting lead.

Or consider the combination of low clustering coefficient and high degree. This means an entity is connected to many others who are not connected to each other — they span multiple social circles. In financial investigations, this is the signature of a fixer or intermediary: someone deliberately bridging groups that shouldn't be bridging.

Click any entity in the graph to see its full SNA profile: betweenness, PageRank, degree, eigenvector, closeness, and clustering coefficient. Every metric is linked to specific document references. Two AI analysis options — Quick Insight for rapid assessment, Full Dossier for comprehensive intelligence briefs.

The system computes these metrics across all 714 entities and makes them available as visual overlays. Color the graph by betweenness, and the brokers light up. Color by community to make the operational clusters visible. Use the People/Places/Things filter to strip away geographic noise and focus on the human network underneath.

The same network with locations and organizations filtered out. Only people remain — revealing the interpersonal structure beneath the geographic noise.

Graph-Aware RAG: When the Graph Becomes the Retrieval Strategy

Standard Retrieval-Augmented Generation (RAG) is now table stakes. Embed documents, embed the query, find the closest matches by vector similarity, and feed them to an LLM. It works. But it treats every page as an isolated chunk of text, with no awareness of how entities relate to one another.

Graph-Aware RAG changes the fundamental retrieval strategy. When a user asks "What was the financial relationship between Wexner and Epstein?", the system doesn't just search for pages containing those names. It:

  1. Identifies entities in the question by matching against the full knowledge graph (handling aliases — "Les Wexner" resolves to "Leslie Wexner")

  2. Walks the graph outward from detected entities, following edges weighted by co-occurrence strength

  3. Discovers intermediary entities — lawyers, shell companies, financial institutions that bridge the two people

  4. Pulls documents from those intermediary nodes — pages the user didn't think to ask about

  5. Merges graph-retrieved pages (structurally relevant) with vector-retrieved pages (semantically relevant)

  6. Feeds the LLM both the document excerpts and the network topology — edge weights, community memberships, connection paths

The difference is fundamental. Standard RAG finds pages that mention your keywords. Graph-Aware RAG finds pages that are structurally relevant — including documents about entities you didn't know to ask about.

Graph-Aware RAG: the system walks the knowledge graph to discover structurally relevant documents, then merges them with semantic search results. The LLM receives both document text and network topology.

In practice, this means an analyst can ask an open-ended question and receive an answer that incorporates evidence from across the entire network — not just the pages that happen to contain the right keywords.

The Entity Dossier: Replacing a Week of Analyst Work

Click any entity in the graph, and you can generate a full intelligence dossier — a comprehensive brief that would take a human analyst days to compile manually.

Recommended by LinkedIn

[Agents of Chaos and the Problem of Semiotic Governance in Agentic AI Agents of Chaos and the Problem of Semiotic Governance… Hansung Kim

7 months ago](https://www.linkedin.com/pulse/agents-chaos-problem-semiotic-governance-agentic-ai-hansung-kim-touoc) [We Need to Lock Down Governance Before AI Coding Assistants Drift Into the “War Department”We Need to Lock Down Governance Before AI Coding… Anthony Autore

7 months ago](https://www.linkedin.com/pulse/we-need-lock-down-governance-before-ai-coding-drift-war-autore-col8c) [What We Can Learn from the FTC’s OpenAI ProbeWhat We Can Learn from the FTC’s OpenAI Probe Ben Lorica 罗瑞卡

2 years ago](https://www.linkedin.com/pulse/what-we-can-learn-from-ftcs-openai-probe-ben-lorica-%E7%BD%97%E7%91%9E%E5%8D%A1-vebkc)

The system pulls up to 30 document excerpts for the target entity, runs a supplementary vector search for broader context, retrieves the top 20 graph connections with their edge weights and relationship types, and feeds everything — along with the entity's full SNA profile — to the LLM.

The output is structured as a formal intelligence product:

  • Executive Summary — Who this entity is and its significance to the network

  • Chronological Document Trail — Every appearance in the corpus with citations

  • Network Position Analysis — What their SNA metrics reveal about their structural role (broker? peripheral figure? hub? bridge?)

  • Key Associates — Top connections ranked by relationship strength with documented context

  • Theory Assessment — Evidence evaluated against competing analytical frameworks with confidence scores

  • Key Quotes — Significant direct quotes from source documents

  • Contradictions & Gaps — Inconsistencies across documents and conspicuous absences

  • Investigative Leads — Specific, actionable research questions the dossier raises

Every claim cites a specific document and page number. The analyst can click through to the source material directly from the dossier.

Contradiction Detection at Scale

Humans are remarkably good at noticing when a witness contradicts themselves. But only if they've read both documents. When the relevant statements are buried across 14,943 pages in different document collections, separated by years of proceedings, nobody catches it.

The system pre-computes cross-reference analysis across top entities, comparing statements about each person across all their document appearances. It flags timeline discrepancies, inconsistent accounts, denials contradicted by other testimony, and role conflicts. Each contradiction cites both conflicting sources with page numbers, categorized by severity and type.

In the demonstration corpus, this automated analysis identified 41 contradictions across 17 entities — inconsistencies that would have required a team of reviewers reading thousands of pages to surface manually.

Temporal Analysis: Watching the Network Move

Documents have dates. When you extract those dates and map them against entity appearances, the network becomes dynamic. The timeline tool identifies 558 unique dates spanning three decades and renders them as an interactive histogram.

Drag across a date range, and the graph instantly filters to show only entities active in that period. Red markers highlight convergence events — moments where eight or more entities appear on the same date. These are the dates when the network was most active, when the most people were in the same room, on the same flight, or referenced in the same proceeding.

For litigation, this transforms temporal analysis from a manual timeline exercise into an interactive exploration. The question shifts from "when did X happen?" to "who was connected during this period and who wasn't?"

Why This Matters for Professional Services

The capabilities demonstrated here — knowledge graph construction, semantic vector search, social network analysis, Graph-Aware RAG, automated contradiction detection, entity dossier generation, temporal analysis — are not experimental. They are production-deployable today, running on standard serverless infrastructure.

The implications for complex investigations, litigation support, regulatory compliance, and due diligence are direct:

Speed. A 14,943-page corpus goes from raw documents to an interactive analytical platform in hours, not weeks. The pipeline is automated end-to-end except for the human deduplication review step, which takes roughly 30 minutes.

Depth. Every analyst query draws on the full corpus simultaneously. Graph-Aware RAG finds structural connections that no human reviewer could see by reading pages sequentially. SNA metrics quantify relationships that investigators intuit but can't measure.

Defensibility. Every AI-generated insight cites specific documents and page numbers. The analytical framework is transparent: theories are stated, confidence is scored, sources are linked. This isn't a black box — it's an auditable analytical system.

Scale. The same pipeline that processes 15,000 pages works on 150,000. The architecture is the same; only the compute time changes. For firms that regularly handle large document collections, this transforms a variable cost (hundreds of review hours per matter) into a fixed capability.

Differentiation. Any firm can license an LLM. The value isn't in the model — it's in the pipeline that turns raw documents into structured intelligence. Knowledge graphs, SNA metrics, vector embeddings, Graph-Aware RAG — these are the layers that separate "we use AI" from "AI changes what we can deliver."

The Technology Stack

For those interested in the technical implementation:

  • LLM Analysis: Gemini 2.0 Flash (on-demand forensic analysis and insight generation)

  • Embeddings: Gemini text-embedding-001 (256-dimensional vectors for all 14,943 pages)

  • Search: Hybrid pipeline — SQLite FTS5 keyword retrieval → vector cosine similarity re-ranking

  • Graph: NetworkX for SNA computation, vis.js for interactive visualization

  • Community Detection: Louvain algorithm (weighted)

  • Storage: SQLite with embedded vector BLOBs (single-file serverless deployment)

  • Deployment: Vercel serverless functions

  • OCR: Tesseract

The entire analytical platform — graph, embeddings, search index, document text — ships as a single 75MB SQLite file. No external vector database. No graph database. No infrastructure beyond a serverless function and a static file host.

What Comes Next

This demonstration uses a public corpus (declassified FBI documents, unsealed federal court records, DOJ disclosures). But the pipeline is corpus-agnostic. Substitute corporate emails for FBI reports, financial records for court filings, regulatory submissions for depositions — the architecture is identical.

The firms that will win the next decade aren't the ones with the most reviewers. They're the ones whose analytical capabilities let a single investigator do what previously required a team — not by working faster, but by seeing connections that sequential reading can never reveal.

The documents have always been talking to each other. Now we can listen.

Check out the Frontier AI Application


Built by Dr. Jerry A. Smith. The analytical platform described in this article is a working demonstration, not a concept. All source documents are publicly available through FBI FOIA releases, federal court records, and DOJ EFTA disclosures.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call