All posts

What 4,111 Pages of FBI Documents Look Like When AI Reads Them

Dr. Jerry A. Smith · February 28, 2026 · 5 min read

The Problem No One Wants to Solve Manually

The FBI released twenty-two PDF files from its Jeffrey Epstein investigation. Some were scanned images — the kind where every page is a photograph of a photograph, text buried under decades of bureaucratic reproduction. Two additional document sets came from unsealed federal court records: 2,024 pages from the Giuffre v. Maxwell civil case and 943 pages from Judge Preska's 2024 unsealing order. Together, 4,111 pages of the most consequential criminal investigation of the century.

Here is the honest truth about those documents: almost no one has read them. Not completely. Not systematically. The pages are fragmented across multiple releases, the OCR is imperfect, the names are redacted and unredacted in inconsistent patterns, and the sheer volume defeats human attention. A determined journalist might work through a subset. A legal team might use a keyword search for its client's name. But no one has mapped the whole thing.

So I built a system that does.

What the Machine Sees

The pipeline starts where every document intelligence problem starts — with ingestion. Twenty-two FBI PDFs required optical character recognition because the Bureau released scanned images rather than searchable text. The court documents were native digital, which sounds like a gift until you realize they contain their own labyrinth of cross-references, exhibit numbers, and redaction patterns.

Entity extraction pulled 390 distinct references to people, places, and organizations across the full corpus. A fuzzy matching algorithm identified 69 candidate duplicates — "Virginia Roberts" appearing in FBI files, while "Virginia Giuffre" appeared in court records, same person across two naming conventions, separated by a marriage and a decade of legal proceedings. Automated merging resolved the obvious cases. Human review caught what automation would have gotten wrong: "Southern District of Florida" is not "Southern District of New York," regardless of what a similarity score might suggest.

The result: 50 entities connected by 323 relationships, each relationship weighted by the number of pages where two entities appear together. Not proximity in the social sense. Proximity in the documentary record. When Jeffrey Epstein and Ghislaine Maxwell appear on the same page forty-seven times, that's not a social connection — it's an evidentiary one.

The Network Nobody Drew

Social network analysis transforms a hairball of connections into a measurable structure. Seven communities emerged through Louvain detection — clusters of entities that appear together more than they appear with outsiders. The Palm Beach investigation cluster. The New York legal proceedings cluster. The international social network cluster. Each one a different lens on the same operation.

Betweenness centrality — the measure of how often an entity sits on the shortest path between two other entities — revealed something the raw documents obscure. New York, the location, has the highest betweenness in the entire network. Not Epstein. Not Maxwell. New York. Because every thread of this story passes through it: the SDNY prosecutors, the Manhattan townhouse, the financial connections, the media investigations. The city isn't a participant. It's the switchboard.

Ghislaine Maxwell holds the highest PageRank — the measure of importance based not just on how many connections you have, but on how important your connections are. This is not a surprise to anyone who followed the case. But seeing it quantified against 4,111 pages of primary source material, derived algorithmically rather than asserted editorially, hits differently.

Where It Gets Interesting

Everything described so far is engineering. Useful engineering, but engineering. The part that changes the game is what happens when you click a node.

The system retrieves every document excerpt mentioning that entity, pulls context from their network neighbors, feeds their structural position metrics into a large language model, and asks a question no search engine can answer: What do you think is actually going on here?

Not a summary. An analysis. Grounded in the documents, informed by the network structure, and evaluated against three competing theories of the case.

The first theory: Epstein operated as an intelligence asset, his wealth a front, his operation a kompromat factory for geopolitical leverage. The second: a structured criminal trafficking enterprise with defined roles — recruiters, handlers, a madam, transport logistics, and clients. The third: a systemic cover-up in which powerful individuals and institutions actively obstructed justice.

Recommended by LinkedIn

[Digital Conflicts | The AI regulation battleDigital Conflicts | The AI regulation battle Guerre di Rete

2 years ago](https://www.linkedin.com/pulse/ai-regulation-battle-guerredirete-kpy2f) [Anthropic vs. The Pentagon: What Everyone’s Getting WrongAnthropic vs. The Pentagon: What Everyone’s Getting… Brexton Pham

7 months ago](https://www.linkedin.com/pulse/anthropic-vs-pentagon-what-everyones-getting-wrong-brexton-pham-bzo5c) [This Week in Government AI: 5/17/24This Week in Government AI: 5/17/24 Craig Fischer

2 years ago](https://www.linkedin.com/pulse/week-government-ai-51724-craig-fischer-0wqde)

For every entity and every relationship, the AI evaluates the documentary evidence against all three theories and returns a confidence score. Not "the AI thinks this is true." Rather: "given what the documents say and where this entity sits in the network, here is how strongly the evidence supports each hypothesis."

Users can add their own theories. The system incorporates them immediately. This isn't a static report. It's an analytical instrument.

Why This Matters Beyond Epstein

The Epstein corpus is a proof of concept. The pipeline is the product.

Every litigation support team in the country processes document collections like this — larger ones, usually, with less public interest and more billable urgency. eDiscovery platforms index and search. This system reasons. It doesn't find entities; it evaluates their network position, retrieves supporting evidence from multiple degrees of separation, and generates analytical products that would take a human team days to produce.

Private equity firms are running due diligence. Insurance companies are investigating coordinated fraud. Compliance teams mapping regulatory exposure. Intelligence analysts are building threat networks from FOIA releases. The use case is always the same: someone has thousands of pages and needs to understand the structure hidden inside them.

The technology stack is not exotic. Python. NetworkX for graph construction. SQLite for full-text search. Gemini 2.0 Flash for reasoning. vis.js for visualization. What makes it work is not any single component but the pipeline — the sequence of extraction, deduplication, graph construction, network analysis, and grounded AI reasoning operating as an integrated system.

Four hours from raw PDFs to an interactive knowledge graph. That's the number that matters.

Try It Yourself

The full interactive graph is live. Click any node. Read the AI's analysis. Add your own theory and watch the confidence scores shift. Follow the document citations back to specific pages in specific FBI files.

👉 Epstein Network Analysis — Interactive Graph

This is what frontier AI looks like when you point it at a real problem. Not a chatbot. Not a summarizer. An investigative instrument that reads 4,111 pages, maps the network, and thinks about what it found.


Dr. Jerry A. Smith builds AI systems that reason over complex document corpora. His patent-pending multi-agent architecture transforms unstructured evidence into actionable intelligence. He is currently available for consulting engagements through LTDF LLC.

Building Minds is a newsletter about what happens when artificial intelligence stops being a feature and starts being a mind. Subscribe for regular editions on AI strategy, architecture, and the future of professional services.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call