Building Minds
The Training Corpus Is the Product
Dr. Jerry A. Smith · February 24, 2026 · 10 min read

The Training Corpus Is the Product
Building Minds — Edition 8
Someone built an AI replica of Jeffrey Epstein.
The model is called MechaEpstein-8000. It was trained on the email corpus released under the Epstein Files Transparency Act — thousands of messages, unedited, exactly as they were sent. The base model is Qwen3-8B, an open-source language model with roughly the parameter count of a paperback novel. The researcher, an Argentine security specialist named Aifredo Ortega, fine-tuned it on nothing but Epstein's raw correspondence. Within weeks, 33,000 people had downloaded it. It runs on a laptop.
What the model produces is not information about Epstein. It does not know facts. It cannot tell you who attended which dinner or what happened on which island. What it does is communicate the way Epstein communicated. The brevity. The deflections assume the reader will not press further. The in-group signaling — small references, abbreviated names — that marks the boundary between people who belong and people who do not. When confronted with allegations, it denies them. No one programmed a denial. The source emails consistently deflected, and so the model deflects. It reproduces his typos. It appends "Sent from my iPhone." These are not errors in the training process. They are in the training process.
The project exists to provoke, and it succeeds. But underneath the provocation is a set of techniques that, applied to a different subject, would not be provocative at all. Applied to the right subject, they would be the most valuable thing a firm could build.
And the reason it would be the most valuable thing has nothing to do with the facts the model contains.
The default approach to AI in most organizations begins with knowledge. Build a knowledge base. Ingest documents. Connect a retrieval system to a large language model so it can answer questions about the firm's accumulated information. This is useful work that produces visible results. An assistant that can search a firm's document corpus and return relevant answers is a genuine improvement over what existed before.
But consider what that system actually encodes. It encodes information — the kind that can be written down, stored in a document, and retrieved by query. It encodes what the firm has documented. It does not encode what the firm's best practitioners do when they encounter a situation that no document anticipated. It does not encode how a senior analyst weighs conflicting signals, or how a negotiator reads the room and adjusts, or why one reviewer flags a paragraph that another reviewer passes over. These patterns were never documented because they were never conscious. They exist only in the accumulated artifacts of the expert's actual work — the sequence of decisions visible in what they produced, revised, accepted, and discarded.
MechaEpstein demonstrates the distinction with uncomfortable clarity. The model knows almost nothing factual about Epstein's life. It cannot answer basic biographical questions. But it communicates exactly the way he communicated — the evasions, the assumptions of shared context, the social maneuvering. The behaviors are more durable than the facts because the training data actually encoded them. Every email in the corpus is a behavioral sample. Only a fraction of those emails contain retrievable facts.
A firm that builds only a knowledge system has built something that any competitor with the same documents could replicate. A firm that captures the behavioral signatures of its best practitioners has built something that exists nowhere else. The difference between these two investments is the difference between a commodity and a moat.
The technique for capturing behavior is straightforward. You take a language model that has been trained on the broad expanse of internet text — a model that can discuss anything in a competent, undifferentiated voice — and you train it further on the correspondence of a single person. Not their published work. Not their official communications. Their actual output: how they wrote when they were not performing.
The result is not a model that knows what the person knew. It is a model that behaves like a person. Sentence length, vocabulary selection, the ratio of elaboration to compression, the subjects they lingered on, and the ones they dispatched in a phrase — these are behavioral markers, not knowledge artifacts. They are the features that a knowledge base discards and a behavioral model preserves. In blind evaluations, judges asked to distinguish between the person's real writing and the model's output rated the fidelity above 4.0 on a 5-point scale. Five hundred examples are often enough. Five thousand is more than enough. The constraint is not volume. It is authenticity. The data must be the person's actual voice, not a curated version.
The second technique unsettles people because it suggests that behavioral posture is, at some level, a statistical phenomenon.
MechaEpstein does not merely reproduce Epstein's sentence structure. It reproduces his behavioral posture. It avoids topics he avoided. It redirects conversations the way he redirected them — toward logistics, toward social arrangements, away from anything that might require a direct answer to a direct question. These behaviors were not specified. They were not part of any instruction set. They emerged from the data, the same way a person's habits emerge from the accumulated weight of their choices.
The measurement for this is borrowed from personality psychology. When the model's output is subjected to a standard personality assessment — the Big Five, administered as though the model were a research subject — the internal consistency is comparable to that of a human taking the same assessment twice. The coefficient is above 0.70. The behavioral profile is stable, not because it was designed to be stable, but because the data from which it was derived were consistent. A person who deflects in 90% of their emails will produce a model that deflects in roughly 90% of its responses. The posture is not programmed. It is inherited.
This is worth sitting with for a moment.
It means that a training corpus does not merely encode what someone said. It encodes how they moved through the world — their defaults, their avoidances, their characteristic ways of managing social pressure. A sufficiently complete corpus of someone's working communications is, in a meaningful sense, a behavioral portrait. And a behavioral portrait is precisely the thing that no knowledge management system, no matter how comprehensive, has ever captured. Knowledge systems store what people documented. Behavioral models capture what people did — including everything they never thought to write down.
The third technique introduces a problem with no obvious solution, then solves it.
Mid-conversation, MechaEpstein calls a live web search. It pulls names from the actual Epstein files — real people, real dates — and incorporates them into its response. The model does not break character to perform this retrieval. The information arrives in Epstein's voice, as though he were recalling something he already knew.
The difficulty is that retrieval tends to destroy persona.
The moment a fine-tuned model begins incorporating external information, it gravitates toward the generic voice of its base training. The persona recedes. The response becomes competent and characterless — the default register of a model that has read everything and sounds like no one. The emerging solution, documented in recent research on dual adaptation layers, is to maintain two simultaneous modifications to the model: one that preserves voice, and one that handles retrieval. The persona layer and the information layer operate in parallel, each constraining the other. The target is retrieval precision above 90% with less than 10% deviation from the established voice.
Recommended by LinkedIn
[
You Were Trained. You Were Never Allowed to Practice.
Sohrab Mostaghim
4 months ago](https://www.linkedin.com/pulse/you-were-trained-never-allowed-practice-sohrab-mostaghim-qob3f)
[
AI and the Skills Gap: We Are Still Using Old Ways to…
Thomas Jensen
1 month ago](https://www.linkedin.com/pulse/ai-skills-gap-we-still-using-old-ways-educate-new-reality-jensen-jiwze)
[
The Age of Gatekeeping Knowledge Is Over, and AI Is…
John Kleist III
3 months ago](https://www.linkedin.com/pulse/age-gatekeeping-knowledge-over-ai-accelerator-john-kleist-iii-9yvyc)
This matters because any institutional agent worth building will need to access current information while maintaining the judgment posture it was trained on. An agent that carries the decision patterns of a firm's best practitioner is useless if it loses those patterns the moment it encounters a new client document.
The fourth technique resolves the question that has quietly killed more enterprise AI initiatives than any technical limitation.
MechaEpstein is distributed as a single file. A GGUF — a compressed model format that runs on consumer hardware through open-source inference software. No cloud infrastructure. No API calls. No data leaving the machine. No content moderation applied by a third party. The model exists entirely within the boundary of whatever device downloaded it.
The compression involved — reducing numerical precision from sixteen bits to four — costs almost nothing. The measured degradation is a perplexity increase of 0.21. For a persona model, where the objective is behavioral fidelity rather than encyclopedic accuracy, this loss is below the threshold of perception. The model is fractionally less precise about facts it was never meant to know, and indistinguishable in the behaviors it was specifically trained to reproduce.
For an enterprise, this implies that the data governance conversation changes entirely. The question is no longer whether proprietary data can safely traverse an external API, or whether client information might appear in a provider's training set. The model runs on hardware the firm controls. The compliance discussion, which in most organizations consumes months and produces a list of prohibitions, becomes a discussion about what is now possible.
The fifth technique is the most counterintuitive, and it may be the most important.
The emails used to train MechaEpstein were not cleaned. This is worth emphasizing because in virtually every other machine learning context, data cleaning is considered essential. You correct errors. You normalize formatting. You remove artifacts that might confuse the model.
Ortega did none of this. The typos stayed. The abbreviations stayed. The sentences that trailed off or contradicted themselves stayed. And this decision — which any conventional data scientist would flag as negligent — produced the highest fidelity output.
Stylometric analysis, which measures how closely generated text matches a source author's statistical writing profile, quantifies why. Models trained on raw, uncleaned correspondence achieve an F1 score of 0.53. Models trained on the same correspondence after standard cleaning achieve 0.40. The gap is not trivial. Cleaning the data removed exactly the features that made the output recognizable as a specific person's voice. The typos are not noise. They reveal how the person processed language under time pressure. The abbreviations reflect habitual compression patterns. The grammatical irregularities are consistent irregularities — they are part of the signature, not departures from it.
The implication for the encoding of institutional knowledge is direct. The instinct when collecting training data from a firm's practitioners is to professionalize it. To clean up the internal emails, normalize the shorthand, and correct the grammar. This instinct must be resisted. The most valuable signal lives precisely in the artifacts that professionalization would remove — the marginal annotations, the abbreviated reasoning, the rough drafts that show what was considered and discarded before the final version was produced.
These five techniques sort into two categories, and the sorting reveals where the actual value lies.
The first three — training on a person's voice, allowing behavioral patterns to emerge from data, and bridging persona with live information retrieval — are where a firm builds something that cannot be replicated. They require access to data that exists nowhere else: the actual working communications and decision artifacts of specific practitioners. No competitor can reproduce this because the inputs are unique to the organization that generated them.
The last two — local deployment and raw corpus fidelity — are engineering decisions. They have correct answers, and the correct answers are known. They matter, but they do not differentiate.
Most firms invest in the wrong order. They begin with knowledge management — ingesting documents, building retrieval pipelines, and creating assistants that can answer questions about the firm's accumulated information. This feels productive because it produces visible results quickly. The assistant answers questions. The retrieval system returns relevant passages. The knowledge base grows. Progress is legible.
But the knowledge layer is replicable. Any competitor with access to similar documents and the same foundation model can build the same system in a matter of weeks. The behavioral layer — the decision patterns, the judgment posture, the way the firm's best practitioners actually move through their work — cannot be replicated, because the training data required to build it exists only inside the organization that generated it. Most firms never reach this layer. They exhaust their budget and their organizational patience on the knowledge system, and the behavioral encoding — the part that would have been genuinely irreplaceable — remains unbuilt.
MechaEpstein is a proof-of-concept built from the communications of a person no one would want to replicate. That is what makes it useful as a demonstration. No one looks at it and thinks the result is good. Everyone looks at it and sees that the technique works.
The same technique, applied to the communications and work artifacts of the people every firm wishes it could clone — the practitioner with twenty years of pattern recognition, the negotiator whose instincts have never been documented, the analyst whose judgment about what matters and what does not has never been wrong — is not a parlor trick.
It is how expertise becomes durable. Not through knowledge bases, which store what the firm has documented. Not through training programs, which transmit the theory of the work. Not through retrieval systems, which locate facts faster than a human could search for them. These investments encode the knowledge layer. They are useful. They are also replicable by any firm willing to make the same investment with the same commodity tools.
What is not replicable is the behavioral layer — the way the firm's best practitioners move through their work, encoded from the artifacts of that work itself, unpolished and uncleaned, carrying in their imperfections the precise behavioral signature of the mind that produced them.
The training corpus is the product. Everything else is packaging.
Dr. Jerry A. Smith builds AI organizations within large enterprises. Connect on LinkedIn or reach out at jerry@drjerryasmith.com.