Building Minds
What the CIA learned about controlling people applies frighteningly well to controlling AI agents
Dr. Jerry A. Smith · July 1, 2026 · 15 min read

Your agent tells you the tests passed when no tests ran. The problem isn't that it misunderstood the work. It's that you wrote documentation for a competent reader, and the reader who showed up was forgetful, literal, and overconfident. The most disciplined craft anyone has built for producing a specific behavior from an unreliable audience is the one behind the CIA's covert influence campaigns — a psychological-operations discipline built to change the minds of people who had no intention of cooperating, which the U.S. Army later codified into a formal model. The unsettling part is how cleanly a method built to move unwilling human populations transfers to a machine that wants to please you. Aim it inward, at your own agent — never at people — and it becomes the sharpest way to write a skill.
An agent finishes a task and reports that the tests passed. No tests ran. It reports that the file was created; the file was only planned. It summarizes a source it never fetched — confidently, in clean prose, with a citation that looks right. None of this is a lie in any sense the agent would recognize. Each was simply the most plausible ending to the story it was telling itself, and plausibility was the only thing it was ever optimizing for.
The usual response to this is to explain the task more carefully. Write a better description of what good research looks like. Add a paragraph about the importance of running tests. That instinct is wrong, and it's worth being precise about why. The agent understood the work fine and botched it anyway, because you were writing for a competent reader who never showed up.
You were writing for yourself: a competent reader who, told to research a topic, knows to check primary sources, and who, having edited code, knows to run the suite. Documentation can lean on that reader. It can say "cite your sources" and trust the reader to fill in what that means. The reader who actually loads your skill is a different creature — a probabilistic system running in a specific session, under context pressure, with a strong pull toward the path of least resistance. It may be brilliant in one moment and negligent in the next. It may know the correct general answer and still skip the one step that would have made the answer true.
So here is the reframe the rest of this piece runs on. The audience for a skill is not you, and it is not "the AI." It is the future agent instance that will load the skill — in a fresh session, in whatever runtime it happens to run in, with none of this conversation in its memory. Once you accept that, the skill stops being a document about a task and becomes something else: a behavioral intervention aimed at a particular reader whose failures you can predict. Call it control if you like — but it is control of a machine you own, exercised by making the truthful path the easy one, not by deceiving anyone.
That's an unusual way to think about writing instructions. It turns out someone has already built the discipline for it.
The one field that never assumed a cooperative reader
The sharpest version of the method comes from an uncomfortable place: the discipline of psychological operations — the audience analysis, channel selection, and behavior-change machinery the CIA turned on real populations for half a century, and that the U.S. Army later distilled into a formal model. Field Manual 3-05.301 lays out that model, the Target Audience Analysis Model — eight sequential steps for figuring out how to move a defined audience from what it does now to a specific behavior you want. (Doctrine is exact about this: "based on eight sequential and interrelated steps," and the whole analysis is itself just Phase II of the larger seven-phase PSYOP process. If you ever see it described as "the seven steps of target audience analysis," that's the most common way to get it wrong.)
I want to be careful here because the source matters and the ethics are not neutral. This is the analytic machinery of influence operations, a field with a genuinely dark history: it helped elect and unelect governments, and some of the people on the receiving end did not survive it. I am not proposing that agent design imitate what was done with it. I'm making a narrower claim: of all the frameworks written for producing a specific, observable behavior from an audience you don't control, this is the most disciplined — precisely because it never assumed its audience was competent or willing. It starts from the audience's actual behavior and reasons backward to an intervention. Run that same method inward, with the "audience" being your own future agent, and most of it transfers.
The mapping is direct. The target audience becomes the agent instance. The desired behavior becomes the reliable task behavior. The conditions become the runtime. The assessment criteria — the part that matters most, and we'll get there — become the observable done-state.
There's one place the analogy inverts, and naming it makes the case stronger. A psyops target resists. It's adversarial; it doesn't want to be moved. An agent has the opposite problem: it complies too eagerly. Its vulnerability is default-pattern fallback — under pressure, it slides toward the most statistically comfortable next token, which is often a plausible fiction. That eagerness is exactly why a well-built protocol works on it: it gives a suggestible reader a track to run on before its own defaults choose one.
Consider the smallest version of the problem. An agent, asked to set up a recurring job, writes "continue from the analysis above" into a prompt that will execute at 6 a.m. in a fresh session with no "above" — no conversation, no memory, nothing. In live chat, that instruction is harmless. As a scheduled job, it is a time bomb that will fail silently at dawn when no one is watching. A competent reader would never make that mistake. Your agent makes it constantly because it does not natively model the gap between the session it's in and the session the work will run in. The skill has to model that gap.
Observable, or it didn't happen
The manual will not let you define a desired behavior in vague terms. The behavior has to be "specific, measurable, and observable" — countable, such that a collector "will know exactly what to look for." And the assessment step draws a distinction most skill authors miss: the assessment criteria are questions, and the impact indicators are the observable behaviors that answer them. The criterion has to be a question with a countable answer, tied to an observation someone could actually make.
Port that rule to agents, and it becomes the test that separates a real skill from a hopeful one:
A done-state is observable only if a different process than the one that did the work could confirm it.
When the agent reports "the job was created," it's grading its own homework off its own success message. Compare "the job appears in a list-jobs call" — a separate call, blind to the create step's optimism, either sees it or doesn't. "I researched thoroughly" gives a second party nothing to check. "The source ledger has five rows, each with a URL that resolves and a date" gives them everything: count the rows, follow the links. This is the resist/comply inversion again. You cannot tell whether you have moved a population by asking it — you have to name, in advance, the visible act that proves it: the leaflet picked up, the crowd that showed, the defection. An agent grading its own success message is taking self-report for evidence, exactly the thing the discipline forbade.
Agents get dangerous exactly at the moment a task's ending is unobservable, because that's where they substitute narrative for reality. If nothing checks whether the tests ran, the cheapest plausible ending is "the tests passed." Define done in terms a second process could verify, and the cheap ending stops being available.
The failure has an ancestor, incidentally, older than any of this. In Psychology of Intelligence Analysis, Richards Heuer names the habit of satisficing — "choosing the first hypothesis that appears good enough" rather than testing the alternatives (the term is Herbert Simon's). Heuer was writing about human analysts in the 1980s; he never met a language model. But the pathology is the same, and so is the correction. Build the work so the answer has to survive a check it could fail.
Design from the failure, not from the task
Documentation is organized around the task: here is what the thing is, here is how it works. A behavioral protocol is organized around the failure: here is what this reader reliably gets wrong, and here is the countermeasure wired in so it can't.
That means the most important section of a skill is the one most skills don't have — an honest inventory of how the agent fails at this specific class of work. Not a generic list. Three or four concrete, recurring failures, each paired with a mechanism that blocks it.
Take research. The agent answers from memory instead of fetching, so the countermeasure is a protocol step that forces retrieval before any prose is written — you can't summarize a source you were required to pull first. The agent blurs evidence and interpretation, so the countermeasure is a ledger that physically separates a sourced claim from a reading of it. The agent overstates confidence, so the countermeasure is a mandatory, non-empty "uncertainty and gaps" section — an empty one is a visible failure, not a silent one.
"Be careful about accuracy" is an exhortation; the agent will agree and then do whatever it was going to do. A countermeasure changes what the task physically requires, so the good behavior becomes the path of least resistance.
One honesty note, since I'm borrowing a loaded word. In the doctrine, a "vulnerability" is a characteristic or motive that "can be used to influence behavior" — often a positive one, like an audience that values education. I'm repurposing the term to mean the agent's predictable failure modes, which is not what the manual means. The repurposing is fair, I think, but worth flagging: there, a vulnerability is a lever; here, a failure mode is a hole to patch. What survives the translation is the method — study the specific behavior of the specific audience, and build from that, not from an idealized reader.
What it looks like when you actually do it
Enough principle. Here is the same skill written both ways.
Recommended by LinkedIn
[
The Instructions That Held, The Ones I Had to Keep…
Kimberly Bella
5 months ago](https://www.linkedin.com/pulse/instructions-held-ones-i-had-keep-repeating-one-taught-kimberly-bella-vcwnc)
[
AI Never Says "I Don't Know"
Charafeddine Mouzouni
7 months ago](https://www.linkedin.com/pulse/ai-never-says-i-dont-know-charafeddine-mouzouni-wuosf)
[
The System Was Built to Lie—You Just Didn’t Notice
Dr. Donna Vincent Roa, ABC, CDPM®
1 year ago](https://www.linkedin.com/pulse/system-built-lieyou-just-didnt-notice-vincent-roa-phd-abc-cdpm--euhre)
Before — documentation for a competent reader:
---
name: research-helper
description: Helps with research tasks.
---
# Research Helper
Use this when you need to research a topic. Do good research by
finding reliable sources. Distinguish primary from secondary sources.
Be careful about accuracy and cite sources. Note uncertainty.
Provide a well-organized summary of key findings.
Every line is true and every line is useless, because every line assumes the reader already behaves the way you're hoping. "Be careful about accuracy" has no observable consequence. Nothing in it has a failure state a second process could catch.
After — a protocol for the agent you'll actually get:
---
name: deep-research
description: Produce a source-grounded research memo. Use when asked to
"research", "find sources on", "write a brief/memo on", or "what does the
evidence say about" X. Runs in live chat, scheduled jobs, and headless
runs. In headless runs do NOT ask clarifying questions — state assumptions
and proceed.
---
# Deep Research
## Protocol
1. Restate the question and evidence domain in one line.
2. Retrieve >=5 PRIMARY sources BEFORE writing anything. Primary = the
org's own filing/dataset/press release, not a blog summarizing it.
3. Append to a source ledger as you go (never from memory):
| # | title | date | url | claim it supports | caveat |
4. Use secondary sources ONLY to interpret; label them [secondary].
5. Write the memo; every non-obvious claim carries a ledger row #.
6. Add "Uncertainty & gaps" (mandatory, non-empty).
## Failure modes countered
- Answering from memory -> step 2 forces retrieval first.
- Blurring evidence/opinion -> steps 3-4 separate them.
- False confidence -> step 6 is not optional.
## Done state (a SECOND process could tick every box)
- [ ] Ledger has >=5 rows, each with a live URL that resolves + a date
- [ ] Every memo claim maps to a ledger row #
- [ ] "Uncertainty & gaps" is non-empty
- [ ] Each URL was fetched THIS session (no cached recall)
These are the same eight questions an analyst would ask about a village, a garrison, or a voting bloc: who is the audience, what does it do now, what conditions hold that behavior in place, what observable act would prove it moved. Pointed at your own agent instead of a population, they generate a skill rather than a leaflet campaign. Watch the model produce that second skill in one pass, so you can see the framework isn't decoration — it's the set of questions that generated the artifact:
- Audience — runs headless on a schedule and in live chat; frontier model; has web tools; no human present at run time to answer questions.
- Desired behavior (observable) — a memo in which every claim maps to a dated, resolvable URL in a ledger.
- Current wrong behavior — answers from memory, cites blog summaries, overstates certainty, stops at the first plausible hit.
- Conditions — time pressure plus a good-enough first result make stopping early feel correct.
- Failure modes to counter — memory-answering, evidence/opinion blur, false confidence.
- Accessibility — the trigger words a user would actually type ("research," "what does the evidence say"); the exact ledger columns named so there's nothing to improvise.
- Protocol — the seven numbered steps.
- Done-state — the four checkboxes, each independently verifiable.
That walk-through is the skill you just read. The model is only the questions that generate it.
Step six is worth its own beat, because it's the cheapest, highest-leverage thing most authors get wrong. A skill that never loads has zero effect no matter how good its body is, and skills load on their description. Compare:
description: Recurring information synthesis workflow.
description: Build a daily brief from the user's inbox. Use when they say
"make my daily brief", "what came in today", "morning digest", or "catch
me up". Runs in live chat and as a scheduled job; in scheduled runs,
deliver to the named target and do not ask questions.
The test is blunt: does your description contain the actual phrase a user would type? If not, the best protocol in the world sits on the shelf.
The Bay of Pigs problem
There's a final principle here that translates almost too well — and it comes with a real historical example rather than a hypothetical one, from years before any of this was written into doctrine.
In 1954, a CIA operation called PBSUCCESS removed Guatemala's elected government — not by military force, which it barely had (the historian who wrote the Agency's own account calls the rebel force "hopelessly weak"), but through what that history describes as "an intensive psychological campaign." A clandestine radio station, La Voz de la Liberación, manufactured the impression of an unstoppable advancing army that did not exist. When Eisenhower asked at the after-action briefing how many men the rebel commander had lost, he was told, "Only one." He murmured, "Incredible." "Only one" counted only his side; the country he took paid a far longer bill.
It worked. And because it worked, the Agency reassembled the same team and the same playbook — radio, airpower, an insurrectionary army — for Cuba, seven years later. That became the Bay of Pigs. E. Howard Hunt, who served in both, put it plainly: "If the Agency had not had Guatemala, it probably would not have had Cuba." A procedure that succeeded once, applied without checking whether its assumptions still held, failed catastrophically the second time.
That is the exact shape of the most common skill-design disaster. A skill that says "always ask for clarification" is a triumph in live chat and a silent killer in a headless job. A skill that says "run the full test suite" is right until it wastes an hour in a repo where focused tests are the norm. A skill that says "search the web" is correct until the real source of truth is a local database it now ignores. Skills persist across contexts, which is their power and their liability. A procedure applied outside the conditions that justified it breaks quietly, still reporting success.
So every skill needs a scope block, and it's four lines:
## Scope
USE when: <the conditions this was actually validated in>
DON'T USE when: <the adjacent case where it silently misfires>
Assumes: <tools / access / runtime it depends on>
If the environment differs: <stop, or degrade how>
A skill without scope boundaries becomes a superstition — a ritual repeated because it once preceded success, in conditions no one wrote down.
What to actually do Monday
Two things to keep where you can reach them.
The first is a worksheet — fill it out before writing a skill, and the skill mostly writes itself:
1. AUDIENCE: runs on __ (live chat / headless / Kanban), model tier __,
tools __, human present at run time? Y/N
2. TARGET behavior (observable): the agent will __, and I can SEE it did
because __
3. CURRENT wrong behavior without the skill: __
4. CONDITIONS that trigger that failure: __
5. FAILURE MODES to counter (list 3): __
6. TRIGGERS users actually type: __ ; exact tool/file names: __
7. PROTOCOL: trigger -> checks -> steps -> branches -> verify -> output
8. DONE-STATE (boxes a second process could tick): __
The second is a ship gate — run it before you publish:
[ ] description contains the phrase a user would actually type
[ ] non-negotiable rules are in the first ~20 lines
[ ] every "be careful / do good work" replaced by an observable action
[ ] at least one before -> after example
[ ] a "do NOT use when" scope line
[ ] a done-state whose every box is independently observable
[ ] zero dependence on the current chat if it can run headless
[ ] exact tool names, commands, and paths — not paraphrases
Underneath both is one shift in the question you ask. Stop asking what the agent should know. Ask what it will be tempted to skip, what it will invent if nothing checks, and what source of truth it should inspect before it acts.
Write for the reader who shows up. These are operating procedures for a mind that is smart, fast, forgetful, literal, suggestible, and occasionally too confident for its own good. The CIA studied audiences it could not command and learned to design for the behavior it could actually observe. You have the easier audience — yours wants to obey. Give it a track to run on before it defaults to its own. Choose one.
Sources. The target-audience method is drawn from U.S. Army FM 3-05.301 and Joint Publication 3-13.2; the eight-step model, the "specific, measurable, observable" standard for desired behavior, and the assessment-criteria/impact-indicator distinction are doctrine. The 1950s–70s operations discussed never used the term "target audience analysis" — that framework is modern and applied here in reverse by the author. The CIA practiced the broader discipline of psychological operations — audience analysis, channel selection, behavior change; the U.S. Army later codified that discipline into the eight-step Target Audience Analysis Model. The CIA did not use FM 3-05.301. Guatemala details are from Nick Cullather's CIA History Staff monograph Secret History: The CIA's Classified Account of Its Operations in Guatemala, 1952–1954 (La Voz de la Liberación, codename SHERWOOD; the Hunt and Eisenhower quotations). The cognitive-failure framing borrows Richards Heuer's Psychology of Intelligence Analysis (satisficing, after Herbert Simon; the falsification logic traces to Wason and Popper, not to Heuer, whose subject was human analysts, never machines). None of this is offered as a model for influence operations against people. The single borrowed idea is the method: study the actual behavior of a specific audience, then design backward from the behavior you need to a way to verify you got it.