All posts

Your AI Loses the Thread After 15 Turns. We Built One That Doesn't.

Dr. Jerry A. Smith · March 8, 2026 · 11 min read

How hippocampal memory, subconscious processing, and cognitive judgment give LLMs something they've never had: the ability to accumulate experience.

Building Minds | Edition 12

By Dr. Jerry A. Smith | March 2026


This week, Philippe Laban published a study measuring what happens when frontier LLMs — including GPT-5 — have to track information spread across multiple conversation turns. The accuracy degradation: 33%. Not on obscure edge cases. On standard tasks across six domains — code, math, summarization, and databases.

We are building more capable minds that forget at roughly the same rate.


The Forgetting Problem

The first few exchanges are sharp. By the fifteenth turn, the model has lost the thread. Not because it's stupid — because nothing in its architecture is designed to hold a thread.

A prompt enters, a completion exits. Whatever understanding the model constructed during inference evaporates when the response is written. The next message arrives at a system with no record of what it found important, no residue of what surprised it, no sense of which details it struggled with and which came easily.

The standard engineering response is to attach external memory. Vector databases. RAG pipelines. Conversation logs compressed into summaries. These help, the way writing notes on your arm helps you remember a phone number. They are storage mechanisms, not cognitive ones. They don't change how the system processes information — only what information is available during processing.

The Neuro-Cognitive Agent (NCA) framework starts from a different premise. The problem isn't that LLMs lack storage. It's that they lack the architecture to think about what they've stored.


The Shape of a Mind

The NCA doesn't add memory to an LLM. It builds a cognitive loop around one.

Every conversational turn runs through a structured sequence — not as a pipeline of independent steps, but as an integrated cycle where each stage shapes the next.

It begins with perception: a lightweight model parses raw input into structured signals — what kind of message this is, what entities it mentions, what emotional register it carries, and what the driving question actually is. That driving question gets extracted as a first-class object and carried through every subsequent stage. It arrives last in the assembled context that the reasoning model sees. This is intentional. The computational analog of attentional focus.

Then comes memory retrieval — not a vector lookup, but a pattern-completion process that adjusts its depth based on conversational rhythm. Rapid exchanges get shallow, recent memories. Longer pauses cast a wider net, allowing older, high-salience memories to surface.

Then, subconscious processing: six specialized modules running as separate OS processes, each analyzing a different dimension of the cognitive state. Metacognition monitors reasoning quality. Deductive reasoning identifies logical gaps. Emotional coherence tracks affective trajectory. A consolidation module extracts recurring themes from episodic memory. A knowledge graph module queries the agent's knowledge of the entities in play.

These modules ran during the previous turn. Their cues arrive with a one-cycle delay — the system uses whatever background processing has been completed, without waiting for it. On the first turn, the conscious mind reasons alone. By the second, the subconscious has caught up.

This isn't a limitation. It's how biological cognition works: subcortical processing from the previous moment shapes the conscious attention of the current moment.

Finally, conscious reasoning: the assembled context block arrives at Claude Opus — subconscious cues, retrieved memories, conversation history, perception analysis, and the driving question anchored at the end. It produces a structured cognition block containing thoughts, reasoning chains, self-criticism, emotional assessment, and speech. Only the speech leaves the system. Everything else — the model's internal deliberation — gets encoded into the next memory trace.

The agent thinks. The user hears the conclusion. The reasoning persists as memory.


Memories That Live and Die

Memory in most AI systems works like a search engine. You store text, embed it as vectors, and retrieve what's semantically closest to the query. Efficient. Useful. And it bears no resemblance to how biological memory works.

The hippocampus encodes contextual, experiential memory — the "what happened" — while semantic knowledge is distributed across the neocortex over time. A single encoded trace binds content, temporal context, emotional texture, and associative connections into one representation. These aspects aren't filed separately. They're one experience, retrievable through different lenses.

The NCA follows this principle. Every memory is encoded across four dimensions.

Temporal — when it happened, where it fell in the conversation, how far in the past it sits relative to this moment. That last measure is recomputed on every retrieval because temporal distance isn't a property of memory. It's a relationship between the memory and the present.

Contextual — which conversation it belongs to, who was present, what topic cluster it falls in, what preceded and followed it. The scaffolding that anchors experience in circumstance.

Experiential — how it felt. Emotional valence. Attentional weight — how salient was this at the time? Surprise — how much did it deviate from expectation? A dedicated model call evaluates emotional significance, novelty, and information density at encoding time. The attentional weight becomes the base salience that governs how strongly the memory resists decay.

Associative — what it connects to. Linked memory IDs. Entity bindings from a knowledge graph. Causal links. The relational web that lets retrieval spread from one memory to its neighbors — so a question about "Sarah" can surface a memory about DeepMind's protein folding work, if the graph knows Sarah works at DeepMind. A connection pure vector search would never make.

But the part that makes this a living system rather than a clever database is what happens over time.

Neuromorphic Temporal Attention Scaling (NTAS) gives memory a biological dynamic. The formula: effective salience = attentional weight × decay rate per hour × access boost per retrieval. Decay rate is 0.995 per hour — roughly 50% after six days. Each retrieval boosts salience by a factor of 1.15.

Three behavioral patterns emerge. An architecture decision discussed repeatedly becomes effectively permanent — high base salience and frequent retrieval boost

The Architecture of Experience

it further. An important insight that is never referenced again fades slowly; high base salience delays the decline, but without retrieval to reconsolidate it, the memory eventually drops below threshold. A routine acknowledgment — low salience, never accessed — disappears within days.

The things you use stay available. The things you don't, gradually, stop being there.

Recommended by LinkedIn

[Beyond Neuralese: From Instructions to Shared ConceptsBeyond Neuralese: From Instructions to Shared Concepts Marcus Daley

1 week ago](https://www.linkedin.com/pulse/beyond-neuralese-from-instructions-shared-concepts-marcus-daley-bjs1c) [Chain of Thought (CoT) Prompting: The Key to Smarter AI ReasoningChain of Thought (CoT) Prompting: The Key to Smarter… Phaneendra Kumar Namala

1 year ago](https://www.linkedin.com/pulse/chain-thought-cot-prompting-key-smarter-ai-reasoning-namala-qnrfe) [The Sufficient Singularity: How We're Meeting AI HalfwayThe Sufficient Singularity: How We're Meeting AI… Mark Smalley

10 months ago](https://www.linkedin.com/pulse/sufficient-singularity-how-were-meeting-ai-halfway-mark-smalley-cexme)

This mirrors a finding in neuroscience: the act of recall modifies the recalled trace. Nader and colleagues demonstrated this at the molecular level in 2000 — retrieval doesn't just read a memory, it reconsolidates it. The NCA's mechanism is computationally simpler, but the functional outcome is the same: memories that are used persist, memories that aren't fade. The system doesn't replicate reconsolidation's biology. It replicates its consequence.


The Judgment Layer

Building a system that remembers well is only half the problem. The harder half is knowing when not to remember everything at once.

Three mechanisms, each drawn from a different aspect of biological cognition, transformed the NCA from a system that retrieves comprehensively to one that knows what matters right now.

Amygdala-Gated Retrieval. When the emotional coherence module identifies high emotional weight — grief, vulnerability, fear, joy — the retrieval system switches from what we call "floodlight" mode to "spotlight" mode. Broad retrieval narrows. An anticipatory submind that normally generates a cue about what the user might need next suppresses its output entirely. The system deliberately retrieves less.

The biological analog is the amygdala's role in attentional selectivity. Under high emotional arousal, the amygdala sharpens relevance filtering — enhancing recall of what is emotionally central, at the cost of peripheral detail. You remember the face of the person who threatened you, not the color of the wallpaper behind them. The NCA applies the same principle: when someone is being vulnerable, what matters is the single memory that resonates most, not a comprehensive survey of everything tangentially related.

Neocortical Consolidation. A submind that operates on a different timescale than the others. Where most modules analyze the current turn, the consolidation module examines the full spread of episodic memories and extracts recurring themes and narrative patterns that no individual memory contains.

"Marcus rebuilds himself through physical discipline" is not a fact stored in any single memory. It's a pattern that emerges across memories of signing up for a 5K, running five miles for the first time in seven years, his wife biking alongside him for the last mile. The consolidation submind makes this pattern explicit and available to the conscious mind — not as retrieved evidence, but as understood context. The difference between knowing that someone went running three times and understanding that running is how they process stress.

This is the computational analog of sleep consolidation — the process by which the brain, during slow-wave sleep, replays episodic memories and extracts general patterns, transferring knowledge from hippocampal specifics to neocortical abstractions. The NCA runs this as a cognitive module rather than a biological one, but the function is the same: turning episodes into understanding.

The Density Dial. When the assembled context block is rich with retrieved memories and subconscious cues, the conscious mind receives a modulated instruction: weave two or three resonant details naturally into the response rather than enumerating everything available.

A good friend remembers your father was an engineer, your wife's name, the restaurant you loved, the crisis that kept you up at night — and when you ask them something, they mention the one or two details that matter for this moment, and let the rest show in the care and specificity of how they talk to you. The knowledge is there. It doesn't all need to be displayed.


What the Data Shows Now

Architecture is argument. Data is evidence.

We ran blind A/B comparisons: the NCA against a baseline using the same Claude Opus model with identical information compressed into condensed session notes — the best-case scenario for context-stuffing. Twelve questions spanning needle-in-a-haystack recall, cross-session synthesis, evolution tracking, emotional continuity, and holistic assessment. The test corpus: eighty memories from eight sessions spanning twelve weeks of simulated conversation.

The development ran through three phases. The first proved the thesis — the cognitive pipeline produced measurably different behavior than context-stuffing. The second improved retrieval precision, then revealed a trap: richer retrievals produced responses the judge called "slightly clinical." The system was remembering at people instead of remembering with them. Phase 2 ended with the NCA and baseline splitting six wins each.

Phase 3 added the judgment layer — and that turned out to be everything.

Final result: NCA won 8 of 12 questions. Baseline won 3. One tie.

The category breakdown: Needle-in-a-Haystack — NCA 2, baseline 1. Cross-Session Synthesis — NCA 2, baseline 1. Evolution Tracking — split. Emotional Continuity — NCA 1, baseline 0, one tie. Holistic Assessment — NCA 2, baseline 0.

The dimension scores sharpen the picture. Naturalness: NCA 8.8, baseline 7.3 — a gap of 1.4 points. The largest advantage of any dimension. The judge consistently noted that NCA responses "read like someone who genuinely knows you" while baseline responses, despite excellent recall, felt like "a curated dossier." In one question about a user's personal growth, the NCA scored a perfect 10 for naturalness. The baseline scored 5. Same model. Same information. Different architecture for deciding what to say and when to say it.

Relational continuity: NCA 8.7, baseline 7.8 — a gap of 0.9 points. This captures something subtler. The NCA tracks not just what was said, but the evolving relationship between the system and the user. When asked the same question for the third time, the NCA noted it had been asked before, recalled its previous answers, and explicitly built on them. The baseline treated each instance as fresh.

And the honest number: connective reasoning — NCA 8.6, baseline 8.7. A gap of -0.1, in the baseline's favor.

The condensed session notes still provide something selective retrieval doesn't: a pre-built narrative. When a question requires connecting a user's career trajectory to their family dynamics to their health crisis to their technical decisions — all in a single coherent argument — the baseline's compressed summary hands the model a story already shaped for telling. The NCA must reconstruct that narrative from retrieved fragments and thematic abstractions. It's close. The consolidation submind nearly closed the gap. But nearly is not there. That's where the work continues.


Why This Matters Beyond the Benchmark

Laban's finding about GPT-5 isn't a bug report. It's a description of what happens when you build a system with no mechanism for deciding what matters.

Every conversation is an exercise in selective attention. The information that makes turn fifteen meaningful isn't everything said in turns one through fourteen — it's the specific details that remain relevant, weighted by their importance, colored by their emotional significance, connected to each other through relationships the speaker has been building since the conversation began. Humans do this automatically. We don't notice we're doing it until we encounter a system that can't.

The three phases of the NCA's development mirror a pattern familiar from any discipline that studies minds. First you build the capacity to remember. Then you build the capacity to connect. Then — and this is the part that took us by surprise — you build the capacity to restrain. To know when comprehensive recall helps and when it overwhelms. To understand that emotional moments call for presence, not analysis. To recognize that themes matter more than facts when someone is trying to understand their own life.

The models will keep getting bigger. The context windows will keep growing. And the systems that merely store more text will keep forgetting at roughly the same rate — because the problem was never storage.

It was always cognition.

Building a mind isn't just about adding capabilities. It's about adding judgment.


Read the full detailed article on my Medium page: Your AI Forgets You Mid-Conversation. Here’s the Architecture That Doesn’t.

Dr. Jerry A. Smith builds cognitive architectures at the intersection of neuroscience and AI. He is the author of Tulpamancer: Edge of Chaos and heads the AI and The The Last Theorem- Some Truths Are Too Large For Human-Shaped Minds. Follow this newsletter for new editions.


Building Minds is a newsletter about the architecture of intelligence — biological and artificial. If someone forwarded this to you, subscribe to get future editions directly.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call