Building Minds
The Cognitive Stack: All Cortex, No Hippocampus
Dr. Jerry A. Smith · April 23, 2026 · 10 min read

Building Minds | Edition 21
Four rungs of memory architecture — and the query classes the wrong rung cannot serve at any price.
By Jerry Smith | April 2026
An executive. An assistant. A Monday morning.
The executive asks the assistant — an AI deployed across the company for seven months — to summarize the last conversation they had about the Chen account. The assistant responds cheerfully and fluently. It also responds incorrectly. It names two people who were not in the meeting. It cites a decision that was reversed three weeks later. It quotes a number from a document the company stopped using in October.
What happened since October? the executive asks.
The assistant does not know what October is.
This is not a hallucination problem. It is not a model problem. No amount of parameter scaling fixes this, and RLHF does not help. The assistant failed because it has the cognitive architecture of a brilliant amnesiac — all cortex, no hippocampus. Trained on the sum of human knowledge, rented by the minute, unable to remember what happened on Tuesday.
The enterprise AI conversation in 2026 is still organized around model choice. The argument that will define 2027 is about context architecture — what kind of memory the system has and how that memory is organized.
Four rungs are visible now. Each rung unlocks a query class that the rung below cannot serve at any price. That is a stronger claim than "exponential capability curve," and an easier one to prove.
Rung 1 — The Monolithic Context
The default. A chunk of content in, intelligence out. The ChatGPT box you are still using. The Claude chat with PDFs attached. The custom GPT with a system prompt and a persona.
Rung 1 works for bounded tasks. Summarize this document. Answer this question. Draft this email. That is the task class it was designed for, and in that class, it is remarkable.
The ceiling is not secret. RULER — a COLM 2024 benchmark of 17 long-context language models — found that only about half of the models advertising a 32K context window can actually sustain performance at 32K (arXiv:2404.06654). Even among those that do, Lost-in-the-Middle documents a roughly 20-point drop in accuracy when the relevant information sits in the middle rather than at the ends of the window (arXiv:2307.03172). Long is not the same thing as readable.
Worse, context rots. Production practitioners measure up to 40% performance degradation from bloated windows, and well-instrumented production agents log a $0.52 overhead per message once their fixed overhead passes 45,000 tokens. The agent is burning money before the user types.
Rung 1 has no memory. It has an input. Every session begins at zero. Whatever understanding is assembled during inference is gone before the next message arrives.
Rung 2 — Vector + Graph Retrieval
The current production standard. Vector retrieval (Pinecone, Weaviate, LightRAG). Knowledge graphs (Neo4j, GraphRAG, LightRAG-graph-mode). External stores the model queries before it answers.
This is where most of enterprise AI actually lives in 2026.
The receipts are real. Microsoft's GraphRAG — evaluated on 1M-token sensemaking corpora — wins 72 to 83% of head-to-head comprehensiveness judgments against vector RAG (arXiv:2404.16130). HippoRAG, a hippocampal-indexing retrieval method, lifts multi-hop QA recall by 20.9 points while running 10 to 30× cheaper and 6 to 13× faster than iterative chain-of-thought retrieval (arXiv:2405.14831). On legal-domain queries — where relational structure matters most — LightRAG wins 84.8% of head-to-head evaluations against vector-only baselines.
But Rung 2 has a subtler ceiling. Everything it retrieves lands back in the same undifferentiated context window. The agent does not know which chunks are facts about the world and which are facts about you. It cannot tell when something happened. It cannot tell whether a sentence is a decision, an observation, or a hypothesis.
The cleaner way to say it, borrowed from Edition 11: a vector database performs approximate nearest-neighbor search. It does not spread. Rung 2 can find the needle. It cannot follow the thread.
Rung 3 — Episodic + Semantic
In 1972, the psychologist Endel Tulving drew a distinction cognitive science has never stopped using. Episodic memory is what I know about me — events I experienced, in sequence, with context. Semantic memory is what I know about the world — facts and concepts stripped of the moment I acquired them.
Rung 3 puts that distinction into the architecture. Memory is no longer a flat pile of chunks; it is typed by cognitive role. Retrieval asks which store? before what query? Provenance is preserved — episodic memories carry timestamps, semantic memories are consolidated. When the stores disagree, the system has a principled way to resolve it.
The headline number comes from Zep, a temporal-knowledge-graph memory system published January 2025. On LongMemEval — a benchmark purpose-built to test cross-session episodic reasoning — Zep delivered +18.5% accuracy while cutting latency 90% versus baseline (arXiv:2501.13956). That is the rare "up and left" on a cost-quality plot.
The benchmark matters because it is what enterprise chat-memory products actually fail at: commercial systems instantiating ChatGPT or Coze on GPT-4o drop 30 to 64% on the same task versus the oracle condition (arXiv:2410.10813).
Rung 3 is also where the open-source community is visibly converging. LangChain's tripartite memory model (semantic / episodic / procedural) and Zep's episodic/semantic split independently re-derived in 2025–26 what Tulving taught cognitive psychology in 1972.
Rung 4 — Cognitive Processing Blocks
Rung 3 has two buckets. The brain has more than two.
Rung 4 decomposes memory along the full functional architecture of cognition: working memory (the scratchpad of the moment), episodic (the me-story), semantic (the world-facts), procedural (how-to patterns), perceptual (pre-parsed sensory), affective (emotional valence), metacognitive (knowing what I know), and attentional (the gate on all of it). Each typed store has its own write rules, read rules, and forget rules.
This is the rung where memory stops being a database and starts being an architecture.
The receipts are thinner at Rung 4 — and honestly, that is itself the finding. The frontier is real; the controlled benchmarks are still being written. What does exist: Voyager unlocks Minecraft tech-tree milestones 15.3× faster than prior state-of-the-art by adding a persistent procedural skill library — a store for how-to, cognitively distinct from episodic and semantic (arXiv:2305.16291). Generative Agents show via ablation that the interface between episodic memory and semantic generalization — what they call reflection — contributes most to believable agent behavior (arXiv:2304.03442).
Recommended by LinkedIn
[
Old-School Computer Vision vs AI Agents: Why Hybrid…
Alex Gurbych, PhD
7 months ago](https://www.linkedin.com/pulse/old-school-computer-vision-vs-ai-agents-why-hybrid-alex-gurbych-phd-w2wyf)
[
Smart AI, Heavy Baggage: How to Stop Context Creep…
Cory Nation, PE, MBA, MSME
1 week ago](https://www.linkedin.com/pulse/smart-ai-heavy-baggage-how-stop-context-creep-before-cory-498cc)
[
McLuhan, AI, and the Architecture of Agency
Paul Duplantis
4 months ago](https://www.linkedin.com/pulse/mcluhan-ai-architecture-agency-paul-duplantis-bp7te)
And in-house: the Neuro-Cognitive Agent — twenty-nine of thirty-six brain mechanisms implemented, each mapped 1:1 to a specific brain region — beats Claude Opus 4.6 9 wins, 1 tie, 1 loss on eleven relational-context questions. The full mechanism map is browsable at nca-crossref.vercel.app; you can hover any region and see which computational primitive it drives.
At Rung 4, the system does not retrieve memories. It reconstructs experiences.
Why the rungs don't just add — they unlock
Three compounding effects explain what is going on.
Typed retrieval compounds. Each rung adds a dimension to the query. Rung 1 has one bucket. Rung 2 retrieves from many buckets but treats them as one kind of thing. Rung 3 asks which store? before what query? Rung 4 asks which store and which cognitive function? — and can spread activation along memory-type edges the way CA3 recurrent collaterals do in a real hippocampus. Every dimension shrinks the query space multiplicatively. HippoRAG's 10–30× cost reduction at equal or better accuracy is what this shrinkage looks like in practice.
Stores specialize. An episodic store has different invariants than a semantic store: episodic memories are immutable and timestamped; semantic memories are consolidated and deduplicated. A procedural store indexes on patterns; a working-memory store decays by minute. Specialization pays on both axes at once — it speeds retrieval and improves recall — which is exactly the shape of Zep's +18.5% / −90% result.
Cognition emerges at interfaces between memory types. Flat context has no interfaces — there is only one store. Rung 2 has one (retrieval). Rung 3 has three. Rung 4 has dozens, and this is where the valuable behaviors live: planning at the interface of working and procedural memory, reflection at the interface of episodic and semantic, identity at the interface of episodic and metacognitive. You cannot buy these behaviors by making the context window bigger. You can only buy them by giving the architecture somewhere for the interfaces to exist.
That is the step-function claim. Each rung adds capability classes the rung below cannot access at any token price.
The honest ceiling
This is where most memory pieces go soft. Here is the other side.
Typed memory is not free. MemoryAgentBench — the cleanest head-to-head across memory architectures — reports MemGPT scoring 30.6% on Accurate Retrieval while vanilla Embedding RAG scores 65.0%. When the routing is miscalibrated, coordination overhead eats the gains. This is the memory-axis analog of Berkeley's MAST finding for multi-agent systems: the failure modes are architectural, not parametric.
Long context sometimes wins. Google DeepMind's Self-Route paper (arXiv:2407.16833) shows that on many factual-QA benchmarks, given sufficient resources, flat long-context outperforms RAG on average accuracy. Databricks's internal evaluations find models saturate and then degrade past a context size (GPT-4 after 64K, Llama-3.1-405B after 32K). The correct framing is not "higher rung always wins." It is the right rung for the query class. Cross-session temporal reasoning, sensemaking over a corpus, institutional continuity — the gap is enormous. Synthetic needle-in-a-haystack at small scales — Rung 1 is fine.
And Rung 4 has no controlled benchmark beating Rung 3 yet. Voyager and Generative Agents are existence proofs, not head-to-head victories. No published LLM+ACT-R or LLM+SOAR hybrid has beaten Zep on LongMemEval. The frontier is real; the receipts are still being written.
The thesis survives. The hedge is simply that each rung has a query-class envelope, not a universal superiority.
What a CTO is actually paying for
If you are buying AI, you are buying a memory architecture. Three questions to bring to your next vendor meeting.
What happens to this conversation in six weeks? If the answer involves retraining, re-uploading, or re-prompting, you are on Rung 1 or Rung 2 — a typing tool that forgets every Monday. Rung 3 and higher retain episodic continuity across sessions. Most enterprise AI does not.
When the system contradicts itself, how does it decide which version to trust? Rung 1 has no answer — no history. Rung 2 cannot tell; every retrieved chunk carries the same epistemic weight. Rung 3 can tell, because episodic memories are timestamped and semantic memories are consolidated. Rung 4 can also tell why — because the metacognitive layer tracks the system's own confidence.
Where is the institutional learning? If your firm cannot describe what the AI has learned about your operations, customers, and domain over the last six months, the answer is nowhere. It has been typing, not learning. Rung 3 stores experience in typed memory. Rung 4 consolidates it into procedural skill libraries — the same mechanism Voyager's agent uses to get 15.3× faster. This is the difference between a cheap assistant and an institutional asset.
The cost story has also changed. Rung 1 token costs are converging toward zero. Rung 2 infrastructure is a solved engineering problem. Rung 3 is where the cost concentrates — and where the capability you cannot buy at Rung 1 at any price actually lives.
All cortex, no hippocampus
That is the architectural diagnosis for most AI deployed in 2026 — trained on the sum of human knowledge, rented by the minute, unable to remember what happened on Tuesday.
The model is rented. The memory is owned. You cannot buy the second by paying more for the first.
Four rungs are visible now. Each one unlocks a query class that the rung below cannot reach. The firms that will compound an AI advantage over the next two years will not be the ones with the biggest context windows. They will be the ones who built the memory — episodic, semantic, procedural, metacognitive — that turns model inference into institutional intelligence.
This is the quieter work. It does not show up in the demo. It shows up eighteen months later, when the firm on Rung 3 can describe its customers, its contracts, and its own operations in a way no one on Rung 1 can. The firm on Rung 4 can reason about why.
Ask the vendor what happens to the conversation in six weeks.
Ask yourself the same question.
If your firm is working through which rung fits your problem — or what it takes to build a typed memory architecture that actually retains what your business learns — reach out. I'm Dr. Jerry Smith at Verity Vantage Group, and that conversation is what we do.
Building Minds is a newsletter about cognitive architecture for the age of agents. The premise across every edition is the same: durable AI advantage is not being built in the models — it is being built in the structures around them. The memory systems, the coordination protocols, the governance layers, the temporal intelligence, the topological anatomy. Each edition investigates one axis of that architecture with a single question in view: what does it actually take to turn a model into an organization?