All posts

The Economy of Attention

Dr. Jerry A. Smith · April 14, 2026 · 6 min read

Building Minds — Edition 20

Attention is a finite thing. This is easy to miss, because we use the word in so many senses that its technical meaning has become obscured. Inside a transformer — the mathematical object that every modern language model is built from — attention is something quite specific: it is the operation by which the model weighs each token in its input against every other token, in order to decide what each word means in the presence of the others. Attention is the whole of what the model does. And it is, in the literal sense, a budget.

The budget is not metaphorical. The cost of attention scales with the square of the input's length. If you double the number of tokens, you quadruple the cost. Every token you give the machine competes, for a finite sum of consideration, with every other token it has also been given. Irrelevant tokens are not free. They draw down the machine's attention in proportion to their presence, and the relevant tokens receive whatever is left.

Two machines sit on my desk, clustered together through a piece of software called EXO. They are M3 Ultras, and between them they run a thirty-five-billion-parameter model with a seventy-billion-parameter model held in reserve. When I ask them to reason about my code, they almost never see more than eight thousand tokens of context on a single turn. People who run systems like this tend to describe the limitation with something close to shame. The cloud models, they note, have windows of a million tokens or more.

I would like to suggest that the shame is misplaced. The small window is not a failure. It is a more honest instrument.


Consider what it would mean for attention to be truly abundant. Suppose you could give a model all four hundred thousand lines of your codebase at once, and pay nothing for the privilege. Would it read the way you hope it would? The attention budget is not suspended because the window has grown. The budget has simply become easier to exhaust without noticing. In 2023, Liu and his colleagues demonstrated something that ought to have been the end of the argument: when the relevant signal is placed in the middle of a long context, the model's ability to locate it collapses. The window will technically accept the input. The machine will nonetheless fail to find the idea inside it.

This result did not stop the sale of larger windows. It only made them harder to evaluate.

Here is what happens when you paste a whole file into a local model that holds eight thousand tokens. The model reads every word, weighs every word against every other word, and if the relevant function occupies the last forty lines of an eight-hundred-line file, the surrounding noise draws the machine's attention away from where the answer lives. The prompt becomes, in a precise sense, a meditation exercise the model is failing. And because the machine sits on my desk and costs me only electricity, I notice the failure immediately. The answers come back confident and wrong. The cluster does not charge me for this. It simply returns a worse result and keeps doing so until I understand what I have been asking of it.

A cloud model with a million-token window will absorb the same mistake and often produce a decent answer anyway. It will do this at a cost that appears only later, on a monthly statement, and by the time the cost is visible, the lesson is not.


There is a better unit of work than the file, and the machine will teach you what it is if you listen.

Recommended by LinkedIn

[Newsletter Volume 8, AI, Mathematics, and Financial LiteracyNewsletter Volume 8, AI, Mathematics, and Financial… Prof Roxanna Ezenekwe

1 month ago](https://www.linkedin.com/pulse/newsletter-volume-8-ai-mathematics-financial-literacy-ezenekwe--m8yue) [Explaining Funding Round Momentum With Complexity TheoryExplaining Funding Round Momentum With Complexity… Daniel W. Dippold

2 years ago](https://www.linkedin.com/pulse/explaining-funding-round-momentum-complexity-theory-daniel-w-dippold-tyu0f) [Peter Senge: The Fifth Discipline at Thirty-Five — Lineage, Surge, and ScalePeter Senge: The Fifth Discipline at Thirty-Five —… Sheila Damodaran

7 months ago](https://www.linkedin.com/pulse/peter-senge-fifth-discipline-thirty-five-lineage-surge-damodaran-35t4f)

A codebase of four hundred thousand lines contains, by most measures, something like thirty thousand addressable symbols — functions, methods, classes, the named things the code actually refers to. For any given change I might ask the machine to make, the number of these symbols that matter is small. Often four. Sometimes two. Almost never a dozen.

The file is a convenience of the filesystem. It is the shape in which a human editor opens code on a screen. The machine has no screen. What the machine has is the ability to consider relationships between named things. If the question is about the method UserSession.validate, the meaningful neighbors of that method are its caller, the class it belongs to, and the tests that pin down its contract. They are not the other fourteen methods in the same file. The file was an accident of organization. The neighbors are the thing.

What follows is not mysterious. You build a map of the code as a graph — tree-sitter will produce one, a language server will produce one, the tools for this are older than the models. You embed each symbol, along with its signature and its documentation, into a vector space. When the machine is asked to consider a change, you retrieve the symbols by their semantic closeness to the task, expand a step or two along the structural edges of the graph, and stop. You hand the machine a tight, coherent slice. Four symbols. Perhaps six. The machine sees the idea and not the file.

Eight thousand tokens is enough for this. Thirty-two thousand is ample. A million-token window, from here, looks less like a capability than a way of selling patience as comprehension.


I want to be careful about the claim I am making. This approach does not solve every problem. A refactor that touches ten call sites will struggle when only three of them are semantically close to the query; the graph has to be walked more patiently, and the walking is still an open problem. Code that does not yet exist cannot be retrieved; greenfield generation remains difficult for different reasons. The hardest case is the refactor whose correct answer is deletion — the removal of a layer of abstraction that should never have been there — because the graph has no edge labeled this should not exist.

These are real limits. They are not, however, limits of the window. They are limits of how the system decides what belongs inside the window on any given turn. The frontier is not capacity. The frontier is selection.


What the cluster has taught me, over months of use, is something I could have said before I owned it but would not have believed in the way I do now.

The machines that pay their own electricity bills are the ones that tell you the truth about what you are spending. The cloud is not lying; it is only making the economy invisible, and invisibility is a strange kind of teacher. My local cluster cannot afford to be sloppy, and because it cannot, I cannot. I have learned to ask it smaller questions, and to prepare the ground more carefully before asking. I have learned that the correct question is almost never how much can this machine hold. The correct question is, what is the smallest slice of the problem that would permit an answer? This is a different kind of discipline than the one cloud usage produces. It is closer, I think, to the discipline of thinking.

The window was never the measure of what these systems understand. The window was the measure of how patiently they could be asked to hunt for meaning inside noise. Smaller windows reveal the hunt. Larger ones let us pretend the hunt is not happening.

The cluster on my desk has been telling me this for some time. I suspect the lesson was available before the cluster arrived, and that the machines actually made me quiet enough to hear it.

— Jerry

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call