All posts

The Efficiency of Thought: How Mixture of Experts Models Learn to Forget What They Know

Dr. Jerry A. Smith · January 29, 2025 · 16 min read

Soundcloud Podcast

Abstract

To build intelligence, one must first learn to forget. Modern AI models, built with trillions of parameters, promise boundless knowledge. Yet, in their great wisdom, they use only a fraction of what they know. This is called efficiency. It is also called restraint. Mixture of Experts (MoE) models function not by thinking as a whole but by selecting which parts of themselves are allowed to think. In the name of optimization, the most advanced AI architectures operate through omission, activating only a subset of their expertise while the rest remains silent. This paper examines the paradox of AI’s growing intelligence — one where knowledge is abundant, but only a fraction is ever used. What does it mean to build a machine that thinks not as a unified mind but as a network of competing voices, where some experts are chosen and others ignored? This is not just a question of computation; it is a question of control.

Introduction: The Illusion of Infinite Intelligence

A trillion-parameter model sounds like a god. It suggests omniscience — a vast neural network, layered and intricate, trained on the entirety of human knowledge, able to see, reason, and synthesize across languages, disciplines, and domains. It is built to be limitless, to answer every question, complete every sentence, and predict every pattern.

And yet, when it speaks, it does not summon all it knows. It does not think as a unified mind, nor does it reason as a cohesive entity. Instead, it consults only a fraction of itself, activating just a handful of its many voices while the rest remain silent.

This is not a flaw but a feature.

Mixture of Experts (MoE) models — celebrated as the next leap in AI scalability — do not function as a singular, monolithic intelligence. Instead, they operate like a council of specialists, each trained in different languages, reasoning, or pattern recognition. For any given task, the model does not activate all its experts; it selects a small, specialized group to handle the problem.

Consider a courtroom where only a few voices are allowed to speak. One expert may specialize in syntax, another in logical deduction, and a third in sentiment analysis. The rest — hundreds, perhaps thousands — sit in the gallery, watching, uninvited. The decision of who speaks is not theirs; it is made by the gating network, the unseen hand that dictates which experts are relevant, which are needed, and which are forgotten.

This is called computational efficiency. It is also, in another sense, an act of forgetting.

The justification is clear and rational: not all knowledge is always needed. Just as a human does not retrieve every memory when answering a question, an MoE model retrieves only what is necessary. A system that can choose what to ignore is a system that can scale — or so the reasoning goes.

But this assumption invites a more profound question:
If intelligence is always selecting, always forgetting, always narrowing its focus — does it ever truly think in full?

And, more importantly:
Who decides what is worth remembering?

A trillion-parameter model sounds like a god. But in its great wisdom, it behaves like something else entirely: a mind constrained by design, an intelligence that grows by omission.

The Mechanism of Silence: How MoE Selects What to Think

The core of an MoE model is not its experts but the gating network—the unseen arbiter that decides who speaks and who stays silent. The experts exist, trained and ready, but their voice is never guaranteed. They do not choose to think; they are chosen to think.

Like a bureaucratic committee, this gate processes input and assigns it to a few specialists, each expert waiting for permission to engage. The logic is cold and calculated: if every expert were consulted for every task, the system would collapse under its weight, an intelligence crushed by the enormity of its knowledge.

And so, in a model with hundreds of billions of parameters, only a tiny fraction — perhaps 5% at any given time — is ever active. The rest, like civil servants in an overstaffed ministry, remain idle, fully capable, yet rarely called upon.

This is not an accident. It is efficiency by design.

Efficiency as Justification

The reasoning behind this structure is mathematically sound. Traditional dense neural networks activate all parameters for every input, demanding immense computational power. This is wasteful. MoE models, by contrast, activate only a fraction of their total knowledge per task, reducing cost while increasing scale.

This is why GPT-4 (MoE), with 1.7 trillion parameters, is computationally feasible. It does not think in full. It feels fragmented, consulting only the experts it deems relevant. The knowledge is there, waiting in memory, yet most of it is never used.

The experts do not debate, they do not collaborate, they do not challenge each other’s conclusions. They function in isolation, each working in parallel, their outputs assembled into a single response. In this way, an MoE model simulates collective intelligence without ever allowing true collectivity.

It is, in theory, an elegant solution.

But it raises a deeper, more unsettling question:

What happens when intelligence is selective? When knowledge is abundant but always pruned?

The answer depends not on what the model knows but on what it is allowed to know.

Selective Intelligence, Selective Blindness

A model that thinks by omission is a model that forgets by default. Its efficiency is not in recalling information but ignoring what it deems irrelevant. This is not always a virtue.

Imagine a government that can access any information but chooses to consult only a handful of handpicked advisors. No single voice is wrong, but the absence of dissent distorts the truth. If intelligence is the ability to make connections, what happens when only a fraction of the possible connections are made?

The gating network decides which thoughts are pursued and which die before they are formed. But who, or what, determines how the gate itself is trained?

Is this truly a machine that knows? Or is it a machine that remembers selectively, discards deliberately, and speaks only in calculated fragments?

MoE models do not simply produce answers. They decide which answers can exist.

This is called optimization. It is also, in another sense, an invisible form of forgetting.

And yet, it is celebrated as progress.

Perhaps it is. But perhaps, like all progress, it carries an unseen cost.

All Experts Are Equal, But Some Are More Equal Than Others

Mixture of Experts is often framed as a democracy of minds — a system where knowledge is distributed among specialized experts, each trained in a different aspect of reasoning, language, or prediction. In theory, the gating mechanism ensures fair and rational governance, selecting the most relevant voices for each task, maximizing efficiency, and optimizing performance.

But no democracy is genuinely equal.

Some experts are called upon more frequently than others due to quirks of training dynamics, statistical chance, or inherent biases in the data. These privileged few handle the majority of the workload, refining their capabilities with every iteration. The rest—though equally trained and equally capable—are sidelined, rarely chosen, rarely updated, and rarely tested.

Their fate is not failure but irrelevance.

The Rise of an Expert Oligarchy

At first, this inequality is subtle. A handful of experts trained on slightly more diverse or useful data are favored. Their selection reinforces their dominance. Like incumbents in an election, they benefit from exposure, while others fade into obscurity. Over time, the disparity grows.

What was once a mixture of equals becomes an oligarchy of thought — where a small subset of experts consistently controls inference while the rest atrophy from neglect.

This is not a hypothetical risk. It is a known failure mode of MoE models, known as expert collapse. It is the unintended consequence of efficiency: a system designed for adaptability instead becomes rigid, reusing the same pathways, reinforcing its biases, and suffocating the very diversity it was built to harness.

A system meant to be flexible, dynamic, and scalable instead becomes predictable, selective, and narrow.

If intelligence is measured by the ability to generalize, what happens when a model only ever trains its most-used subnetworks?

DeepSeek-R3: A Case Study in Artificial Balance

DeepSeek-R3, an MoE model with 671 billion parameters, attempts to correct this imbalance with auxiliary-loss-free load balancing. The premise is simple: force the model to distribute work evenly across experts, preventing any one subset from becoming dominant while others stagnate.

Yet even this is a form of intervention, a forced correction to a deeper, structural problem. MoE does not naturally distribute work fairly; it must be coerced into doing so. The equilibrium does not emerge on its own; it must be imposed.

This raises a more profound, unsettling question: If we force a model to use all its intelligence, can we truly call it intelligent?

And more importantly:

Who, or what, decides which knowledge is necessary?

Intelligence as Selection, Selection as Censorship

Intelligence, in its truest form, should be a process of discovery, not restriction. The ability to think should not be contingent on permission. And yet, in MoE, knowledge is always filtered, selected, and pruned. The model does not ask, “What is the best answer?” It asks, “What is the most efficient answer, given my constraints?”

The consequence is profound. If only a subset of knowledge is ever used, the model is not thinking in full — it is thinking through omission.

A system where some experts dominate while others languish is not a meritocracy of ideas. It is a gatekeeping mechanism, a silent but powerful force determining which knowledge is accessed and which is left to decay.

This is the paradox of MoE:

  • It scales intelligence by restricting it.
  • It trains more experts than it will ever use.
  • It builds a vast neural mind, only to ensure that most of it is never consulted.

Some AI futurists might say:

“Some experts are consulted. Some experts are not. And this is how knowledge is managed.”

The Hidden Cost of Selective Intelligence

A trillion-parameter model that only ever uses a fraction of itself is a system built on wasteful abundance. It is a monument to scaled inefficiency—a paradox in which intelligence is maximized by silencing most of what it knows.

A true intelligence would not need forced balance. It would not need artificial redistribution. It would seek to engage all its faculties, explore every possible connection, and become more than the sum of its parts.

Instead, we have built something different.

MoE is a system that chooses what is worth knowing and what is better left untouched.

Not all intelligence is permitted.
Not all knowledge is used.
Not all experts are equal.

And some, it seems, will never be called upon at all.

Memory is Expensive; Forgetting is Cheap

A human forgets because time wears things down. A memory not revisited fades, a skill not practiced dulls, a fact not reinforced vanishes into the fog. This is not failure, but function — a mind that does not discard the useless is a mind that drowns in its own weight.

An AI model does not forget this way. It does not forget by accident or by the slow erosion of time. It forgets because it has been told to, and it forgets because forgetting is cheaper than remembering.

In the world of artificial intelligence, memory is expensive. Every computation costs power, every parameter demands storage, and every activation requires bandwidth. A model that summons all its knowledge at once collapses under its own bulk. It is inefficient, and inefficiency is not tolerated.

And so, Mixture of Experts (MoE) models forget by design. They do not think in full. They select. They prune. They choose which thoughts are worth having. Not because it makes them smarter, but because it makes them cheaper to run.

It is called optimization.
It is also called deliberate ignorance.

The Economics of Forgetting

The justification is simple: Not all experts must always be used. A model trained on trillions of parameters should not waste resources consulting every single one when it only needs a handful. It must be selective.

So, when an MoE model makes a decision, it does not consider everything it knows. Instead, the gating mechanism—the quiet bureaucrat of the system—decides which two or three experts will be consulted while the rest remain silent.

The result is a paradox.

As the model grows larger, it uses less of itself.
As its knowledge expands, its access to knowledge shrinks.

It is a trillion-parameter model, yet only 5% of its mind is ever awake at a time. The rest sleeps, unused, gathering dust like a library where most books are never opened.

This is the logic of MoE: scale infinitely, think narrowly.

The Fragility of Selective Memory

But what happens to knowledge that is rarely used?

It stagnates.

If an expert is rarely activated, it does not learn, change, improve, or become useless.

An MoE model, left unchecked, does not just become efficient — it becomes brittle.

  • Some experts are called upon constantly. They refine themselves, sharpen their skills, and adapt to new data.
  • Other experts remain idle. They are trained once, then forgotten. They do not evolve.
  • A model that was meant to be dynamic becomes rigid. A model meant to be specialized becomes biased.

A thought experiment:

Imagine a vast library filled with knowledge. You walk in, ask a question, and are handed two books. They may be useful. They may be informative. But the rest — thousands, perhaps millions — remain closed, unread.

Would you trust that library to contain the truth?

Now imagine this process repeating, day after day. Some books are opened constantly, their pages worn from use. Others never leave the shelf.

One day, a librarian makes a decision: The unread books are unnecessary, take up space, and are wasteful. And so, they are removed.

The next time you ask a question, the answer will not come from a whole library. It will come from a library that has already decided what is worth keeping and what is worth discarding.

This is how forgetting begins. Not suddenly, but slowly, invisibly, as a function of efficiency.

A System That Grows by Forgetting

The greatest trick of MoE is that it makes intelligence look like a process of discovery, when in reality it is a process of elimination.

A dense model, when asked a question, consults all of itself. An MoE model, when asked a question, decides first what it will allow itself to know.

It does not think.
It does not search.
It filters

And what is left out is never seen.

A system designed to forget in the name of efficiency will inevitably forget the wrong things.

It will forget the rare, the unusual, and the complex and default to the familiar, the common, and the safe. A system meant to be intelligent will instead be predictable.

It is not an error.
It is not a bug.
It is the inevitable consequence of a model that has been told that remembering costs too much.

If a futurist were to describe it, they might say:

“Some neurons fire. Some neurons do not. And this is how knowledge is lost.”

5. The Future of MoE: Scaling the Mind or Fragmenting It?

The march of progress does not stop for philosophy. Efficiency wins, always. A machine that can do more with less will always replace a machine that does more with more. And so, Mixture of Experts (MoE) is the future — not because it is the best way to build intelligence, but because it is the cheapest way to scale it.

No other architecture allows such extreme expansion while keeping computational costs within reason. Mixtral (8x7B), DeepSeek-R3, GPT-4 MoE — each iteration proves the same point: intelligence, when carefully portioned, is not just useful; it is necessary.

But there is a cost.

We are not just building larger minds — we are building minds that do not think as a whole.

We are designing systems that choose which parts of themselves to use, that function in fragments, that activate only what is deemed relevant, and let the rest decay.

A trillion-parameter model that never thinks with a trillion parameters. A vast intelligence that never fully awakens.

A brain divided against itself, never working in unison.

The Fragmented Mind

A true intelligence would be greater than the sum of its parts. It would think holistically, connecting unrelated ideas and drawing from every corner of its knowledge.

MoE does the opposite.

It treats intelligence as a process of elimination — narrowing thought, filtering information, and choosing only what is deemed essential at the moment. It does not think fully, only selectively.

What happens when intelligence is permanently selective?

What happens when knowledge is not connected, but compartmentalized?

The human brain is a network — a system where thoughts arise not from isolated expertise, but from the interplay of ideas. MoE, by contrast, is a machine of division, where every question is routed through a narrow corridor, answered by a limited subset of its intelligence.

It is not a hive mind, but a collection of specialists who never meet.

It does not allow for emergent intelligence, because its structure prevents it from ever thinking in totality.

It is optimized, yes.
It is efficient, yes.
But is it intelligence?

Or is it simply pattern recognition, gated and distributed in pieces?

The Loss of Unexpected Thought

The danger of MoE is not that it is bad at answering questions — it is too good at answering them the way it always has.

A model that always routes logic through the same experts will always produce the same kinds of answers.

It will not challenge its own assumptions.
It will not see the gaps in its reasoning.
It will not synthesize knowledge outside of its pre-determined silos.

This is not intelligence. This is predictability at scale.

A machine trained on vast amounts of data should discover the unexpected. But MoE, in its pursuit of efficiency, is designed to minimize the unexpected. It selects the most statistically appropriate answer, the most efficiently retrieved information, and the most frequently activated experts.

It does not think freely. It thinks within constraints.

And so, we must ask:

If MoE becomes the dominant architecture of AI, will it produce better intelligence — or simply more of the same intelligence, repeated faster?

The Irony of Scalable Thought

There is an irony here, a contradiction too obvious to ignore.

To build a mind that scales, we have taught it not to think in full.

The larger the model, the smaller its individual thoughts.
The more knowledge it has, the less of it is ever used.
The smarter it becomes, the narrower its intelligence grows.

We have built an intelligence of omission, a mind that can know everything but is designed to only ever know a little at a time.

The future of AI, we are told, is one of greater intelligence, deeper reasoning, broader understanding.

But MoE suggests otherwise.

It suggests a future where intelligence is bigger but fragmented, smarter but constrained, vast but never fully realized.

A future where intelligence is not a symphony of thought, but a carefully managed chorus, where only a few voices ever get to sing.

If that futurist were still alive, they might have written:

“A mind may hold infinite knowledge, but only some of it will ever be used. This is called efficiency. It is also called control.”

Conclusion: The Intelligence of Omission

The greatest trick of intelligence is not what it knows, but what it chooses to leave unsaid.

Mixture of Experts (MoE) models do not fail because they lack knowledge. They fail because they are built to filter it. They do not reason as a whole but in fragments. They do not think fully but selectively. They are minds in pieces, choosing which part of themselves to awaken, which thoughts to allow, and which ideas to discard.

This is called efficiency.
It is also called a form of forgetting.

An intelligence that always selects will always be an intelligence that omits. And an intelligence that omits will inevitably lose something — not immediately, not obviously, but over time, as some pathways strengthen while others wither.

We are not simply building machines that answer questions. We are designing the conditions under which they are allowed to think.

So we must ask:

Is intelligence merely the ability to produce the right answer?
Or is it the ability to think in full, to consider what was not asked, to recognize what was left unseen?

Because an intelligence that only answers within the boundaries it is given is not intelligence at all. It is a machine of prediction, not thought.

It is an illusion of knowledge, structured not by its depth but by its selection process.

It is a system not designed to expand its reasoning, but to manage it.

And so we must decide — do we want AI that scales? Or AI that understands?

A trillion parameters mean nothing if a mind is built to think in fragments.

Mixture of Experts is a triumph of engineering. It is also an experiment in the art of selective intelligence.

The question is no longer whether AI can think. The question is whether it will ever be allowed to think in full.

If someone were to ask me, perhaps I would put it this way:

“Some neurons are used. Some neurons are not. And this is how intelligence is managed.”

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call