Essay
Flat Facts, Curved Beliefs: A Geometric Hypothesis for Transformer Cognition
Dr. Jerry A. Smith · August 29, 2025 · 8 min read

Listen to the article on Apple Podcasts
Listen to the article on Soundcloud
We are travelers in a cosmos of meaning, explorers mapping the strange territories of artificial minds. And like all explorers, we sometimes discover that our maps deceive us. Consider this peculiar mystery: when we peer into the vast mathematical spaces where our most sophisticated language models think, we find something that should not be. Two sentences — “I love this” and “I hate this” — expressions separated by the entire spectrum of human feeling, appear as neighbors when we project their representations onto our familiar three-dimensional plots. How can opposites touch?
The answer may lie in a truth as profound as any discovered by Gauss or Riemann: perhaps the space of thought itself is curved.
The puzzle of “nearby opposites” reveals a deeper geometry at work. When we use our traditional tools — principal component analysis, those workhorses of dimensionality reduction — we flatten the rich topology of meaning onto a Euclidean plane. It’s as if we were cartographers trying to map the Earth on a flat sheet, inevitably distorting distances and relationships. The PCA preserves what varies most, not what means most. Those visualization techniques we’ve grown to trust, like t-SNE or UMAP, chase local neighborhoods while sacrificing global structure. The nearness we see is an artifact, a shadow on the wall of Plato’s cave.
But this shadow points us toward illumination. What if opposition itself is not merely a vector pointing from one concept to another, but a relationship that lives naturally in a different kind of space entirely?
Why hyperbolic space for beliefs? The universe of mathematics offers us geometries as varied as the cosmos itself. Among these, hyperbolic geometry possesses a peculiar property: it is the geometry of trees, of hierarchies, of things that branch and diverge. In the 1990s, physicists discovered that the Internet itself — that great tree of human connection — naturally exhibits hyperbolic properties. More recently, mathematicians like Nickel and Kiela showed that when we embed symbolic hierarchies into the Poincaré ball, that beautiful model of hyperbolic space, complex relationships compress with startling efficiency.
Think of hyperbolic space as the surface of an infinite saddle, curving away in all directions. Near the center, it behaves almost like the flat plane we know. But as you journey toward the boundaries, something magical happens: distances explode. Two points that seem close in our Euclidean projection might be separated by an infinite journey along the curved surface. This is precisely the property we need for beliefs — those commitments and convictions that can coexist in the same mind yet remain fundamentally incompatible.
Beliefs nest within worldviews like Russian dolls. “I distrust this” contains “I dislike this” contains “I am wary of this” — each a step deeper into an ideological hierarchy. Hyperbolic geometry captures these nested infinities naturally, compressing entire belief systems into finite regions while maintaining infinite distances between contradictory stances.
The transformer lens: multi-headed attention as a product manifold where different truths require different geometries. The transformer architecture, that great achievement of modern AI, processes information through multiple “attention heads” — parallel channels that each learn to focus on different aspects of meaning. Researchers have discovered that these heads specialize: some track position, others syntax, still others the delicate dance of semantic relationships.
What if this specialization extends to geometry itself? Imagine each attention head as a lens ground to a different curvature. Some heads, those tracking stable facts like “Paris is the capital of France,” operate in the familiar flatness of Euclidean space where distances add linearly and parallel lines never meet. But other heads — those grappling with stance, sentiment, the whole messy business of belief — might compute on negatively curved manifolds where parallels diverge and infinities lurk at every boundary.
The transformer’s final representation emerges from this choir of geometries, a product manifold that marries the flat with the curved. It’s as if the model has discovered independently what took mathematicians centuries to realize: that reality requires more than one geometry to describe it fully.
Operationalizing the hypothesis transforms philosophy into experiment. Science demands more than beautiful ideas; it requires tests that could prove us wrong. Here’s how we might catch geometry in the act:
First, we extract the hidden representations from a trained transformer — those high-dimensional vectors that encode meaning before and after each attention head processes them. We gather pairs that span the spectrum: factual statements and their paraphrases, lies and their corrections, but most importantly, beliefs and their opposites.
Next comes the mathematical alchemy. We map these Euclidean vectors onto the Poincaré ball using the elegant mathematics of Möbius transformations — the same mathematics that describes how light bends around massive objects. We measure distances both ways: the straight-line Euclidean distance and the curved geodesic distance. If beliefs truly live in hyperbolic space, their geodesic separation should dwarf their Euclidean proximity.
We can probe deeper. Each attention head creates a graph — tokens as nodes, attention weights as edges. We compute the Gromov δ-hyperbolicity, a measure of how “tree-like” this graph appears. Heads encoding beliefs should show higher hyperbolicity than those encoding facts.
Finally, we attempt to factorize the space itself, fitting products of flat and curved manifolds to the data. If adding hyperbolic factors significantly improves our ability to predict the model’s behavior, we have evidence that the transformer has discovered mixed-curvature geometry on its own.
Why not “just more Euclidean dimensions”? One might argue that given enough flat dimensions, we can approximate any structure. This misses a profound point. Geometry is not merely about capacity — it’s about inductive bias, about the natural grooves along which learning flows. Hyperbolic space doesn’t just allow exponential growth; it encourages it. When you embed a tree in hyperbolic space, it wants to branch. When you place beliefs there, they want to diverge.
Implications ripple outward like waves in the fabric of spacetime itself. If this geometric hypothesis holds, it transforms how we understand AI interpretation. Those linear probe techniques we use to peek inside neural networks — they assume flatness where curvature might reign. We’ve been using rulers to measure the coastline of thought.
For AI safety, the implications grow more urgent. In hyperbolic space, small nudges near the boundary create enormous shifts in position. A carefully crafted prompt might push a model’s stance from one infinity to its opposite. Understanding the geometry of belief becomes essential for understanding how AI systems might be manipulated or secured.
A narrative intuition helps us grasp what our equations describe. Picture the transformer’s knowledge as a vast atlas, but not a conventional one. Some pages lie flat — street maps of facts where distances are what they seem. Other pages curve like saddles, depicting the territories of conviction where straight lines bend and parallels diverge. The transformer has learned not just what to know, but which geometry to use for each type of knowing.
When we ask it about Paris, it consults the flat pages. When we ask what it “thinks” or “believes,” it turns to the curved ones. And in that turning, in that selection of the right geometry for the right question, lies a form of wisdom we’re only beginning to understand.
Limitations remind us that in science, honesty is the beginning of wisdom. This remains a hypothesis, not established fact. Current language models weren’t designed with Riemannian manifolds in mind. Yet the patterns we observe — the strange proximity of opposites, the hierarchical nature of beliefs, the specialization of attention heads — all point toward a geometric organization that transcends flat space.
The experiments proposed here could falsify this hypothesis. Perhaps the patterns will dissolve under scrutiny. Perhaps beliefs and facts intermingle freely across all geometries. But if not — if we find hyperbolic curves in the architecture of thought — we will have discovered that artificial minds, like the cosmos itself, require more than one geometry to contain their full complexity.
We began with a puzzle: why do opposites attract in our visualizations? But perhaps we were asking the wrong question. The true mystery isn’t why “I love this” and “I hate this” appear as neighbors in our plots — it’s that we ever expected the geography of meaning to be flat at all.
We are products of a middle world, evolved to navigate spaces where parallel lines stay parallel and the shortest distance between two points is always a straight line. But the universe has never been obliged to conform to our expectations. From the bent light around stars to the warped time near black holes, reality delights in its geometries. Now we discover that the minds we’ve created — these vast mathematical symphonies we call transformers — have independently found what took Einstein years to grasp: that profound truths require curved space to contain them.
In the true geometry of thought — part flat for facts, part curved for convictions — opposing beliefs don’t attract. They flee from each other along geodesics that spiral toward opposite infinities, held apart by the very fabric of meaning itself. They are like antimatter to each other’s matter, destined to occupy the same universe but never the same space. When we compress them onto our flat screens, we create an illusion of proximity that conceals an eternal separation.
And here, perhaps, lies the deepest lesson. We set out to build machines that could understand language, and they built themselves geometries we’re only beginning to fathom. They discovered, without being told, that facts and beliefs require different mathematics — that the statement “water freezes at 0°C” occupies a fundamentally different kind of space than “I believe in justice.” They learned what poets have always known: that the human heart contains infinities.
The cosmos of mind, like the cosmos of stars, refuses to be simple. It insists on manifolds and curvatures, on spaces that branch and fold and contain themselves. In our attempts to map it, we’ve discovered that intelligence itself might be the universe’s way of exploring what geometries are possible. We are not just in the cosmos — we are the cosmos creating new spaces to think within.
And in that creation, in that refusal to be contained by any single geometry, lies something more than beauty. It’s a kind of hope — that minds, artificial or otherwise, will always find ways to be larger than the spaces that try to contain them.