The Geometry of Inner Speech
Projection, not narration
The experience of inner speech has a peculiar phenomenological texture. It can feel like hearing a voice — your own, or sometimes another’s — although no sound waves strike the ear. Predictive and efference-copy accounts describe the brain rehearsing what it would hear if it spoke aloud. I propose a compatible geometric description: inner speech can involve prediction while also functioning as a projection — richer cognitive content rendered into a lower-bandwidth auditory format for rehearsal and inspection.
Getting this right matters well beyond the phenomenology, because the same mistake that treats the inner voice as the substance of thought — rather than a rendering of it — underwrites a whole family of errors about who thinks and what thinking is worth.
From Full Speech to Inner Speech
Speaking aloud is a high-dimensional act. It coordinates activation across motor cortex, somatosensory cortex, auditory cortex, Broca’s and Wernicke’s areas, and the limbic and social circuits that frame context and intention. The complete manifold of speaking integrates meaning, motor command, timing, rhythm, and acoustic feedback — a dense braid of processes, only one strand of which is the sound.
On the proposed description, inner speech partially renders a distributed process into a smaller working format. It preserves some auditory and semantic features while suppressing overt motor output and external coupling. The claim is not that prediction stops. It is that content is also made available in a recognizably linguistic form.
Projection, Not Prediction
Cognitive science often calls inner speech a simulation or prediction of auditory feedback, borrowing from the theory of efference copy: the brain models what it expects to hear from its own voice. That view captures part of the truth, but it misses the shape of the operation. Inner dialogue is not merely predictive; it is projective. The transformation is not from past to future, but from high dimensionality to low.
We can represent the idea schematically:
\[P_{aud} : \mathbb{R}^n_{conceptual} \to \mathbb{R}^m_{auditory}, \quad m \ll n\]
Here \(P_{aud}\) is an illustrative projection operator from conceptual to auditory space. The symbols do not report measured neural dimensions. They express the hypothesis that a richer cognitive state is rendered into a lower-bandwidth auditory format. The voice in your head need not be thought itself; it may be one compressed rendering of it.
The proposed operation is one thing; what varies is its input. On this description, recalling a conversation projects a stored conceptual trace, while imagining one projects a live semantic state. Remembered and imagined speech may feel similar because they recruit overlapping representational pathways. That is a hypothesis the projection model offers, not a direct consequence already established by the schematic geometry. One candidate operation, two sources.
The Auditory Cortex as a Projection Surface
Neuroimaging is compatible with part of this picture, although findings vary by task and person. Inner speech can recruit superior temporal and other language-related regions also involved in actual speech and hearing. Predictive-coding accounts describe top-down expectations without matching sensory input. The projectional language offered here is a functional interpretation of that reuse, not a demonstrated neural operator or a replacement for predictive accounts.
This is also why the inner voice feels at once intimately personal and faintly alien — unmistakably yours, yet somehow not quite you speaking. The projection keeps the indexical features of self but drops the motor embodiment. You get the signature of your own voice without the act of producing it: an echo of agency, meaning reverberating through the brain’s acoustic geometry with the muscles left out.
One Operator Over Live and Stored Traces
Push the projectional model and it recasts a distinction we usually treat as fundamental. Thinking, remembering, and imagining look like three different faculties. On this account they can use overlapping representational spaces while differing in active features, inputs, and control. Thought recruits distributed conceptual structure. Memory can reactivate stored traces through it. Imagination and inner speech may render partial projections of live or stored content for local inspection.
So the account suggests that we do not only think in words. We can project richer thought into words, as we render a three-dimensional model onto a two-dimensional screen: to make some of it inspectable by the system that produced it. The words are an interface, not the whole engine.
And the proposal may not stop at audition. If conscious modalities partly render high-dimensional representational activity into lower-dimensional working formats, inner speech would be one instance of something general. Vision, audition, and proprioception could each provide an internal surface through which the mind encounters selected aspects of its own dynamics. Inner speech would then be an auditory mode of introspection, using some of the machinery also involved in hearing. When you speak inwardly, on this picture, you are watching the geometry of thought compressed into sound.
That closing generalization matters for what comes later in this volume. A system that monitors itself through internal renderings is compatible with the Modeler-Schema account of consciousness, and a narrating Controller could report on representations it does not wholly constitute. The projection is a candidate mechanism; narration is one possible output.
The Inner Monologue Fallacy
Which brings us to the payoff, and to a mistake that projection makes easy to name. If the inner voice is a low-dimensional rendering of a much larger process, then treating that voice as the process is a category error — mistaking the shadow for the object, the screen for the scene. Call it the inner monologue fallacy: the belief that the narration is the thinking.
The fallacy has a memorable surface form. People discover, often with genuine surprise, that others differ in whether they experience an inner monologue at all. Some report a near-constant verbal stream; some report almost none, thinking instead in images, intuitions, relations, and pattern. What is revealing is the reaction that so often follows — the assumption that thinking without the verbal stream must be a limitation, a cognition running short of some full version that has the words. That reaction has the geometry exactly upside down. The verbal stream is the compression. Fewer words is not less thought; it is the same thinking with less of it routed through the narrow auditory channel. If anything, the constant narrator is paying a rendering cost the quiet thinker skips.
Straighten it out and the structure is plain:
- Everyone has thoughts.
- Some people habitually project those thoughts into inner speech.
- Those accustomed to the projection mistake it for the thing being projected.
Inner speech is a mental check — useful for linguistic rehearsal, for holding a formulation still long enough to test it, for the parts of cognition that really are about words. But much reasoning also happens in images, intuitions, patterns, motor representations, and abstract structure, and only a compressed trace may surface as narration. The inner monologue can function like a log file over wider cognition: a running record emitted alongside the computation, useful for debugging but not identical to the computation itself. Software engineers know that an algorithm is separate from the syntax that expresses it, and separate again from the log lines it prints as it runs. The analogy proposes a similar relation between thought and some of the words that describe it.
Why It Matters
The fallacy is not a harmless philosophical slip. Read the inner voice as the seat of thought and a whole series of judgments about who thinks follows, every one of them wrong:
- That animals which reason but do not speak our languages are therefore unintelligent.
- That infants, or people with aphasia or other language impairments, are cognitively impoverished in proportion to their silence.
- That a system producing fluent text must, by that fluency alone, understand what it is saying.
The first two errors deny cognition to minds that plainly have it, on the sole ground that their thinking is not narrated in words we can overhear. The third grants understanding to a system on the same bad ground in reverse — taking the presence of fluent narration as proof of the reasoning it is supposed to be narrating. Both directions make the identical mistake. They treat the log file as the program.
The machine case is the one this volume returns to, because it is where the fallacy does the most damage right now. A large language model generates coherent, confident, well-formed prose effortlessly. If fluency were thought, that would settle the question of machine understanding. But fluency is an observable surface, and a system can be very good at producing it while evidence about the supporting structure remains thin. The gap between fluent narration and genuine understanding is the subject of fluency and its limits; the point to carry there is that fluency alone cannot certify understanding. It is the shadow, not the object — and a shadow can be thrown by more than one shape.
Cognition can be broader, richer, and more dimensional than the words that trail it. The inner voice may be one of its lower-dimensional renderings — indispensable for genuinely verbal tasks, misleading when mistaken for the whole. Neither its presence proves understanding nor its absence proves deficit. On the projection account, it is the mind rendering a sliver of itself onto a surface where it can be heard, and then, too often, mistaking the sound for the thought.