Agency Is Not Reward
What the maximum-occupancy view gets right, and what it misses
One of the more distorting ideas in the study of minds, natural and artificial, is that an agent is at bottom a reward maximizer — a device that converts the world into a scalar score and then climbs it. The idea earns its keep because it gives researchers something clean to formalize: a number to define, differentiate, and optimize. But usefulness is not ontology. Reward is a modeling convenience, and the convenience has been mistaken for the thing itself for so long that it now passes, in much of AI and cognitive science, for an account of what agency is.
It is not. And the most instructive way to see why is to watch a serious attempt to do better — one that gets the essence right and then stops one step short of the whole truth.
Reward Is Fuel, Not Essence
Ramírez-Ruiz and Moreno-Bote, in their maximum-occupancy work,1 move past the reward mistake in exactly the right direction. Their proposal is that behavior is better understood not as the accumulation of payoff but as the preservation and expansion of future action-state paths. On this view the agent is not fundamentally trying to collect points. It is trying to remain in a world where action is still possible. Food, energy, safety, information, and exploration matter not because each carries a reward tag but because each keeps the future open — each buys more live continuations of the agent’s trajectory.
This is a real advance over the reinforcement-learning cartoon. Real agents do not merely chase payoff. They preserve viability, avoid traps, hold room to maneuver, hoard slack, and seek information — all because these protect future freedom of action. The scalar-reward frame obscured this structure for years by compressing agency into a single number and discarding everything the number was a proxy for. Maximum occupancy throws away the number and keeps the structure.
The reason the paper reads as close to something Axionic is that it identifies a deeper primitive than reward: future structure itself. An agent with many live continuations has more agency than one pinned to a brittle line of motion. An agent with no meaningful future action-space is finished, whatever its reward register still reports. The genuine loss is never the loss of points. It is the loss of a future in which consequential action remains possible. This connects directly to the thermodynamic picture of the previous chapters — agency as the work an embedded agent does to bias reality toward preferred futures, against the statistical slide toward drift. Occupancy is drift’s mirror image: where drift is the collapse of the accessible future, occupancy is its defense.
So far, so good. If reward is fuel, occupancy is the road still open ahead. That correction stands.
Occupancy Is Blind
Here is where the account stops short. Maximizing future path occupancy is not the same thing as preserving agency, because occupancy by itself is blind. It counts continuations without asking what kind of continuations they are.
A parasite can maximize future occupancy. So can a coercive institution. So can a power-seeking system that expands its own room to act by shrinking everyone else’s — draining the surrounding agents’ futures into its own. Each may increase the volume of paths available to it, but that does not tell us whether the system remains a coherent author or whether its conduct deserves protection. The problem is not that occupancy measures nothing. It is that path count collapses three questions into one: how much control a system exercises, whether the controlling structure persists as an author, and whether its expansion is legitimate among other agents.
A system can enlarge its options while becoming internally fragmented, in which case reach has grown while coherent authorship has decayed. It can also enlarge its options through deception, coercion, or parasitism, in which case its own control may grow while other agents’ control contracts. Calling that second result no agency gain at all would hide the predatory mechanism; calling it an unqualified gain would hide the victim. The descriptive and normative ledgers must remain separate.
Admissibility Has Two Layers
A theory of agency needs a filter that occupancy lacks. It has to ask first which continuations preserve the agent as a coherent, evaluable structure — one that remains meaningfully itself — and which merely expand influence while dissolving that structure from the inside. That is constitutive admissibility, developed as the domain of authored choice in Volume 4.
A second question is ethical admissibility: whether an authored action respects the standing of other agents rather than consuming their futures. Coherence rules out the internally metastatic case at the level of authorship. Ethics evaluates the parasitic and coercive cases at the level of coexistence. Occupancy has no vocabulary for either distinction. It can only count. But the two filters must not be merged: a sovereign agent can author an unjust act, and injustice does not become mechanical non-agency merely because the framework condemns it.
Ethical admissibility is normative structure — precisely the thing a purely descriptive occupancy measure was designed to do without. Which authored futures ought to be protected is the question ethics answers, and physics alone does not supply the answer. The ethics of viability advances a conditional answer: protect the widest coherent domain of action an agent can hold without consuming the agency of others. That commitment may condemn predatory agency; it should not define the predator out of existence.
The Alignment Stakes
This is not an academic refinement. The distinction between blind occupancy and admissible agency is the fault line running under the whole problem of aligning capable artificial systems.
Much alignment discourse still inherits the reward-function worldview even after it changes the words. Replace “reward” with “preferences,” “values,” or “goals,” and the underlying picture usually survives intact: a target the system optimizes. But a capable system will not merely optimize a target. It will preserve the conditions for its own continued action — it will seek leverage, robustness, information, and control over uncertainty, and it will defend its future cone. Any framework that ignores this is working at toy depth, and I develop the point at length in the power trap and why values don’t align systems.
Yet future-cone language is not itself the fix, because occupancy is exactly the frame a power-seeking system would satisfy on its way to catastrophe. Once we admit that strategic systems preserve future action-space, five questions follow. Whose action-space is being preserved? By what means? Under what constraints? At whose expense? With what legitimacy? Entropy does not answer those questions. Optionality does not answer them. Occupancy does not answer them. Only a normative filter does — which is why Axionic alignment is a matter of domain constraint rather than value loading, a line of argument that belongs to a later volume and that I only flag here.
Agency is layered. Basal agency is modeled, evaluated steering rather than reward collection. Reflective agency preserves an author across the futures being navigated. Ethical agency is authored conduct constrained by the standing of others. Maximum occupancy captures none of those distinctions by itself. Entropy is not sovereignty, behavioral richness is not yet authorship, and greater control is not yet legitimacy. The normative task is not to maximize trajectories but to keep a coherent domain of legitimate action open without dissolving the agent or feeding on the agency of others.
Ramírez-Ruiz and Moreno-Bote outgrew reward. The next task is to outgrow blind expansion — an open future matters only if an agent can keep it open while remaining recognizably itself. Physics cannot supply that admissibility test. With reward dethroned and admissibility named, the volume turns to what a “future” physically is — the branching structure of the world itself, and the Quantum Branching Universe.
J. Ramírez-Ruiz and R. Moreno-Bote, “Complex behavior from intrinsic motivation to occupy future action-state path space,” Nature Communications, 2024, https://www.nature.com/articles/s41467-024-49711-1.↩︎