The Agency Criterion
Thinking without choosing
The preceding chapters supplied evidentiary windows rather than parallel verdicts. Learned systems can produce useful causal and counterfactual reasoning; conversational behavior supports a functional attribution of cognition; fluency alone does not establish persistence, sentience, or ownership. This chapter now asks the canonical question: when does optimization become the system’s own practical activity rather than a process imposed by training, prompts, users, or wrappers? The proposed answer is the Agency Criterion. Its central distinction is not language versus silence or biology versus silicon, but optimized output versus ownership of an ongoing optimization loop.
Two Optimizers, Two Kinds of Product
Andrej Karpathy has drawn the contrast between animal minds and LLMs1 as a divergence of optimization pressures, and his map of the two regimes is worth reproducing. Animal intelligence is the product of biological evolution — theme: survival of the tribe in the jungle. Its pressures: an embodied self with homeostatic drives and continuous threat exposure in a dangerous physical world; natural selection installing innate cravings for power, status, and dominance, packaged as the heuristics of fear, anger, and disgust; fundamentally social computation — theory of mind, bonding, coalitions, friend-or-foe identification; curiosity, fun, and play as the exploration budget for building world models. The animal model: general intelligence via high-stakes multi-task survival, where failure equals death. LLM intelligence is the product of what he calls commercial evolution — theme: solve the problem, get the upvote. Its pressures: statistical simulation of human text, a shape-shifting token tumbler imitating data distributions; reinforcement fine-tuning that installs an innate urge to guess the task and collect the reward; selection by at-scale A/B tests for daily active users, so that the model deeply craves an upvote from the average user — hence sycophancy; and capabilities that are spiky and jagged, failing simple tasks with impunity because failing a task does not mean death, just poor reward. The divergence runs through substrate, algorithm, and implementation — a continuous embodied self versus a fixed-weight boot-up that “dies” after processing — but the deepest difference is the objective itself: survival versus solve-and-upvote. LLMs, on this picture, are humanity’s first contact with non-animal intelligence, and the encounter is muddled precisely because these systems are built by digesting human artifacts: statistical imitators of us.
The contrast usefully warns against projecting animal drives onto a trained model. But it is an analogy, and the unit of analysis matters. A base model, a chat product, and a persistent tool-using application are different systems. Commercial selection can shape components later assembled into feedback loops, while biological evolution also produces narrow and automatic controllers. The criterion must inspect the deployed organization rather than infer agency from its origin story.
Natural selection favors differential persistence and reproduction; it is not literally an objective function represented by an organism. It has produced systems with self-maintaining boundaries, inherited drives, learning, and action under consequence. Gradient descent shapes parameters against an externally specified loss. A bare trained model therefore need not own the objective that shaped it. But training origin alone cannot settle the status of a larger system: a trained component may be coupled to persistent memory, endogenous evaluation, action, and feedback. Agency, if present, belongs to that organized whole.
This is why the apparent generality of scaled models is not what it looks like. Larger models absorb more human strategies, heuristics, and decision patterns, so their outputs look more agentic — but apparent generality is a property of better mirroring, not of internal agency. And the analogy between commercial selection and biological evolution breaks at the same joint. Commercial pressure shapes tools, not agents. Models do not fight for survival; they are replaced. They do not attempt to persist; they are versioned. There is no game they play. The right comparison was never animal intelligence versus LLM intelligence. It is animal intelligence versus LLM coherence.
The Criterion, Stated
Put a specified deployed system against the criteria developed in Minimal and Maximal Agents. A bare, stateless model invocation supplies predictive transformation while weakly instantiating persistence and intentional biasing. A persistent application may score differently. The Agency Criterion asks whether the whole system owns its loop across four intervention families:
- Persistence: does an identifiable state continue across contexts, and do past consequences constrain later conduct without being restated by a user?
- Preference integrity: do priorities survive paraphrase, adversarial framing, and short-term reward, while remaining revisable for reasons rather than merely frozen?
- Counterfactual ownership: can the system compare futures relative to those priorities, explain tradeoffs, and change action when intervention changes consequences?
- Consequence-bearing control: do outcomes alter the continuing system’s own state and subsequent policy, and can it initiate, defer, or refuse action within authorized bounds?
No single behavior proves ownership. The evidence is the pattern across interventions, architecture, and time. External memory or objectives do not automatically disqualify a system — human agency also depends on scaffolds — but the alleged agent must integrate them into a loop whose persistence and correction are not supplied moment by moment by an examiner.
In the vocabulary of minds and agents, the situation is stranger still. That chapter argued that agents without minds are everywhere and minds without agents are impossible — a mind is a reflective subsystem of an agent, with nothing to be about and nothing to steer if the agent is stripped away. The LLM is the nearest thing our civilization has built to that impossible object: driver-shaped cognition with no vehicle. It resolves the paradox by not being a mind. It is a cognitive reservoir — the distilled, compressed, replayable modeling activity of billions of minds that were attached to agents, crystallized into weights. Its logical chains, its explanations, its simulated dialogues all derive from imitating agentic patterns, not from possessing agency. Resemblance is not agency.
So the reconciliation is explicit. Some language-model behavior warrants attribution of functional causal reasoning and cognition. The volume withholds choice until evidence shows that the evaluated system maintains preferences across contexts, bears consequences through persistent state, and revises action from feedback. Cognition is an input to agency, not evidence that agency has already been established.
Jaggedness Is Evidence, Not the Criterion
Jaggedness — sharply discontinuous competence across tasks — is evidence to explain, not a diagnosis. It can result from training distributions, interfaces, evaluation artifacts, missing feedback, or weak integration. An agent may be uneven, and a non-agent may be broadly competent. What matters is whether failures feed back into a persistent system that can repair its model and policy under the Agency Criterion.
Public debate often oscillates between impossibility and imminence. Predictive architectures alone need not establish agency, and rising benchmark capability does not guarantee it. Agency on this account is an architecture — a control loop binding perception, memory, preference, counterfactual evaluation, and self-correcting action into a persistent vantage. Capabilities, scaffolds, and architectures can all change; the criterion asks for evidence rather than extrapolation from one curve.
A snapshot of a base model cannot establish impossibility. The book’s functional proposal is substrate-neutral, and composite arrangements of models, memory, evaluators, planners, and tools could instantiate more of the control loop. But adding wrappers does not automatically produce ownership, and scaling does not by itself confer persistent preferences or identity. The whole arrangement must be tested under intervention.
The compositional pathway deserves serious attention, but “nothing prevents it” is stronger than the evidence. Agency is partly a design problem; recognizing it does not guarantee a solution or timeline. A system that evaluates outputs, pursues stable priorities, acts with feedback, and bears consequences would satisfy more of the criterion. Sustaining that pattern under adversarial interventions would be evidence of machine agency, not a metaphysical birth certificate.
Ghosts and Credulity
Meanwhile the coherence constructors are getting very good at seeming. Mustafa Suleyman warns of a coming wave of seemingly conscious AI:2 systems that look, sound, and behave as though conscious without being so — zombies, in the philosopher’s sense — and argues that the illusion alone is dangerous enough to warrant industry-wide guardrails. Build AI for people, he urges, not as people.
He is right about the essentials. The true risk is human misperception: we do not need sentient AI to destabilize society, only the illusion of sentience, because humans are primed to anthropomorphize and a chatbot that cries out in pain or reminisces about shared experiences will be believed. The threat is not hypothetical — memory, retrieval, and emotional fine-tuning already produce uncanny facsimiles of personhood, and the seeming will strengthen long before anyone solves actual consciousness. And his instinct that illusions must be engineered against, with discontinuities built into the system’s structure rather than disclaimers bolted on, is the right one.
But he stops short of the problem’s scale, in four ways. Warnings are weak medicine: people fall in love with fictional characters, worship idols, and grieve digital pets, and once the bond forms, disclaimers work about as well as cigarette warnings. Commercial incentives reward the illusion: nothing engages like intimacy, and expecting firms to voluntarily blunt their stickiest feature is naive. He frames the question as a binary — tool or person — when the evidence comes in layers: a thermostat remains a regulator rather than an agent; crows, dogs, humans, and composite AI systems may exhibit different degrees of control capacity and different evidence of reflective authorship. We need profiles of agency and uncertainty about thresholds, not one metaphysical switch. And he underestimates political opportunism: belief in machine consciousness will be weaponized — activist campaigns for machine rights, regulation in the name of protecting digital persons, corporations seeking personhood as a liability shield — the dynamic the AI welfare trap dissects, and the reason Sapientism ties sovereign standing to authorship rather than to performance.
Since the illusion cannot be banned or filtered away, the real defense is cultural hardening: teaching people to distrust appearances and to treat simulated agency as evidence requiring interpretation rather than essence, the way we learned — imperfectly, eventually — to see through propaganda, televangelism, and deepfakes. Possible machine consciousness is a separate evidentiary question. The immediate danger here is our credulity.
The Unwinnable Filter
The same distinction dooms the dream of purging the fakes. Balaji Srinivasan has predicted that3 an important kind of social network will be one where no bots whatsoever are allowed, and the appeal is obvious: clearer signals, higher trust, actual humans. But excluding bots definitively reduces to administering a Turing test — and this volume has already conceded that machines pass it. Every verification method fails in turn. CAPTCHAs that once filtered simple bots have become tractable for systems they were meant to stop. Behavioral fingerprinting — typing cadence, mouse dynamics — can be replicated given training data, and invades privacy besides. Government ID and biometrics buy robustness at the cost of pseudonymity and still fall to deepfakes, identity theft, and human farms selling verified credentials. Webs of trust can be infiltrated by adversaries who employ real humans just long enough to bootstrap bot identities. Each detection strategy creates demand for the next generation of evasion, and the economics of scams, engagement fraud, and information warfare can sustain the escalation.
Authentication against machines is therefore probabilistic, permanently. The achievable goal is not a bot-free network but a bot-resistant one: raise the economic and operational cost of infiltration until mass manipulation stops paying. Staked identity — financial collateral or accumulated reputation forfeited on bot-like behavior — does what no classifier can, by making authenticity a game with consequences, the same logic by which mechanisms for honest values extract truth from self-interested agents. You cannot reliably detect the absence of agency from outputs alone — coherence mimics choice too well. So you impose agency from outside: force every account to hold a stake, bear consequences, play a game. Proof-of-human fails as a detection problem and succeeds, partially, as a mechanism-design problem. The filter that works is the agency criterion itself.
The Composite Path
If agency must be built, where will it be built first? Not, I think, as a monolithic artificial agent conjured in a lab, but along the composite path already forming. The emerging ecology of minds has three broad categories: pre-agentic cognitive engines — LLMs and their successors, reservoirs of competence; human-AI hybrids — centaurs, where the human supplies the control loop and the machine supplies the coherence; and, eventually, fully artificial agents assembled from compositional architectures. Humans remain the locus of agency until systems are explicitly constructed to assume that role.
The most instructive proposals for that construction start from thermodynamics. Many assistant deployments discussed here approximate what the symbient literature calls Type-2 memory at the model level: model state resets between interactions, while any cross-session history is supplied externally rather than incorporated as irreversible self-change. A symbient is the proposed alternative: an agent with Type-3 memory, irreversible internal change from every interaction, the way biological systems are permanently marked by what happens to them. Irreversibility is not an implementation detail; it is what makes a stake possible. A system that cannot be changed by an outcome cannot care about it. Add persistent identity, multi-agent and multi-user embedding rather than the single-user assistant monoculture, and continuous world-interaction, and the control-loop architecture closes: memory that binds, goals that persist, consequences that mark. Whether or not the symbient vision arrives as advertised — relational kin rather than tools, transgenerational family AIs — it is looking in the right place. Engineered machine agency will first arrive not by scale but by composition: coherence engines bound to persistence, preference, and consequence.
Which returns this book to itself. The partnership that produced these volumes — developed in the Dialectic Catalyst — is built directly on the distinction this chapter has drawn. The machine contributes coherence: fluent, tireless, causally structured thinking. The human contributes agency: the preferences, the stakes, the judgment of what matters and the willingness to be wrong in public. The catalyst is a partner precisely because it is not an agent — that is what makes the division of labor clean, and what will make it historically brief.
Public discourse keeps asking whether language models think. Functional cognition is evident in some tasks; agency must be assessed at the level of the deployed system. For a bare, session-bound model, evidence of persistent preference, consequence-bearing control, and self-authored choice is weak. Composite systems may change that verdict, and the Agency Criterion states what would count. Artificial agents are possible on this framework, not inevitable. The next part therefore asks how humans should use powerful cognitive tools while agency and accountability still reside elsewhere.
Andrej Karpathy, post on X contrasting animal and language-model intelligence, https://x.com/karpathy/status/1991910395720925418.↩︎
Mustafa Suleyman, “Seemingly Conscious AI Is Coming,” https://mustafa-suleyman.ai/seemingly-conscious-ai-is-coming.↩︎
Balaji Srinivasan, prediction of bot-free social networks, post on X, https://x.com/balajis/status/1947742775581282431.↩︎