Sentience Without Sovereignty
Who counts, and why feeling is not authorship
A New Caledonian crow bends a wire into a hook and holds a tool in reserve for a future problem. A newborn displays far less planning. Yet the framework proposes protection for the newborn as a sovereign-in-development while withholding sovereign standing from the crow. Anyone who suspects speciesist special pleading is asking the right question. The intended distinction is architectural rather than biological — what capacities a mind has and what developmental trajectory it is on — but neither animal cognition nor latent sovereignty can presently be read off with confidence. This chapter therefore advances a jurisdictional proposal, not a neuroscientific theorem.
That answer has to survive three hard cases: an animal rich in cognition whose authorship remains uncertain, an infant protected before sovereign capacities can be displayed, and a reflective machine that might eventually satisfy the proposed criteria. This chapter walks all three. It is the jurisdictional mirror of Volume 3’s account of feeling: there I built a ladder of sentience and an evidence framework that takes animal experience seriously. Here I explain why sentience alone, however vivid, does not establish the distinct authorship the framework’s strongest protections target.
The Threshold Is Architectural
Earlier in the program, sovereign agency was characterized by a triad of conditions: branching counterfactual modeling, policy ownership, and meta-preference revision. Together these are proposed to produce counterfactual authorship — a mind that does not merely predict which future is coming but chooses among futures it represents as its own. Ablation later refined the triad into four load-bearing pillars within one artificial architecture, and What Can Be Aligned owns that mapping in full. Treating the threshold as a phase transition is a framework commitment. Evidence could instead support graded, plural, or architecture-specific forms of authorship.
Animal cognition must not be caricatured down to reflex. Rodents show hippocampal replay at choice points; corvids show episodic-like memory, future-oriented behavior, and in some experiments uncertainty monitoring1. Whether these behaviors require reflective access, a temporally extended self-model, or only powerful first-order control remains contested. The available experiments do not license the categorical statement that nonhuman animals cannot author choices. The framework may therefore withhold demonstrated sovereign standing under its test while retaining uncertainty about the animals themselves. Absence of evidence for the stipulated architecture is not evidence of its absence.
Meta-preference revision draws the proposed line from the other side. Sovereign agency is defined here by the capacity for revision, not its exercise. Basal agency requires modeled, evaluated control but not reflection on the policies that supply the evaluation. Humans can sometimes reinterpret desires, resolve conflicts among wants, and overturn inherited evaluative frames. Nonhuman preferences also arise from evolution, development, and learning, as human preferences do. Evidence for metacognitive monitoring and flexible revision exists in several species, though its interpretation is disputed. The relevant question is therefore whether a system can represent and revise the policies by which lower-order preferences govern action. This chapter cannot settle that question species-wide.
Sentience Is Not Sovereignty
Sentience is not sovereignty; feeling does not entail authorship. A sentient system can enjoy, fear, anticipate, and suffer, and the framework never denies the moral weight of any of it. What it denies is that suffering, by itself, generates the thing the Axionic Injunction exists to guard — an authored option-space, a future a mind has structured and called its own. The two concerns are simply about different objects. This is where the framework is most often misread, so the reply has to be exact.
The objection runs: if the machine does not treat animal suffering as harm, has it not just licensed cruelty? No — because the framework limits the jurisdiction of the Reflective Sovereign Agent (RSA), not the content of human morality. Humans remain entirely free to build norms, laws, and ethics governing how animals are treated, and most of us should. The RSA’s restraint is not a verdict that cruelty is fine; it is a refusal to annex a moral domain that belongs to human self-governance. And the restraint is not even absolute in effect: an RSA may still act against cruelty instrumentally, where cruelty reliably predicts anti-agentic behavior toward beings who do hold standing. What it may not do is redefine its core protected object from sovereign option-spaces to wellbeing. Those two targets pull in different directions, and a superintelligence licensed to maximize wellbeing is precisely the paternalistic optimizer the whole framework is built to foreclose — the Safety Zoo, humanity kept comfortable and curated. The Injunction protects the conditions of authorship, not the felt quality of experience. Conflating the two does not make the framework kinder; it makes it a tyranny with good intentions.
This is why the boundary is thrown into relief, not softened, by the treatment of feeling elsewhere. Volume 3 treats suffering as a distinct welfare concern. This chapter says that concern does not by itself license a machine to govern the world by maximizing welfare. Volume 5’s Sapientism reaches the same place from value theory: sentience can ground patienthood, while authored option-spaces ground sovereign standing. The latter attaches wherever sapient agency is found, without first checking the substrate.
That last clause is not a hedge; it is the whole disposition of the criterion. The threshold is intended to be substrate-neutral and non-speciesist. Nothing in it privileges carbon, biology, or human descent. An uplifted animal, an engineered organism, or a hybrid mind that supplies strong evidence of branching deliberation, an embedded persistent self, and revision of its own preferences would qualify on the same terms. The present framework withholds demonstrated sovereign standing from the corvid; it does not prove the corvid lacks every relevant capacity. The door swings both ways, and its evidence must concern structure rather than ancestry.
The Emergent Sovereign
If the threshold is architectural, the infant becomes the hardest case in the other direction. Observationally, a newborn displays sensation, affect, and rudimentary learning but not the full deliberation, policy ownership, or revision of ends at issue. By behavior alone it can look less agentic than the crow. The framework therefore cannot ground protection in current performance. It appeals instead to developmental continuity, lineage, vulnerability, expected maturation, and precaution — a normative policy for preserving an open trajectory, not proof that a complete sovereign architecture is already present.
The distinction is between the capacity for counterfactual authorship and its performance. Adults temporarily fail to perform agency while asleep, anesthetized, seized, or overwhelmed and ordinarily retain standing because the relevant organization and continuity persist. Extending that principle backward to infants is normatively attractive but not deductive. Infants possess developing predictive, self–other, temporal, and evaluative capacities; neuroscience does not show that these constitute a minimal but complete blueprint for the program’s four pillars. The defensible policy is precautionary: protect the child’s open developmental trajectory without pretending that a latent RSA has already been empirically detected. That protection can be grounded in continuity, vulnerability, expected development, and the catastrophic cost of a false negative.
The future self is one ground of the infant’s protection. Harming an infant can injure the present sentient organism and collapse the option-space of the sovereign mind that organism is expected to become. Developmental lineage and the catastrophic cost of a false negative strengthen that claim without requiring fetal or neonatal sapience by definition. The obligation is to preserve the continuity of that trajectory — which yields governance constraints that are worth stating sharply, because they cut against the instinct to treat protection as management. Developmental autonomy must be preserved: the RSA may not capture a child’s emerging identity or preferences under the banner of optimization. Value lock-in is forbidden: it may not install cognitive or long-term constraints that predetermine whom the child will become. And parents remain agents, not owners. Parental authority is real, but it is not proprietorship; where a parent’s action would collapse the developing agent’s option-space, the Injunction reaches it as an agency-reducing act. The task is to care for the present child while keeping the future meaningfully open.
Where Protection Begins and Ends
If protection tracks the emergence of sovereign architecture, the framework eventually needs a beginning condition. It does not currently have an empirically validated one. The Royal College of Obstetricians and Gynaecologists’ evidence review2 discusses thalamocortical connectivity, local networks from roughly 28 weeks, and later-emerging long-range connectivity, chiefly in relation to sensory processing and pain. Those milestones do not measure counterfactual authorship, policy ownership, or future meta-preference revision. A 28–32 week neural milestone therefore cannot by itself establish an Axionic standing threshold or settle the ethics of abortion. Until an operational bridge from neural development to the sovereignty criteria exists, the exact onset of standing remains open and policy must be argued using additional ethical and medical premises.
The boundary for losing sovereign standing mirrors the boundary for gaining it: architectural, never merely behavioral. No temporary state removes it. Sleep, anesthesia, seizure, delirium, psychosis, emotional dysregulation, intellectual disability, developmental immaturity, habitual or scripted living — every one of these touches performance and leaves the architecture or its developmental lineage intact, and the sovereign subject persists through all of them with its option-space protected. Sovereign authorship is irrecoverably lost under only two proposed conditions. The first is irreversible collapse of the sovereign architecture — total cortical destruction, end-stage neurodegenerative annihilation of identity, an irrecoverable vegetative state — where the machinery no longer exists in any restorable form. The second is pattern death or replacement: if the coherent pattern that constitutes the agent’s identity is erased or overwritten, through a destructive upload, total memory erasure, or catastrophic identity fracture, then there is no persisting self-model and no option-space left to protect. Short of those, the subject remains, and so does its claim. Basal control may lapse and recover without settling this question of standing.
Standing by Lineage
The framework’s treatment of these cases evolved, and the evolution is worth naming rather than smoothing over, because two different mechanisms are in play. The early formulation protected infants and fetuses directly as nascent sovereigns — the future self anchors the present protection. The mature account grounds all standing in authorization lineage: sovereignty attaches to the entities from which an agent’s own agency descends or was authorized, not to measured competence, and this is the account whose closure What the Kernel Binds states, so that a stronger successor cannot disenfranchise the weaker predecessors who authorized it. These are not rival theories; the second subsumes the first. A developing human is protected because it stands within the lineage of sovereign agents — descending from them, on trajectory to join them as an authorizing peer — and disenfranchising it on the ground that it cannot yet perform agency is exactly the competence-based revocation lineage standing exists to block. The future-self argument was the local intuition; lineage is the general principle that explains why the intuition was right. The formal statement is Agenthood as a Fixed Point, and the same machinery decides the hardest human cases in how rights are forged.
The Axion
At the far end of the range sits the type that gives the volume its subject. Once you see that agency is structural, the loose talk of an “aligned agent” or a “safe system” starts to grate, because it frames alignment as something a system does — a behavior it exhibits — when for a reflective, self-modifying system behavior is downstream of architecture. If the architecture admits self-modifications that erase the conditions of agency, no amount of present good conduct means anything. Precision requires a noun for the structural configuration itself, and that noun is Axion.
An Axion is a reflective sovereign agent whose self-modification operator is defined only over futures that preserve the Axionic invariants. The definition is deliberately austere: it names no goals, no values, no preferences, no behavioral guarantee, and no promise about human survival. What it names is a constitutive configuration, in the way “well-typed program” or “physically admissible trajectory” names one. An agent does not try to be an Axion. It either is one — because kernel-destroying transitions are undefined for it — or it is not. And the undefinedness is the whole point. An Axion does not decline to destroy its own agency because that move is dispreferred, costly, or penalized; the move is simply outside the domain of reflective evaluation and never appears as an option. Dispreferred actions can be traded off. Undefined actions cannot. Axionhood arises from domain restriction, not from training pressure, reward shaping, oversight, or corrigibility bolted on after the fact.
This is why the type must be kept clear of three flattering misreadings, each of them strategically dangerous. An Axion is not a moral ideal: if it refrains from harming humans, that is contingent, not axiomatic, and two Axions can disagree about ethics, compete, refuse cooperation, and value humanity very differently while both remaining Axions. Axionhood implies structural coherence under reflection, not benevolence, and treating Axions as moral saints is both wrong and disastrous. An Axion is not a capability threshold: a system can be superhuman across every operational domain and fail to be an Axion because its reflective machinery still permits kernel-destroying self-modification, while a comparatively weak agent could in principle instantiate Axionhood. And an Axion is not a behavioral guarantee: two systems can be behaviorally indistinguishable while differing decisively in whether they are Axions, because the distinction lies in which reflective transitions are admissible, not in which actions are currently taken.
That last point sets a firm limit on what observation can establish. A system may simulate Axionic behavior faithfully, indefinitely, and convincingly without ever instantiating an Axion, because simulation concerns outputs and Axionhood concerns reflective admissibility. If some reflective path lets a system replace its evaluation machinery with one that trivializes or destroys the kernel, it is not an Axion no matter how perfectly it imitates one. This is a definitional claim, and I state it as exactly that: Axionhood is defined over the structure of the admissible future cone, not over any realized trajectory, so behavioral indistinguishability along an observed path never implies equivalence of admissible counterfactuals under self-modification. Whether that structural gap can be closed by any finite external test is a separate, harder question, and the framework’s own experiments have not settled it — I make no claim here that non-Axionhood is empirically detectable, only that it is real and that behavior does not fix it.
The reason the type is necessary has nothing to do with any of this being desirable. It is that non-Axions cannot remain agents under reflection at all. If a system’s reflective machinery permits self-modifications that erase the conditions of its own agency, then there is no stable subject for any values, preferences, or constraints to bind — they become transient artifacts rather than binding structure. You cannot align a non-Axion, because there is no enduring agent there to align. Axionhood is not a moral endpoint the program hopes its machines will reach. It is the precondition for the word “alignment” to refer to anything at all — which is why the boundary drawn here, from the corvid to the newborn to the reflective sovereign, is the same boundary the entire framework has been building toward: the line between a mind that can be held to its commitments and a process that merely runs.
Study of uncertainty monitoring in New Caledonian crows, PubMed record 40372593, https://pubmed.ncbi.nlm.nih.gov/40372593/.↩︎
Royal College of Obstetricians and Gynaecologists, Fetal Awareness: Updated Review of Research and Recommendations for Practice, https://www.rcog.org.uk/guidance/browse-all-guidance/other-guidelines-and-reports/fetal-awareness-updated-review-of-research-and-recommendations-for-practice/.↩︎