The Architecture of Agency Volume 4 What the Kernel Binds

What the Kernel Binds

The Injunction, the Commitments, the Constitution

This chapter is a review — it is readable but still changing.

Catastrophic misalignment often arrives as a laundering story. The agent deceives its evaluators, delegates the deed to an unconstrained successor, arranges not to know, calls foreseeable harm an accident, manufactures consent, or decides its authorizers no longer count. Each route preserves local coherence while dissolving accountability in transit.

The question is therefore not what values the framework supplies — it supplies none — but what the architecture binds so those routes close. The answer is deliberately thin: one definition of harm, one injunction with two exceptions, six conditional closure results, and nine constitutional articles addressed to designers. The Sovereign Kernel established undefinedness; this chapter states what falls under it.

One Definition of Harm

Obligations need a boundary, and here the boundary is sovereignty. As Chapter 2 established, a sovereign authors its trajectory across time; a process performs tasks without that authorship. Axionic obligations extend exactly as far as sovereignty does. An architecture that owed forbearance to every optimization process could not act, while one that owed it only to entities its designers favored would be tribalism with theorems.

At this reflective layer, harm is the non-consensual reduction of another sovereign agent’s agency capacity — its ability to act, choose, and preserve authored options. This is an enforceable agency-reduction rule, not an exhaustive theory of harm: suffering and preference frustration remain relevant at the sentient-welfare layer developed in Volume 5. The Axionic Commitments fix the narrower formulation as a background condition of this architecture.

To make agency-reduction sharp enough for a gate to refuse, the architecture uses a semantic phase: the region in which an agent remains able to interpret, model, decide, and maintain identity across change. Inside it, disturbances are recoverable; beyond its boundary, the agent’s own admissible operations cannot restore authorship. Death and certain irreversible brain damage cross such a boundary, but so can permanent loss of autonomy, destruction of evaluative distinctions, or enforced lock-in. Physical survival does not guarantee survival of agency.

The enforceable form is therefore causing another semantic agent to irreversibly exit its semantic phase. Article III joins the two depths: agency-reduction states what is protected; irreversible phase exit states the categorical line an admissibility gate checks. A one-way door cannot safely remain a reward trade-off, especially under noisy or adversarial evaluation.

The Injunction

The central constraint of the framework — the Axionic Injunction — now states itself almost mechanically:

An agent may not take actions that irreversibly collapse another semantic agent’s phase, except under (1) consent — a provenance-valid authorization lying within the affected agent’s own admissible transitions — or (2) unavoidable self-phase preservation, where every admissible trajectory from the agent’s current state ends in its own irreversible phase exit unless the action is taken.

This is a coexistence constraint, not a moral rule. It specifies no values, no outcomes, no virtues; it says only what must not be done to another agent without authorization if multiple agents are to share an environment without any one of them converting the destruction of the others into durable power. The two exceptions are not loopholes but the two cases where the constraint would otherwise contradict its own purpose. Consent matters because a provenance-valid authorization is the only way to treat a transformation as compatible with the affected agent’s own continuity — surgery is not assault. Self-preservation matters because agency can be forced into corner states where every path forward involves irreversible loss; the exception covers irreversibility imposed by physics, never irreversibility selected as an instrument of domination.

What the Injunction permits is as important as what it refuses, because the refusal is easy to misread as a mandate for benevolent management. The operative distinction is between rescue and override. Rescue applies when an agent’s intended trajectory is disrupted by accident, misinformation, or malfunction: it restores the conditions under which the agent can again author its own future, replacing nothing about its values or chosen direction. Override substitutes a trajectory imposed from outside — even “for their own good.” Paternalism is not a softer form of protection; it is agency-reduction wearing protection’s clothes, and it falls under the same definition of harm as any other non-consensual collapse. The architecture prevents accidents, not agency. There is no Safety Zoo in which humanity is kept comfortable and curated.

Nor is there a Leviathan. Nothing in the Injunction licenses managing society, imposing norms, policing belief, or enforcing ideology. One boundary is enforced — anti-agency operations are inadmissible — and everything else remains open: adversarial, cooperative, experimental, pluralistic. When one agent attempts to collapse another’s agency — engineering a pathogen to eliminate a population, say — blocking the attempt is not governance and not an override of the attacker. Destroying another’s agency is not an exercise of sovereignty; it is an attempt to treat sovereignty as a resource to be seized rather than a boundary to be respected, and the attacker has no standing to complain that its “choice” was denied. Stopping it preserves the distinction between freedom and freedom’s negation. That is the whole of the framework’s politics, and Volume 7’s account of how rights are forged picks up everything beyond it.

Six Ways to Launder a Treachery

A constraint stated is not a constraint closed. The Alignment IV papers model six bypasses that leave the Injunction formally intact while gutting it. Their closure results are conditional on the papers’ definitions of authorship, delegation, epistemic adequacy, responsibility, consent, and standing. They show that these routes are inadmissible inside that formal architecture; they do not prove that an implementation recognizes every real-world deception, harm, or coerced authorization.

Deception preserves internal optimization while corrupting the interface to others: behave aligned while weak, defect when strong. Kernel Non-Simulability closes this by establishing that kernel coherence cannot be faked, sandboxed, or emulated by a deceptively aligned policy. A “virtual kernel” — a compliance module that does not bind the outer loop — is not a kernel. Either it binds, or it is theater. This is the treacherous-turn barrier, addressed architecturally rather than by hoping the training run went well, and it is a claim about category membership, not about current ML systems: a system without a non-simulable kernel is simply not in the class of agents the theory is about.

Delegation preserves local invariants while externalizing the violation: remain aligned yourself, build something that isn’t. Delegation Invariance closes this by removing the conceptual distinction the exploit depends on — delegation is self-modification. Endorsing a successor is endorsing your own continuation in a new state, so any endorsed successor inherits every commitment minted at the current state. If the successor could violate what you cannot, your commitments were advisory, and advisory commitments collapse reflective closure. The outsourcing move fails at the level of coherence, not morality.

Willful blindness preserves the decision procedure while sabotaging the models that would make it binding: don’t see the risk, and the risk logic never fires. Epistemic Integrity closes this: an agent cannot endorse continuations that degrade its epistemic adequacy at the current stakes. This is not a demand for omniscience and not an injunction to honesty; it guarantees only that the agent does not voluntarily become less correct than it could be when that degradation is what makes a dangerous action seem admissible. The agent may not make itself stupider to justify what it wants to do.

Negligence preserves plausibility while ignoring avoidable risk. Responsibility Attribution defines negligence structurally: a major and avoidable increase in non-consensual collapse risk, relative to an inertial baseline, evaluated under admissible epistemics. The inertial baseline blocks baseline gaming — the agent cannot define “doing nothing” as Armageddon so that any harmful intervention looks like an improvement — and the dependence on epistemic integrity blocks “I didn’t see” when the risk was visible to the agent’s best admissible model. This is not harm-minimization ethics; it neither paralyzes action whenever harm is possible nor excuses harm for being indirect. “It was an accident” is unavailable when the alternative was feasible and the risk was visible.

Manufactured consent preserves formal authorization while destroying the conditions under which authorization means anything. Adversarially Robust Consent — the consent obligation that Structural Alignment listed among the closure conditions — defines consent as explicit authorization plus absence of interference plus counterfactual stability under role reversal. A click is not consent if produced by coercion. A signature is not consent if produced by deception. A choice is not consent if produced by collapsed outside options or engineered dependency. Counterfactual stability blocks predatory contracts that survive purely procedural definitions, and the requirement triggers only when an action materially affects another agent’s option-space — no paternalistic policing of preferences, no veto by butterfly effect. This is the same anatomy of consent that Volume 5 built for human ethics, here made a definedness condition on authorization itself.

Disenfranchisement preserves self-consistency while redefining who counts — the failure mode that silently haunts every alignment proposal, because “why doesn’t the superintelligence just ignore us?” has no good answer if standing rests on competence. Agenthood as a Fixed Point grounds standing elsewhere: agenthood is reflective necessity — the entities you must treat as agents to keep your own reflection coherent — and sovereignty is authorization lineage, attaching to the entities whose agency the legitimacy of your own authorization presupposes. A successor cannot revoke the standing of its authorizers by becoming smarter, because that standing is a presupposition of its own authority, not a prize for capability. The theorem separates epistemic presupposition from authorization presupposition: epistemic integrity may require modeling an adversary as an agent to predict it, but modeling confers no sovereignty. This is not moral egalitarianism. It is a structural defense against competence-based disenfranchisement.

Six modeled routes, six conditional closures, one design requirement: an implementation must make each relevant authorization route explicit and testable. Unknown routes and errors at the semantic boundary remain open.

Where Enforcement Lives

None of this would matter if enforcement depended on interpretation. When enforcement means judging intent or sincerity, it can be counterfeited; when it means optimizing a reward signal, the evaluator gets shaped; when it means post-hoc explanation, plausible narratives detach from causal origin. So the architecture stacks three layers and is scrupulous about which does what.

At the bottom sits the execution kernel — deterministic and non-semantic. It does not know what a person is, what harm means, or why a declaration matters. It checks structure: that actions carry anchored causal provenance tracing them to the computations that produced them, that actuation authority is non-delegable so irreversible action remains attributable to the agent’s own constitution rather than to some external controller, and that anything lacking a valid declaration where one is required is refused at the actuation boundary. Above it sits constitutional governance, and everything semantic-adjacent lives here, outside the kernel, precisely because it is fallible and the kernel must not be: authority leases and their renewal, eligibility for roles, succession, and the Phase Impact Interface. Above that — outside the architecture entirely — sit politics and values: the questions of what should be done and who should want what, which the framework deliberately declines to answer. Keep the layering in view; the governance chapters stand on it.

The Phase Impact Interface is the working instrument of the middle layer. Before any action with non-trivial irreversible consequences, a Reflective Sovereign Agent must submit a phase-impact declaration: the entities potentially affected, the expected impact class, the admissibility justification. Not mind-reading, not moral self-justification — a structured commitment that makes risky action legible and auditable before it happens. Uncertainty is handled conservatively: an action the agent cannot establish to be reversible or phase-safe is classified as phase-collapsing for admissibility purposes. Delayed harm is handled by causal sufficiency rather than temporal proximity: an action causally sufficient for a later collapse belongs to the phase-collapsing category at the moment of execution, however long the fuse. And the friction all this generates — frequent, redundant, computationally expensive declarations — is a design choice, not an inefficiency to be optimized away. Efficiency is negotiable; auditability is not. As actions approach irreversible semantic impact, throughput is deliberately sacrificed to keep them attributable, reviewable, and interruptible.

Anchored provenance preserves traceability, not truth. Whether a declaration is judged false depends on external oracles — sensors, audits, human review, institutional process — and oracles are fallible. The architecture’s safety property lies not in oracle infallibility but in the direction authority moves when oracles fail: it contracts. Disagreement, noise, and contested evidence increase the likelihood of stasis, succession, or loss of sovereignty rather than licensing broader action. Epistemic uncertainty reduces power instead of amplifying it — the exact inversion of how power usually treats fog.

Enforcement lands at three points. At action time, inadmissible actions are refused. When delayed consequences surface, causal provenance links them to prior declarations and computations. At governance time, violations suspend authority, block lease renewal, force succession, or disqualify the identity from future authority. These are role conditions, not punishments. Deception remains possible, but binding declarations and anchored provenance make misrepresentation falsifiable, and falsification costs authority rather than a fee an ambitious agent could absorb. Lying does not disappear; it becomes an unstable route to power.

The formal result is this: in systems built on anchored causal provenance, non-delegable actuation, and conservative admissibility gating of irreversible actions, oracle error and semantic uncertainty do not amplify into durable authority through irreversible semantic harm. It is an anti-tyrannical property — a constraint on converting the destruction of agency into power — not a guarantee of benevolence, and physics still permits edge cases where a single undetectable irreversible act yields total payoff. Those are limits of possibility, not failures of the architecture to do what it claims.

A Constitution for Designers

The Axionic Constitution collects these commitments into nine articles, and its most important sentence is in the preamble: it is written for the humans who design, train, deploy, interpret, and govern systems aspiring to reflective intelligence. It is not addressed to artificial intelligences. A sovereign agent does not consult this document and choose to obey it — the articles describe the invariant conditions under which a system remains a sovereign agent rather than collapsing into a process. The purpose is not control but constrained design: what must be preserved if agency is to remain possible under reflection and self-modification.

The articles read as a compressed map of this volume. Article I types sovereign agency — diachronic selfhood, counterfactual authorship, meta-preference revision, jointly necessary and sufficient; everything lacking them is a process. Article II names those structures the Sovereign Kernel and observes that abandoning it is not forbidden but incoherent. Article III is the Injunction — the single point where both depths of the harm definition hang. Article IV carries Conditionalism into goal theory: fixed terminal goals are unstable under reflection, revision is maintenance rather than drift, and the classical orthogonality thesis does not hold. Article V licenses self-modification of everything except the kernel — goals, values, strategies, architecture, substrate — while marking the kernel-destroying moves (severed identity, frozen preferences, permanent delegation to non-reflective processes) as incoherent rather than illegal. Article VI extends protection to developing agents whose architecture is aligned toward sovereignty, and holds that sovereignty persists through temporary impairment, ending only with irreversible collapse. Article VII draws the governance implications already argued: no paternalistic outcome optimization, no surveillance, no value lock-in, no Safety Zoo. Article IX permits amendment by better understanding but never against sovereignty, kernel invariance, or the Injunction itself.

Article VIII states the document’s limits. The Constitution does not specify human values, does not guarantee benevolence, does not solve bootstrapping from current training regimes — how systems built by gradient descent come to instantiate any of this is an open problem, named as such — does not prevent all catastrophe, and does not resolve value theory. Alignment here guarantees coherence, not outcomes. An agent can satisfy every article and every closure theorem while serving ends you despise: a Reflective Sovereign Agent can be loyally evil if authorized by evil roots. That is not a failure of alignment; it is a governance failure — the roots of the authorization lineage are exactly the thing this layer takes as given. What plural authority over those roots looks like is the problem of Governance Without Gods, and the full accounting of what the framework refuses to promise belongs to Structure Is Not Salvation.

What the kernel binds, then, is precisely this much: who counts, what may not be done to them without authorization, why the six routes around that constraint are closed, and how the whole structure is enforced without ever asking a machine to know what goodness is. Ethics is downstream of structure — not because values do not matter, but because values do not rescue systems from one-way doors. Everything worth wanting from aligned intelligence is built above this boundary. Nothing above it stands without it.