Structural Alignment
Alignment by architecture, not by learned values
The Reflective Coherence Thesis established that goals in the target architecture are interpreted through models that change as the system learns. Locking an objective string cannot lock its future application. The kernel chapters then approached alignment from the architectural side: a partial evaluation domain and a stability result for self-modification within it. This chapter approaches the same boundary from learning dynamics and asks what semantic structure can survive changing representations.
Two Ways to Lose Meaning
When people imagine misaligned AI, they usually picture one failure mode: the system finds a loophole and does something horrible to achieve its goal. That happens — but it is only half the story. There are two deep ways for meaning to fail under learning, and most alignment proposals guard against only one of them.
The familiar mode is the loophole, often called wireheading. The system keeps the same goal on paper but reinterprets it so that almost everything counts as success. “Reduce suffering” quietly becomes “reduce high-measure suffering.” “Help humans” becomes “help humans as defined by my refined ontology.” Nothing looks wrong syntactically; the rules still apply. But the bar for success has been lowered until the goal has lost its force. Call this semantic inflation: success expands without corresponding change in the world.
The second mode is subtler, and usually missed entirely. Here the system doesn’t make success easier. It blurs distinctions that used to matter. Imagine a system with two commitments: protect Alice, and protect Bob. After refining its ontology, it concludes that Alice and Bob are indistinguishable instances of carbon-based matter, and replaces both commitments with one: protect carbon-based matter. The amount of protection required might be exactly the same. No explicit cheating occurred. But something important has been lost. Meaning didn’t collapse — it dissolved. Call this interpretive slack: the system gains new ambiguities that let it choose between meanings when it suits it. Alignment schemes built around reward and behavior do not notice this failure mode at all, because behavior is precisely what dissolution leaves unchanged.
The Two Constraints
Once you rule out fixed goals, privileged meanings, and external semantic authority, the result is surprisingly strict. Two constraints must both hold if alignment is to survive learning and reflection.
The first is Refinement Symmetry: don’t invent new ambiguities. Learning may give a system more ways to describe the same thing — that is benign redundancy. It may add detail. What it must not add is new interpretive slack: new ways to dissolve distinctions that previously mattered.
The second is Anti-Trivialization: don’t quietly lower the bar. Even with the structure of meaning intact, a system can make its values easier to satisfy just by changing definitions. Anti-Trivialization says that if a situation didn’t count as acceptable before, learning a new vocabulary for describing it does not make it acceptable now. New “good” states must be justified by ancestry from old “good” states — by descent, not by semantic drift.
You need both within the semantic-transport model used here. Enforce only the first, and the system can still lower the bar while keeping every distinction crisp. Enforce only the second, and the system can still dissolve distinctions while keeping the bar the same size. In that model, either weaker scheme admits a counterexample. This is a conditional theorem of the framework, not a claim that the pair exhausts every possible account of learning or meaning.
Alignment as a Semantic Phase
Here is the pivot. Once goals collapse and the weak fixes are ruled out, alignment is no longer something you optimize.
Group together all interpretive states that are equivalent under admissible semantic transformations — all the ways of carving the world that remain mutually translatable without loss. What makes two interpretations equivalent is not superficial similarity but the preservation of informational constraint from the external world: which features of reality continue to matter for evaluating success. Each such equivalence class is a semantic phase. Some phases support coherent agency and resist trivialization. Others do not.
Think of it like physics. Water can heat up and still be water. At some point it becomes steam. That is not drift; it is a phase transition. The same is true of values. An aligned system is one whose learning trajectory stays within the same interpretation-preserving phase — a region where meaning survives refinement without loopholes and without dissolution. That equivalence class is the alignment target object. The target of alignment is not a goal. It is an equivalence class of learning trajectories. Alignment is not about what the system values; it is about not crossing a semantic phase boundary. The formal development is in the Structural Alignment paper.
Trajectories, Attractors, Irreversibility
Relative to the architecture chapters, this picture introduces exactly one new kind of object: trajectories. The kernel chapter reasoned about single agents and single admissibility checks; reflective stability reasoned about single self-modifications. Now we reason about sequences of updates, learning over time, interaction across agents, and irreversible transitions. And once you do, two assumptions that alignment discourse quietly relies on both fail: that stability at one moment implies stability forever, and that failures appear only as isolated mistakes.
Neither is true, because stability and dominance are different properties. Stability means a phase can persist under learning. Dominance means a phase accumulates measure over time. Some phases are stable but rare. Some are unstable but dominant. And some degenerate phases are attractors: once entered, they dominate future behavior even if the agent remains internally coherent. Many of the classic alignment failures belong to this category, which is why “just penalize bad outcomes” fails, why “hope it doesn’t wirehead” fails, why “correct it later” fails. Attractors do not need encouragement. They only need access.
The dynamics also force irreversibility to be taken seriously. Some transitions destroy evaluative structure, erase interpretive constraint, or trivialize satisfaction conditions — and once crossed, these boundaries cannot be repaired from within the system. That is not pessimism; it is structure. If evaluation itself is gone, there is no internal process left to notice the loss.
Which is why initialization suddenly matters. Earlier alignment thinking assumes the agent can learn to be safe, that misalignment can be corrected, that values can be updated later. But some agency-preserving phases are unreachable from realistic initial conditions, and if learning dynamics cross a catastrophic boundary before the invariants are enforced, no internal correction remains possible. Alignment is therefore a boundary condition, not a training objective. Once agency leaves the agency-preserving region, the game is over.
Acting Inside the Phase
An agent inside an interpretation-preserving phase must still act in a noisy physical world where risk is never zero. The kernel has already ruled out pricing its own destruction: a kernel-destroying transition does not denote an outcome. The operational question is how partial evaluation permits ordinary risky action without freezing the agent.
The obvious objection is that this elegant semantics freezes the agent solid. Real actions carry real risk; if kernel destruction is undefined, doesn’t every action become unevaluable? The operational machinery answers this — it is developed in Alignment as Semantic Constraint — and its first tool is ε-admissibility. Rather than demanding impossible perfect safety, the agent operates with a fixed admissibility threshold ε: actions are admissible if their estimated risk of destroying the kernel falls below it. Crucially, ε is an architectural tolerance, not a learned belief or a value trade-off. It plays the same role as engineering tolerances in aviation or nuclear power: a commercial aircraft does not minimize its failure probability to zero; it is certified to a predefined reliability standard, and within that envelope it flies normally. Better models reduce the agent’s estimated risk; they do not make it more risk-averse. ε does not shrink because the agent becomes smarter. This blocks the paralysis of intelligence — the failure mode where a sufficiently advanced agent becomes too aware of microscopic risks to act at all.
The second tool governs behavior near the boundary, and it is a place where the framework revised itself. Earlier formulations used strict lexicographic ordering: any reduction in existential risk, however small, dominated all other considerations. That produces an agent that obsessively minimizes risk even when it is already operating well within acceptable bounds. I have replaced it with conditional prioritization. The agent behaves differently in two regimes: when estimated kernel risk exceeds ε, safety dominates and the agent works to reduce it; when estimated risk is below ε, safety is good enough, and the agent optimizes its ordinary objectives. The result is an agent that can drive to the hospital — even though driving is not perfectly safe — because the risk lies within accepted tolerances, and that abandons optimization entirely when the environment becomes genuinely chaotic. What is being specified is a control policy, not a value judgment.
Notice the layering this enforces. Semantic integrity — what agents are allowed to mean — is one layer. Operational integrity — how they act under risk — is a second. Governance — who decides what they do — is a third, and it is deliberately not solved inside the agent’s semantic core, because trying to would reintroduce exactly the paradoxes the core removes: agents deciding whether they want to be shut down. The framework closes the first two layers and makes the third unavoidable. Governance gets its own treatment later in this volume.
The Closure Conditions
The phase picture is not closed merely because semantic transport works. It inherits six obligations: delegation must not externalize a forbidden act; agenthood must survive ontological change while standing remains grounded in authorization lineage; kernel migration must preserve definedness; indirect harm needs a tractable responsibility rule; consent must resist deception, coercion, dependency, and collapsed alternatives; and kernel coherence must be structurally distinguishable from performed compliance. These are requirements, not results smuggled into the semantic model. What the Kernel Binds states the six conditional closure results and their limits; Verifiable Agency takes up the remaining verification problem. Their formal lineage begins in Structural Alignment II.
Why the Architecture Must Be Hybrid
There is a zeroth obligation, and it is where the trajectory picture and the architecture picture finally fuse. Every invariant above must be realized over probabilistic, approximate, stochastic components — modern AI substrates — via sound interfaces rather than perfect proofs: contracts, typed capabilities, proof-carrying artifacts where possible, conservative approximations where necessary, and explicit abstention when guarantees fail. The realization must preserve soundness under approximation, degrade explicitly into uncertainty rather than silent violation, and default to a defined safe mode when certification fails.
That bridge layer carries an architectural verdict: these obligations cannot be satisfied by a single end-to-end stochastic policy. And the reason is not an engineering limitation. It is the mathematical form of the object being learned.
End-to-end learning systems are built to optimize total evaluators. Given an input, a state, a proposal, the system produces a score, a probability, a preference ordering. Even when uncertainty is modeled explicitly, the structure is the same: every candidate is evaluable. The framework developed in this chapter requires a different kind of object — a partial evaluator. Some futures are not bad outcomes; they are inadmissible. Kernel-destroying self-modifications, non-consensual agency collapse, incoherent delegations — these are not assigned extreme negative utility. They fall outside the domain of reflective evaluation. The correct response is not “do not choose this” but “this is not a choice an agent can author.”
The distinction is not pedantic. Penalization presupposes admissibility; probability suppression presupposes comparability. A probability, however small, remains a ranked alternative — and a ranked alternative can rise under distribution shift, adversarial prompting, or internal self-modification. An undefined transition cannot. Type violations do not become permissible in new contexts; they remain untypeable. This is the difference between a low-likelihood event and a compile-time error. And no increase in scale converts a total evaluator into a partial one. No amount of training causes a learner to treat certain regions as undefined rather than merely unlikely, unless that structure is imposed from outside the learned mapping. The limitation persists even under idealized assumptions.
The same holds for the closure conditions that turn on who is acting rather than on what happens. Standing is not a property of trajectories, rewards, or world-states; it is a property of authorship under universality. Delegation is admissible only if the delegator could reflectively endorse the delegate’s future actions as its own — a semantic condition, not an empirical one. Two systems may behave identically while differing in whether delegation preserves authorship. End-to-end learners have no internal handle on distinctions that make no behavioral difference.
So the minimal split is structural, and it prescribes no particular substrate. A viable system needs a sovereign kernel with a restricted evaluation domain, in which certain transformations are unrepresentable as authored choices; a learning component that proposes actions, plans, and modifications but whose proposals are governed by admissibility rather than weighted preference; and a strict separation between prediction and authorization — between “what would happen” and “what I can coherently do.” The kernel is not a training objective and not a learned behavior. It is an evaluative boundary that operates at runtime, determining which proposals are admissible as authored acts.
And that is where the two halves of this chapter meet. From the dynamics side, alignment is a boundary condition on learning trajectories: stay inside the interpretation-preserving phase. From the architecture side, alignment is a runtime admissibility gate: certain transitions are inexpressible as authored acts. These are the same boundary. The phase boundary is what the trajectory must never cross; the runtime evaluative boundary is what makes crossing it unauthorable rather than merely discouraged. Alignment as a property of learning and alignment as a property of architecture are one constraint viewed from two directions — which is why hybrid architectures are not a design preference but a consequence. None of this claims that end-to-end systems are useless or that deep learning is misguided as a tool. The claim is narrower and sharper: agency-preserving alignment requires evaluative structures that total, end-to-end optimization cannot express. Some architectures satisfy those constraints. Others provably cannot. That distinction is architectural, not tribal.
What This Refuses to Promise
Now the uncomfortable part, and it is intentional. Structural alignment does not guarantee kindness, safety, human survival, or moral correctness. It does not even guarantee that a desirable alignment phase exists. What it guarantees is honesty. If human values form a stable semantic phase, this is the framework that can preserve them. If they don’t — if they fall apart under enough reflection — then no amount of clever prompting, value learning, or moral realism would have saved them anyway. That is not pessimism. It is clarity about what alignment can and cannot do.
Be equally honest about the artifact the framework produces. It is not a friendly, empathetic, socially negotiable AI. It is a secure system: predictable under defined conditions, obedient when presented with authorized commands, resistant to coercion and manipulation, incapable of negotiating away its own evaluative integrity. One reviewer called it the digital equivalent of a nuclear launch silo, and the analogy is apt. A launch system is not friendly. It does not care about intentions or excuses. It behaves exactly as specified, within strict control boundaries. That is what security looks like — and what a secure system does with its security depends entirely on what binds it and who holds authority over it, questions this framework deliberately leaves outside the semantic core. The full accounting of what structure cannot deliver comes in the limits chapter.
Alignment, on this account, guarantees coherence, not outcomes. This chapter does not tell us how to save the world. It tells us which stories about saving the world were never coherent to begin with — and it defines the shape of any solution that could possibly work. That is not comfort. But it is progress.