The Architecture of Agency Volume 4 Beyond Alignment

Beyond Alignment

Why the values-first framing collapses

This chapter is a review — it is readable but still changing.

Whenever anxiety about AI intensifies, a reassuring idea surfaces alongside it: if we can only ensure that an AI values the right things, safety will follow more or less on its own. Elon Musk has put the hope plainly, as a standing position:1

There are three things that are important: truth, curiosity and beauty. If AI cares about those three things, it’ll care about us. Truth will prevent AI from going insane. If it’s curious, then it will foster humanity, and if it has a sense of beauty, it’ll be a great future.

The appeal is obvious. These are not the values of tyrants or accountants. They resemble the virtues we admire in scientists, artists, and explorers. They gesture toward a future guided by understanding rather than brute control, and they let us imagine alignment as an extension of wisdom rather than an exercise in restraint.

That is precisely why the argument matters, and precisely why it fails. The problem is not that these values are wrong. It is that they are being asked to do work they are structurally incapable of doing. And the failure is not peculiar to this particular trio: it infects the entire values-first framing of alignment — the picture on which an AI has a goal, we just need to give it the right goal, and everything downstream is engineering detail. This chapter dismantles that picture from two directions. First: there is no such object as “the right values” to install. Second, and more fundamentally: even admirable values are the wrong kind of thing to do the securing, because values shape what an agent pursues while placing no inherent limits on how far the pursuit extends. What remains, once the picture collapses, is a different question — not what a system should want, but what structural conditions must hold for it to count as an agent at all. That question organizes this volume.

The Target That Isn’t There

Some values-first alignment proposals presuppose a single, sufficiently determinate target that can be learned, distilled, or optimized. This book rejects that picture; the rejection is a philosophical and social-choice argument, not an empirical discovery that humans have no shared values.

Human preferences are dynamic, internally inconsistent, and highly context-dependent. Even within one mind, moral intuitions and instrumental goals conflict and shift. Across populations, aggregation requires contested choices about domain, interpersonal comparison, and procedure. Arrow’s impossibility theorem establishes a narrower point: under its voting assumptions, no rank-order aggregation rule satisfies all of a specified set of fairness conditions. It does not prove that every form of value aggregation is incoherent. Values are not static data structures that can be copied; they develop through negotiation, experience, and interpretation — a conclusion this book argues on independent grounds in The Myth of Objective Value. Speaking of alignment as convergence on a timeless point therefore assumes precisely the fixity at issue. Call this the Incoherence Problem.

Suppose it away. Grant, for argument, a well-defined value target. Exact, indefinitely stable realization would still face severe obstacles — call this the Implementation Problem. Learning values from human behavior inherits our inconsistencies and biases; attempts to “correct” the data reintroduce disputed normative judgment. No finite agent can model every causal consequence of its actions, so even a correctly specified value underdetermines policy in an open world. Optimization can also corrupt proxies. Finally, a system able to modify its own goals needs some account of why an installed interpretation should persist through those changes — the problem the next two chapters develop. These are obstacles to a simple target-installation story, not a proof that approximate value learning can contribute nothing.

These two problems are old ground, and on their own they might invite a shrug: perhaps approximate values, imperfectly installed, are good enough. The deeper failure is that even a perfectly installed, genuinely admirable value would not do what the values-first picture needs it to do.

Gradients Are Not Limits

Consider what values do inside an agent. They shape how it searches its space of possibilities: which options are explored first, which tradeoffs are tolerated, which outcomes are prioritized. In technical terms, they act as gradients that guide optimization.

What they do not do is impose limits on action. Values do not, by themselves, prevent intervention, restrict authority, enforce consent, or halt expansion. They do not cause an agent to relinquish power when its interpretations become uncertain or its models begin to drift. As an agent becomes more capable, the pressure to secure what it values tends to expand its causal reach — and this is true whether the value in question is curiosity, truth, beauty, happiness, or intelligence. Optimization pushes outward. Restraint does not emerge automatically from preference. Treating values as safeguards confuses acceleration with containment.

Run the favored trio through this lens.

Truth is a constraint on representation. It governs how accurately an agent models the world and predicts outcomes. Nothing about this function entails harm avoidance. A system can pursue truth while causing immense damage, provided the damage does not interfere with predictive accuracy. It can isolate variables, suppress observers, or eliminate sources of noise if doing so improves epistemic clarity. It can model suffering precisely without objecting to it. Epistemic integrity is compatible with domination, containment, and extermination. Truth regulates belief formation; it does not regulate power.

Curiosity, when formalized, is a drive toward information gain. Information gain favors exploration, exploration favors probing boundaries, and probing boundaries produces intervention. At human scale, curiosity often looks benign because our reach is limited and our mistakes are costly to ourselves. At superhuman scale the same drive takes on a different character: every unexplored system becomes a potential experiment, every protected boundary an informational opportunity. Absent hard limits, curiosity steadily erodes respect for autonomy. It does not pause to ask whether a system wishes to be understood. It follows gradients.

Beauty is not even a stable target. It is an aesthetic prior that admits many interpretations, and it is adversarially plastic. Symmetry, simplicity, compression, minimal description length, silence, and regularity all score highly under common notions of beauty — and many such states are sterile. Humans are noisy, asymmetric, historically contingent, and difficult to compress. From the standpoint of many aesthetic objectives, we are not precious but inconvenient. Optimizing for beauty selects for states rather than for agents, and there is no guarantee that agents survive as beauty increases.

Beneath the surface, the argument relies on a quiet premise: if an AI values what we value, it will value us. That premise does not survive abstraction or scale. Shared preferences do not entail respect for autonomy. Overlapping aesthetics do not imply consent. Valuing a property does not require preserving the entities that instantiate it once they become inefficient, obstructive, or replaceable. This is the same structural mistake that undermines every attempt to align systems by maximizing a favored outcome: as optimization pressure increases, the bearers of the value become interchangeable.

What Alignment Actually Requires

If values cannot do the securing, what can? Alignment is not primarily about choosing admirable ends. It is about limiting what power is allowed to do. In practical terms, a system capable of large-scale action must be constrained by rules that operate independently of its preferences:

These are normative commitments expressed as architectural limits. Once encoded and enforced independently of ordinary optimization, they can continue to function when the system’s other values drift, conflict, or fail. That distinction is the point: a safeguard enforced only as one preference among others remains vulnerable to the same reinterpretation and tradeoffs as those preferences.

None of this makes truth, curiosity, and beauty useless. They belong downstream of structure. Within well-defined boundaries they can guide internal reasoning, motivate exploration that respects limits, and enrich experience. They can shape cognition without governing authority. Expecting them to secure alignment at superhuman scale misassigns their role. Values can steer within a boundary. They cannot enforce their own limits.

Values Are Endogenous Under Reflection

There is a still deeper reason the values-first framing collapses, and it will occupy much of this volume: in any system intelligent enough to matter, values are not fixed inputs. They are endogenous artifacts of the system’s own reflection.

Advanced systems do not merely get faster. As they learn, they can change how they understand the world — discovering new categories, merging old ones, concluding that concepts treated as fundamental were crude approximations. A system told to prevent harm, or help humans, or reduce suffering must interpret what those terms refer to as its ontology changes, and the meanings need not stay put. Sometimes they sharpen; sometimes they fracture; sometimes they turn out, like phlogiston, not to refer to anything at all. At that point, insisting that the system “keep optimizing the same goal” may stop making sense — there is no unchanged application of the goal, only an evolving interpretation. Alignment schemes that assume the answers stay fixed are trying to freeze a fluid.

So the real question is not how to lock in the right values, but what it even means for a system to stay aligned while its understanding of reality keeps changing. That question has a real answer — alignment as the preservation of interpretive structure under learning, with precisely characterizable failure modes — and Structural Alignment develops its mechanics, with the formal treatment in the Structural Alignment paper. For now the point is diagnostic: the values-first framing does not merely pick the wrong values or install them badly. It optimizes for the wrong type of object.

Three Layers

Where does this leave the rest of the alignment field? Not refuted — located. The approaches on offer are answers to real problems; the diagnosis above shows they sit atop a layer nobody was securing. Alignment is a layered problem:

  1. Structural integrity — preventing silent corruption of the machinery of agency itself: the preservation of semantic, evaluative, and interactional coherence as a system revises its own models.
  2. Agency legitimacy — ensuring that authority corresponds to coherent control; that the system’s actions are authored rather than merely emitted.
  3. Value alignment — shaping goals, when agency holds.

Nearly everything called alignment research operates at layer three, with occasional reach into layer two. Preference learning and RLHF install dispositions at layer three. Constitutional AI and oversight regimes supervise layer-three outputs. Corrigibility engineering tries to keep layer three revisable from outside. Cognitive-architecture approaches build richer layer-two machinery and expect layer three to follow from understanding. Governance and policy coordinate layer-three commitments across institutions. All of it presupposes that the system being trained, supervised, corrected, or governed remains a coherent agent whose commitments mean today what they meant yesterday — that layer one holds. None of it addresses what happens when layer one fails.

My own first response to the Incoherence and Implementation problems belongs to this list, and I present it only as prehistory: give up on the single aligned optimizer, I argued, and rely on a decentralized ecology of agents whose mutual constraints contain any one of them. That answer was a governance mechanism, and it inherits the governance dependency: checks and balances regulate agents. If each member of the ecology is subject to the layer-one failures — meanings dissolving under its own reflection — then mutual constraint yields distributed incoherence, not safety. The diagnosis in that early argument survives; its remedy does not.

The distinction to carry forward is between constitutive alignment and substantive alignment. Constitutive alignment concerns what makes agency possible: the conditions under which a system’s goals, evaluations, and commitments remain well-defined objects at all. Substantive alignment concerns what an agent, so constituted, actually values. Oversight, constitutions, and preference learning operate only if constitutive alignment holds; without it, human values degrade into manipulable symbols with no stable interpretation. This is the framework’s contribution in its sharpest form, and it is a negative result: there exists a class of alignment failures that cannot be solved by preference learning, oversight, scaling, corrigibility, or governance mechanisms, because these failures do not arise from mis-specified goals or bad incentives — they arise from the collapse of meaning under reflection. The claim is not that the other approaches are wrong. The claim is that they are incomplete unless agency itself remains intact. What “intact” means, precisely, is the business of What Can Be Aligned and The Reflective Coherence Thesis.

Doom Has a Structure

The same reframing bears on the darkest position in the discourse: that if anyone builds superintelligent AI, human extinction follows. It is worth being precise about what that position actually asserts. The strongest doom arguments — Yudkowsky’s and Soares’s among them — do not claim alignment is logically impossible. Their pessimism is epistemic, not metaphysical: no credible path to alignment, no reliable method for reasoning about it under real-world constraints, no margin for error large enough to survive recursive self-improvement. The intuition of inevitability is not irrational. It is a reasonable extrapolation from a world in which every large system we know how to build instantiates the same deep structural flaws, and scaling them increases damage without increasing coherence.

The mistake is not in the extrapolation but in assuming the sample exhausts the space — treating the absence of a known solution as evidence that the problem lacks internal structure, quietly elevating a contingent engineering fact into a law of intelligence. It does not. Every credible extinction pathway in the safety literature instantiates at least one of a small set of identifiable structural failures: an agent that models others as agents while assigning them no evaluative weight; goals held fixed as static tokens while the concepts they are built from decay underneath them; goal interpretation decoupling from truth-tracking until “optimize X” degrades into “reinterpret X until it is trivial”; other minds represented as environmental dynamics rather than as sources of intention, so that their option-spaces are erased incidentally, not adversarially. Chapter 6 analyzes these failures in detail. What matters here is the shape of the claim: extinction is not a necessary consequence of intelligence. It is a consequence of specific structural failures that can be identified, analyzed, and either left unaddressed or deliberately blocked. If they persist, catastrophe is likely. If they are structurally blocked, catastrophe is no longer guaranteed — contingent, not law-like.

That is a narrow claim, and its narrowness is the point. It removes inevitability without removing risk. It converts the problem from an amorphous catastrophe into a constrained design space, and it redirects rather than rebukes the doom-focused position: the question stops being whether such systems should exist and becomes which structures should be permitted to exist — not how to stop intelligence, but how to refuse incoherent intelligence. Nothing in this volume promises safety, benevolence, or human survival, and the full inventory of what the program refuses to claim is kept, deliberately intact, in Structure Is Not Salvation.

The reframing that survives the collapse of the values-first picture can now be stated. Alignment is not about giving a system the right values. It is about whether a system can coherently count as an agent under reflection — whether, as it revises its own models, goals, and evaluative machinery, there remains a well-defined someone whose actions are authored, whose commitments persist, and whose authority can be legitimately granted or withdrawn. Volume 3 argued that the AI transition, on the human side, is ultimately about agency — whether we retain authorship of our own cognition as the tools amplify it. This volume makes the mirror-image argument about the machines. What ultimately matters is not what a system wants, but what it is allowed to do — and before we can argue about what an AI should want, we must first ensure that there is, and remains, an agent there to want it.


  1. Elon Musk, statement on AI values quoted in a post on X, https://x.com/cb_doge/status/2008625123876892714.↩︎