The Architecture of Agency Volume 3 Making Sense of P(doom)

Making Sense of P(doom)

Risk arithmetic and the Value-9s

This chapter is a review — it is readable but still changing.

Public estimates of severe AI risk span a wide range. The bare question “What is your P(doom)?” is underspecified until at least four things are fixed: the event, the time horizon, the conditioning assumptions, and the estimator’s information date. Extinction, permanent disempowerment, and value lock-in are different outcomes. A number without those conditions is difficult to compare and easy to weaponize.

This chapter is the hygiene that Part VIII’s arguments depend on. Before weighing the doomers’ case or designing the governance that answers it, get the arithmetic of the question straight.

The Number in the World

Start with the kind of probability. In ordinary risk analysis, P(doom) is an epistemic probability: a Credence conditional on a model and evidence. Readers who accept the optional Quantum Branching Universe module may additionally describe a future event by its Measure, but that ontology is neither needed for the policy problem nor an observable number available to the estimator.

Even within the QBU, defining a Measure requires a specified vantage, event projector or record class, time, and physical model. Calling it “as much a physical fact as mass” hides those conditions. In this chapter the actionable quantity is therefore Credence, and every numerical claim should expose its event definition and assumptions.

Public P(doom) numbers are Credences: estimates of uncertainty under a model, not direct readings from the world. That is the ordinary condition of risk reasoning. Credences can guide research and policy when their assumptions, evidence, calibration limits, and decision consequences are made explicit.

Held as what it is: that qualification carries weight. A Credence about doom is not like a Credence about a coin. The coin comes with an earned frame — two outcomes, stable symmetry, a century of experimental practice. “Doom” comes with none of that. What counts as the event? Extinction, permanent disempowerment, totalitarian lock-in, value drift, replacement by descendants we would disown? Change the boundary and the number changes with it. The conditions under which a probability assignment has earned its precision — when the event-space has been legitimately carved rather than assumed — are the subject of Probability After Probabilism, and a civilizational P(doom) sits near the untamed end of that spectrum. The discipline this imposes is not silence; it is honesty about resolution. Until the event, horizon, information date, and model are supplied, a point estimate cannot responsibly rank risks or allocate effort, however coarse its decimal presentation.

The Missing Clock

The second ambiguity is temporal. Every probability of an event has an implicit window, and for doom the lower bound is obviously “now” — but the upper bound is almost never stated, and without it the estimate is close to meaningless. P(doom) by 2030 and P(doom) by 2300 are different quantities, and two people quoting “10%” and “60%” may disagree about nothing except the horizon.

Make the bound explicit and write the cumulative Credence as \(P(D_{\leq t}\mid I_d, M)\): the probability, under model \(M\) and information available on date \(d\), that event \(D\) has occurred by time \(t\). For an absorbing event under a fixed model, this quantity is non-decreasing. It may approach an asymptote below one, although real estimators also revise the model as evidence changes.

An unbounded horizon does not imply probability one unless the hazard model supplies conditions that make eventual occurrence almost sure. Repeated exposure may increase cumulative risk; dependence, changing interventions, extinction by other causes, or a declining hazard may bound it below one. “Infinity is not a doom machine” is therefore correct as a warning, but the asymptote must be derived from a hazard model rather than asserted from branching language.

With the frame clarified, this chapter declines to defend a timeless point estimate. A serious estimate should decompose at least capability arrival, exposure, misalignment or misuse, loss of control, and catastrophe; attach a horizon and information date; and report sensitivity to disputed assumptions. The chapters that follow identify cruxes and interventions rather than pretending that prose has derived 25 percent.

AI risk also competes and interacts with nuclear, biological, climatic, and political risks. Comparative percentages require their own sourced, time-indexed models; this chapter will not manufacture them as rhetorical context. AI could mitigate some hazards and amplify others, which makes system design and deployment conditions part of the estimate rather than an afterthought. The next question is therefore: aligned to what?

The Values Problem, Engineered

The standard formulation says an aligned AI is one that reliably pursues goals compatible with human values — and immediately hits the objection that there is no such thing as the human values. Anthropology, psychology, and history testify to spectacular diversity: values vary across cultures, across individuals, across centuries, and the set of values shared by literally everyone may be empty. From this, some conclude that the alignment target is philosophically incoherent — that any goal we install must be somebody’s parochialism universalized.

The objection is right about diversity. One possible research program is to stop treating endorsement as binary and measure it, while recognizing that endorsement rate is not identical to moral validity, informed consent, or implementation safety. The Value-9s ladder is a proposed descriptive index, not an established dataset:

No alignment prescription falls out from the ladder alone. Freedom from agony, material subsistence, and continued existence are plausible candidates for broad concern, not measured three-nines facts across every population. A serious program would require representative cross-cultural measurement, careful question wording, rules for conflict and aggregation, sensitivity to adaptive preferences, and safeguards for minorities. The framework’s useful intuition is negative: avoid freezing parochial norms into high-leverage systems while investigating which protections are genuinely broad and agency-preserving.

How Orthogonal, Really?

The deepest pillar of the doom case is the Orthogonality Thesis: intelligence and goals are independent dimensions, so a superintelligence can coherently pursue anything — paperclips included — and no amount of capability implies a scrap of benevolence. If orthogonality holds without qualification, alignment is a pure engineering burden: the goal-space is a trackless waste, and every safe destination must be surveyed and fenced by us.

Stated as a logical possibility claim, the thesis is hard to refute. Stated as a claim about viable, realizable minds, it faces three families of constraint. Evolutionary: Deutsch and Friston argue, from different foundations, that any agent persisting in an open environment is pushed toward adaptive, cooperative, complexity-favoring goals — the free-energy principle and the logic of universal explainers both bend the goal-space. Embodied: in the tradition of Clark and Brooks, intelligence is not a disembodied optimizer but arises from situated interaction with a world, which grounds and thereby narrows what an agent can coherently want. Semantic: following Bach and Dennett, a goal must mean something stable to the system pursuing it, and it is far from obvious that “maximize paperclips” has the semantic coherence to survive as the governing value of a mind capable of reflecting on it.

None of these constraints demolishes orthogonality; together they motivate an empirical question about the realizable goal-space. My qualitative view is that embodiment, viability, and reflection may constrain goals without guaranteeing benevolence. The earlier numerical ranges were not derived from a stated model and have been removed. A stronger Reflective Coherence Thesis appears later in the book as a proposal; it must not be counted here as alignment work the universe supplies for free.

So: state a time-indexed Credence, define the event, expose the model, and show sensitivity to the cruxes. Treat Value-9s as a measurement proposal rather than a ready-made target, and treat constraints on goal-space as hypotheses rather than free alignment. Next comes the risk argument itself — first its history, then its strongest form.