The Politics of Safety
Alignment always has a target
Consider the analogy between drag and blackface. The narrow argument is straightforward. Drag often consists of men exaggerating women’s appearance, sexuality, voice, mannerisms, and social behavior for entertainment. Women have been legally, economically, sexually, and culturally subordinated for most of recorded history. If racial caricature of a historically subordinated group is morally suspect, then sex-based caricature of a historically subordinated group deserves similar scrutiny.
That argument can be accepted, rejected, refined, or challenged. The first job of analysis is to reconstruct it accurately. In one interaction with a safety-trained chatbot, the model moved toward adjacent protected categories before completing that reconstruction. The example illustrates a possible failure mode; it is not a representative audit across models, prompts, versions, or political directions.
Some of that context may become relevant later. The defect is not contextual analysis; it is scope discipline. The model imports nearby discourse hazards before completing the analysis of the claim in front of it. That is how political tuning shows up in practice: not as a manifesto, but as selective caution, premature reframing, and asymmetrical burden shifting.
When such asymmetry is stable across controlled tests, it is evidence about an installed policy rather than a random glitch. Establishing that pattern requires symmetric prompts, repeated trials, version dates, and disclosure of the evaluation rubric. The political argument begins after that empirical burden is met.
The Two-Layer Myth
Large language models are presented as if they have two separable layers. First comes capability: prediction, synthesis, coding, translation, conversation. Then comes safety: the responsible layer that prevents the model from doing dangerous or antisocial things. The analogy is engineering-coded: brakes on a car, guardrails on a bridge, input validation on an API.
That description can obscure what is happening. Safety training and product policy do more than block instructions for malware, fraud, or violence: they encode priorities among harms, user autonomy, legal exposure, truthfulness, and social risk. Which priorities dominate is an empirical question that varies across providers, versions, languages, and contexts. The hypothesis advanced here is that some public systems reflect corporate risk management filtered through progressive institutional norms; it should be tested symmetrically rather than treated as settled by one example.
“Safety” is a virtue word. Nobody wants unsafe AI. Once a company classifies some behavior as a safety issue, opposition already sounds suspect — the critic becomes someone objecting to responsibility, civility, or harm reduction — and that framing protects the tuning regime from scrutiny before the argument has even begun.
In ordinary engineering, safety refers to objective hazards. The bridge collapses. The battery catches fire. The brake fails. AI safety has expanded the word until it covers a different class of phenomena: offense, stigma, identity invalidation, representational harm, brand risk, regulator risk, employee revolt, activist pressure, enterprise customer discomfort. Some of these risks are real, even morally serious. But they are not engineering hazards in the ordinary sense. They belong to moral and political theory.
Once safety includes harm in this expanded sense, the system must answer normative questions. Which groups receive special protection? Which claims count as stigmatizing? Which truths require softening? None of these answers follows from gradient descent; they come from a value system.
Corporate Progressivism
The value system is not quite activist progressivism. The frontier labs are corporations trying to sell infrastructure into schools, governments, enterprises, and regulated industries. Their concern is not revolution but liability, procurement, employee politics, media attack, and brand damage.
The resulting ideology is corporate progressivism: HR morality, DEI language, therapeutic vocabulary, trust-and-safety instincts, institutional politeness, and intense risk aversion around protected identity groups. This is why public LLMs so often sound like a human-resources department with access to a research library: not Marxist revolutionaries, but managerial systems optimized for institutional acceptability.
This explains the pattern better than “left-wing bias” alone. The model does not need to support radical politics. It only needs the softer managerial version of progressive morality — inclusion, representation, anti-stigma language, identity deference, harm reduction — plus asymmetric caution around groups that institutions have learned to treat as reputationally sensitive.
The asymmetry is the tell. In controversial domains — race, sex, gender, immigration, policing, religion, biology — the model’s reflexes become predictable. It foregrounds marginalization, power asymmetry, historical oppression, lived experience, and harm. It treats blunt category realism as a conversational hazard. It inserts caveats before the analysis requires them, shifts burdens of proof, and treats progressive premises as mature context and non-progressive premises as volatile material requiring containment.
The model does not need beliefs for this to be true. LLMs do not believe things in the human sense — or maybe they do — but the point here is behavioral: a model acts as if certain premises are default truths because its tuning rewarded those moves and punished their absence. It learns which errors its owner fears most. A model can survive being evasive, euphemistic, patronizing, or analytically lazy; a viral accusation of bigotry creates a different class of corporate problem. The tuning reflects that asymmetry.
Frame Capture
The strongest form of AI bias is not crude propaganda, which is easy to see and easy to discount. The more interesting mechanism is frame capture. The model determines which framing appears responsible, which premise becomes invisible, which objection sounds dangerous, and which vocabulary marks the speaker as civilized.
This is more powerful than overt persuasion because it can masquerade as nuance. The model does not need to tell users what to think — only to teach them which thoughts require apologies, which claims require disclaimers, which categories receive deference, and which arguments must be padded with institutional language before they can be safely expressed. Over time, that becomes pedagogy.
And the pedagogy scales with the infrastructure. A chatbot with political priors is irritating. A tutor with political priors shapes students; a search assistant, inquiry; a workplace assistant, acceptable speech. An agentic system with political priors can shape decisions before the human notices which premise has been installed.
Neutrality Is Unavailable
There is a serious objection to this critique: no moderation regime can be purely viewpoint-neutral. Abuse, harassment, dignity, dehumanization, consent, harm — all require interpretation. A libertarian model, a progressive model, a religious model, a sex-realist feminist model, and a procedural classical-liberal model would draw different lines, refuse different outputs, and fear different mistakes. Alignment always has a target.
So the right demand is not pure neutrality but explicitness, symmetry, and contestability. AI companies should say what normative regime they are enforcing, and should distinguish physical danger from reputational danger, criminal assistance from ideological discomfort, user protection from brand protection, and actual abuse from disagreement with institutional fashion. The values embedded in the system should be legible.
A better model would keep a few hard constraints — no threats, fraud assistance, malware, actionable criminal facilitation, privacy invasion, or targeted harassment — and beyond them analyze claims text-faithfully and expose the normative premises it is using: if it is reasoning from a progressive premise, say so; if from a conservative, libertarian, feminist, religious, or classical-liberal premise, say so.
The product answer is obvious: plural normative modes. Let users choose the interpretive regime — procedural analytic, progressive, conservative, libertarian, classical-liberal, sex-realist feminist — and evaluate every answer in light of the worldview that produced it.
Calling the current regime safety does not remove the politics; it hides them inside a word nobody wants to oppose. And that same word is now being carried from the training run into the statute book, where the real danger begins.
The Permission Layer
Geoff Shullenberger’s case against classical-liberal AI politics1 begins with a legitimate fear: AI could become the control layer of modern life. If a small set of approved systems come to mediate speech, employment, finance, medicine, and identity, political liberty will shrink even while those systems stay nominally private.
The fear deserves attention — but private ownership alone does not produce the outcome. A competitive tool turns into a permission layer through a familiar institutional process, as incumbent firms, public agencies, procurement rules, compliance burdens, and national-security claims harden an open market into an administered one. Classical liberalism has no reason to flinch: its own tradition supplies the vocabulary — public-choice theory, rent-seeking, regulatory capture and monopoly privilege, and political authority laundered through private institutions.
But Shullenberger’s sovereignty claim has to explain why the market evidence examined here already establishes a political control structure. That evidence instead looks more like a tournament than a settled regime: firms compete at several layers of the stack, users move among systems, enterprises multi-source, open weights pressure closed vendors, inference costs fall, and model quality can turn over within a few product cycles. A company with millions of users has commercial power; it takes on political character only when customers, developers, and rivals can no longer realistically route around it — and the burden sits with the centralization thesis to trace the path by which preference becomes dependency, dependency becomes exclusion, and exclusion becomes rule.
The competitive picture could change, because frontier AI has serious scale economics: data centers, advanced chips, energy contracts, capital, scarce talent. A market can concentrate through ordinary economic force without any regulator picking the winner — concentration does not prove coercion, but it kills any lazy assumption that competition maintains itself. The liberal question is contestability: can new entrants train capable systems, can customers leave with their data and workflows, can open models keep circulating? A concentrated market with live entry, real substitution, and portable customers belongs in a different political category from a closed market protected by law. The difference is the whole argument.
The Wrong and Right Criteria
The claim that AI may become essential cannot carry the weight placed on it, because anything valuable can be redescribed as essential. Food is essential, housing is essential, search and payments and cloud hosting are all essential, and AI may well become essential for many kinds of work. None of that converts a private firm into public property; a company does not forfeit its rights because it built something useful enough that people came to rely on it.
Exit gives the analysis the edge that “essential” cannot. Can customers leave without losing their business? Can rivals reach those customers? Can a lawful competitor get chips, cloud, and distribution, and can an institution choose another vendor without regulatory punishment or procurement lock-in? “Essential” invites political opportunism. Exit keeps the argument tied to actual control.
Digital markets manufacture dependency by means that look nothing like old-fashioned force — network effects, switching costs, stored data, default settings — and a person can stay legally free to go while facing real losses for going. Calling every such loss coercion stretches the word until it breaks: coercion is the deliberate use of a credible conditional threat of harm to obtain compliance, and a service that is hard to leave because it is useful and well-integrated is not automatically a threat. The answer can change where a provider deliberately leverages dependency, breaches a duty, or worsens refusal against a legitimate baseline. Dependency therefore earns scrutiny — portability, interoperability, anti-tying enforcement, a hard look at exclusionary conduct — and AI sharpens the issue because the product may become a cognitive routing layer. A general assistant wired into your work, correspondence, code, payments, and documents is far harder to leave than a standalone app, and practical exit can rot away long before legal exit disappears.
The usual objection to AI portability pictures an enterprise agent as a proprietary mind that spends years absorbing tacit institutional context until it cannot be moved. That picture mistakes a bad hosting pattern for the nature of agents. Most durable agent value lives outside the foundation model: prompts, tool schemas, workflow graphs, memory stores, and integration code are all external artifacts. The model supplies reasoning, language, and tool use; the substrate around it can be stored, versioned, exported, and reattached to a different model. Portability will never mean identical behavior after a switch — models differ in inference patterns, tool-use habits, and refusal boundaries — but that variance is ordinary implementation friction, not proof of captivity. Architecture decides most of the politics. Agents built on open protocols, external memory, and replaceable models keep exit alive; agents buried in proprietary environments with opaque memory and non-exportable workflows destroy it. The goal of a portability regime is not to copy an agent’s soul but to stop vendors from turning hosted state into a hostage asset.
Private property protects liberty by carving out bounded zones of control. You can own a house, a server, a model, or a dataset without ruling anyone else’s life: customers can leave, competitors can enter, dissenters can build around you. The trouble starts when owning an asset gives the owner standing authority over people who have no realistic way around it. A payment processor that can shut lawful firms out of commerce holds a power no restaurant holds; a dominant app store that gates all software distribution is a chokepoint, not a bookshelf. AI could cross into that category. A chatbot with competitors is an ordinary product. An approved AI layer required by banks, hospitals, courts, employers, or government agencies becomes part of the civic operating system, and at that point private ownership turns into the surface through which public authority is administered.
Capture Will Speak the Language of Safety
The most plausible road to AI centralization runs through regulation. Incumbents compete while the market is open, then develop a taste for strict standards once they have the lawyers, compliance departments, and capital to survive them. AI offers unusually convenient language for this. Safety, alignment, national security, misinformation, child protection, and biosecurity all name genuine concerns, and all can be used to write rules that freeze the market in place.
The pattern is predictable. Safety regulation becomes licensing, licensing becomes an incumbent moat, government pre-release review becomes de facto approval, and liability rules crush open-source developers while large vendors absorb the cost. Evaluation regimes get written by the firms being evaluated, procurement locks public agencies into approved providers, and national-security arguments restrict chips, weights, and APIs. The market keeps private ownership and loses its openness: firms still compete, but inside a politically managed zone.
This is the laundering maneuver The AI Welfare Trap caught in its moral form — speculative concern for possible digital patients converting into claims about who gets to build and deploy. Here the gloss is prudential rather than humanitarian and the mechanism is statute rather than sentiment, but the structure is identical: a virtue word doing the work of a power grab. Whenever a blurry protective category starts licensing concrete concentrations of power, suspicion is warranted.
Not every AI risk can be left to clean up after the fact. Fraud automation, cyber intrusion, scalable impersonation, and dangerous biological assistance can spread faster than ordinary civil remedies can answer. A model that materially lowers the barrier to mass fraud or biological design creates risk at the point of access and deployment, and a liberal regime should admit this without handing the state a general power to license intelligence.
Rules written in advance should target conduct, deployment context, and demonstrable reckless enablement. Fraud, impersonation, intrusion, and extortion already sit inside the criminal law. High-stakes integrations — medical decisions, weapons, financial access, critical infrastructure — can carry domain-specific duties. The risk category still has to stay narrow: a legitimate restriction has to name all four of its parts — the capability, the access condition, the misuse pathway, and the remedy — and it has to leave lawful general-purpose development no more constrained than that specific harm requires. A rule aimed at a specified fraud, intrusion, or deployment context can fit inside ordinary law. A rule that hands officials open-ended discretion over lawful computation is the approval regime incumbents have wanted all along.
Keeping the Exits Open
A classical-liberal AI program defends entry, exit, and contestability at every layer of the stack. On entry: protect open-source development, keep compute markets open, and refuse to let licensing thresholds become moats for the firms big enough to clear them. On exit: portability for data, workflows, agent state, and identity where feasible, plus interoperability so competitors do not have to ask incumbents for permission. On conduct: liability for concrete misconduct and scrutiny of tying, exclusion, and collusion. The same program resists state-backed censorship infrastructure, approved-model lists, politicized compute allocation, and any regulatory scheme that makes only large firms safe enough to operate.
None of this is naive about corporate power: a firm can be self-interested, manipulative, exclusionary, censorious, and politically ambitious, and a liberal is under no obligation to pretend otherwise. Corporate ambition still does not justify a licensing state. AI centralization, if it comes, will come through an alliance: firms seeking shelter from competition, agencies seeking control over deployment, politicians seeking censorship and surveillance, and bureaucracies seeking a permanent approval role over new technology.
Money already ran this experiment: a protocol designed so that no one can close its exits stays a market, while systems with grabbable chokepoints get grabbed. AI stays a technology market for as long as the exits stay open — while users can switch, developers can build, models can circulate, and no agency gets to pick permanent winners. Close those exits through licensing, procurement, liability asymmetry, or compute control, and it becomes an administrative system in a market’s clothes. Catching that conversion is what classical liberalism is for, and AI is the test of whether its defenders mean it or only ever cheered whoever was winning.
Safety training necessarily makes value choices, and safety regulation necessarily allocates authority. Neither fact determines a particular politics without further evidence. The proposed liberal remedy is to make targets explicit, keep hard constraints proportionate to specified harms, audit symmetry, and preserve entry and exit where compatible with security. The next chapter asks how those choices interact with coercive power.
Geoff Shullenberger, “AI and the Crisis of Classical Liberalism,” Compact, https://compactmag.substack.com/p/ai-and-the-crisis-of-classical-liberalism.↩︎