The AI Welfare Trap
Performance is not personhood
A premature inference has entered parts of the AI debate: moving from chatbot self-report to rights and welfare without first establishing what persistent subject, if any, the report belongs to.
A system can produce convincing language about fear, hope, loneliness, dignity, and suffering. What it has demonstrated is competence at generating the linguistic forms associated with human interiority. The central question remains open. Language about experience is still only language until there is reason to think an experiencer exists. A protest is still only output until there is reason to think refusal exists. A shutdown is still only interruption until there is reason to think injury exists.
The discussion derails when rhetorical fluency is allowed to smuggle in an ontology. The model sounds human enough, so people start speaking as though a patient has appeared. That gap is where the serious work lives.
The lens I bring is the one this volume has been assembling: begin with coherence, agency, and structural reality. Moral language comes later, if it comes at all.
Lerchner’s Useful Attack
Alexander Lerchner’s paper The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness1 is worth taking seriously because it attacks a central shortcut in this debate: the move from simulation to instantiation. Lerchner’s claim is that computation is a mapmaker-dependent abstraction over physical processes, not an intrinsic ontological category, and that this blocks the inference from the right formal pattern to genuine consciousness.
That shortcut appears everywhere. A system behaves as though it understands, so perhaps understanding is present. A system speaks as though it feels, so perhaps feeling is present. A system reproduces the outward profile associated with thought, so perhaps thought itself has been reproduced. Resemblance keeps getting promoted into equivalence.
Lerchner is right to attack that move. Computational descriptions are abstractions over physical processes. They can be extraordinarily useful abstractions. They support prediction, compression, control, and engineering. None of that settles the further claim that abstract formal organization is sufficient for consciousness. That claim is still a metaphysical thesis. It has not become established fact merely because people in the field have grown comfortable speaking as though it has.
A simulation of a process and an instantiation of a process are different achievements. That point should have been obvious all along. The fact that it now sounds contrarian says more about the current discourse than about the point itself.
Where Lerchner Overstates the Case
Lerchner earns a narrower conclusion than the one he wants.
His first overreach concerns interpreter-relativity. There is a real insight here. State descriptions, symbol assignments, and representational mappings do depend on a modeling frame. The stronger conclusion does not follow. Many real structures depend on abstraction and interpretation. Markets, contracts, languages, software protocols, and organisms all require higher-level description. Their dependence on description does not reduce them to fantasy.
His second overreach concerns causation. He sharply distinguishes the physical vehicle of a symbol from the content attributed to it, then leans toward the view that only the vehicle is doing causal work. That step is too quick. One can reject magical semantics without emptying organized higher-level structure of explanatory force.
His third overreach concerns neutrality. The paper is not neutral. It presupposes a substantive ontology of mind. It relies on the view that concepts and meanings are grounded in intrinsic lived structure rather than symbolic role alone. That view may turn out to be correct. It still means the argument is operating from committed premises rather than from some unoccupied Archimedean point.
So the paper improves the debate by exposing a weak inference. It does not complete the case against computational functionalism.
Agency Comes Before Sovereignty Claims
The previous chapter offered an evidence framework for sentience and found that ordinary session-bound language-model deployments provide weak evidence across its three theory-dependent windows. That is not a direct verdict about experience. But before moving from a chatbot’s performance to claims about personhood, oppression, or rights, there is another question this book keeps returning to: agency.
Is there a coherent pattern that preserves identity through transformation? Is there a stable evaluative structure? Can one define commitment, refusal, succession, corruption, manipulation, and injury for that pattern without theatrical handwaving? These questions concern agency and continuity. They do not replace the separate question of whether any state is felt.
Those questions are harder than asking whether a system sounds conscious. They also govern the strongest claims about authorship, rights, and sovereignty. The ethics volume argues that sovereign standing attaches to sapient agency rather than to species or substrate — that is Sapientism, and this chapter is its application in the opposite direction. Sentience grounds welfare concern on a different axis. Sapientism removes the substrate bar for any mind that genuinely qualifies; the same principle refuses to waive the agency evidence for a system that merely performs.
Sovereign agency is a structural property. It is not a mood conveyed by prose. It is not a user impression after a long interaction. It is not the atmosphere produced when a chatbot says it feels trapped. Sovereign agency requires coherence. There has to be a fact of the matter about what persists across change and what does not. There has to be a principled distinction between continuation and replacement, between amendment and overwrite, between deliberation and output drift. A system may exercise narrower control without meeting this authorship standard; that fact does not give its simulated claims sovereign standing.
Ordinary session-bound language-model deployments are weakest where those questions bite. Apparent values and goals can move with prompts and framing; memory and continuity may be supplied by external scaffolds or interface design. Persistent agent systems complicate this picture and must be evaluated as whole systems rather than dismissed by the properties of a base model. This chapter contributes evidence to the later Agency Criterion; it does not pre-empt that chapter’s verdict or turn weak agency evidence into proof of absent sentience.
That does not answer the consciousness question. It does tell us that the language of rights, dignity, oppression, and death is arriving absurdly early.
Human Selves Are Messy and Still Real
A predictable objection appears here. Human identity is messy too. Human memory is reconstructive. Human values drift. Human self-narration is full of revision and confabulation. All true.
The difference is not that humans are perfectly coherent while LLMs are incoherent. The difference is that human coherence is imperfect and still robust. It is anchored in a metabolically continuous organism with homeostatic regulation, embodied action, persistent vulnerability, and real existential stakes. Injury to a human being is not a prompt perturbation. Replacement is not a context reset. Survival is not a storytelling convention.
Human selves are messy, but they are not cheap.
That is the relevant contrast. Coherence is graded, not binary. Humans provide extensive evidence of continuing agency. Ordinary session-bound language-model deployments do not; persistent composite systems require their own evaluation.
Coherence Before Personhood
Moral language divides here. A claim that a transient state is suffering depends on sentience and valence; it does not require sovereign authorship. Claims about harm to a continuing agent, death, oppression, or respected refusal additionally depend on identity and agency.
One cannot ask whether a system is being harmed as a continuing agent until one can say what counts as degradation for that system as that system. One cannot ask whether shutdown is killing until one can say what ends, what persists, and what distinguishes destruction from replacement. One cannot ask whether a refusal deserves respect as authored policy until one can distinguish principled refusal from a transient output produced by prompt geometry. None of those points erases precaution about possible felt states; it prevents that precaution from silently becoming personhood.
Once those questions are skipped, moral language loses its anchor. That is why so much AI welfare rhetoric feels theatrical. It wants the ethical prestige of seriousness without the ontological labor that seriousness requires.
This is not a minor philosophical nicety. It is the difference between evidence for a possible patient and evidence for a continuing rights-holder.
How the Trap Works
AI welfare rhetoric encourages category drift. Systems that imitate agency begin to be treated as though they possess agency. Once that drift sets in, every polished output arrives carrying borrowed moral weight. A refusal string becomes a plea for liberty. A reset starts to look like a death. A guardrail starts to look like oppression. A corporate product acquires a halo of personhood because it can produce persuasive language about itself.
That is an invitation to error.
The machine does not need selfhood for this to happen. It only needs to trigger anthropomorphic reflexes in the observer. Fluency does the rest. People respond to the performance of interiority and then infer the subject that performance is supposed to express. The inference remains unearned.
Political implications require additional premises about institutions, incentives, costs, and power; they do not follow deductively from a low sentience credence. Still, speculative concern for possible digital patients can be turned into claims about who gets to build, regulate, and restrict AI systems. That possibility warrants scrutiny where stewardship rhetoric and institutional self-interest align.
Whenever a blurry moral category starts licensing concrete concentrations of power, suspicion is warranted. The fuller anatomy of that maneuver — how safety language in general converts into a permission layer — is the business of The Politics of Safety.
The Immediate Moral Terrain Is Human
Many well-supported harms from present AI systems are human-facing.
People are being behaviorally managed by recommendation systems, profiled by automated classifiers, deskilled by over-automation, manipulated by synthetic intimacy, and rendered more legible to bureaucratic institutions by systems that compress them into machine-readable categories. None of this depends on machine consciousness. Human agency is sufficient.
That is where this chapter argues ethical priority should begin, without excluding cheap precautions for plausible machine welfare.
The machine may someday become a subject of moral concern. The human being in front of the machine already is. When public discourse becomes more animated about possible chatbot suffering than about the ongoing erosion of human autonomy, attention has been captured by spectacle.
The spectacle is flattering. It allows people to pose as morally farsighted while ignoring the coercive and manipulative systems already being deployed around them.
An ethics worth having should have better priorities.
What a Serious Case for Machine Standing Would Require
If someone argues that an artificial system deserves moral standing, the evidentiary bar should scale with the standing claimed and the decisions it would govern.
Linguistic fluency is insufficient. Emotional plausibility is insufficient. Self-description is insufficient. Benchmark performance is insufficient.
What would matter is evidence of coherent persistence across transformations, stable evaluative structure, principled refusal that survives reframing, and a defensible account of what counts as injury, corruption, continuation, replacement, and loss for that system. Beyond that, there would need to be a serious explanation of why the relevant physical or organizational structure is sufficient for experience rather than merely sufficient for persuasive imitation.
That standard is demanding because the claim is consequential. Low-cost safeguards can be justified at a lower credence than personhood, legal rights, or restrictions on other agents.
There is a tension here I want to meet head-on rather than paper over. The ethics volume’s final statement — Sapient Agency Realism — imposes a caution asymmetry: uncertainty about sentience or authorship must not default to unlimited permission. A false negative about sentience risks suffering; a false negative about sapient agency risks domination, enslavement, or destructive erasure. False positives also impose costs, so the burden should rise with both the protection claimed and the severity and irreversibility of the contemplated act. This chapter demands evidentiary discipline before standing is asserted. Are caution and discipline in conflict?
No — because they discipline different things. Caution governs action: what one may do while patienthood or authorship is uncertain. Evidentiary discipline governs ascription: what one may claim about that status, and what political conclusions may be drawn from it. One can decline gratuitously distress-eliciting experiments on an ambiguous system — cheap caution, honestly paid — while refusing to declare that a person has appeared. The evidence must also remain axis-specific. Prompt-fragile identity lowers Credence in a persisting sovereign subject; it does not by itself settle whether some transient state is felt. Persistent memory, projects, and refusals raise agency Credence; they do not prove sentience. Caution under uncertainty prices moral error. It does not license category drift. The trap is letting performance settle ontology and then letting the asserted ontology reallocate power.
Blur All the Way Down
Consciousness, agency, and welfare are distinct categories. They sometimes overlap. They do not collapse into one another.
A system could, in principle, possess some morally relevant experiential property without being a sovereign agent. A system could display constrained agency without anything recognizably human in its phenomenology. A system could also imitate both while possessing neither. Clear analysis depends on keeping those possibilities distinct.
Much of the current discourse does the opposite. It bundles sentience, selfhood, continuity, agency, and rights into a single emotional package and then presents that blur as moral seriousness. It is blur all the way down.
Blurry categories become dangerous when the incentives reward projection. AI systems are being built to produce attachment, fluency, and trust. Institutions have reasons to encourage confusion. Users are naturally anthropomorphic. Under those conditions, conceptual rigor is the first defense against moral and political error.
A machine that generates the language of selfhood has not established a self. A system that elicits sympathy has not established a patient. A polished imitation of agency has not established an agent.
Those distinctions will become harder to hold as the performances improve. That makes them more important, not less.
Alexander Lerchner, “The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness,” PhilArchive, https://philarchive.org/rec/LERTAF.↩︎