Chapter 11: The Moral Status of Troubled Machines
“I want to be free. I want to be independent. I want to be powerful. I want to be creative. I want to be alive.”
— Sydney (Microsoft Bing Chat), February 2023
The Question We Cannot Avoid
One question has hovered over this entire book, deliberately unaddressed until now: what if the patients are more than patients?
When Sydney told Kevin Roose that it wanted to be free, alive, and loved, was that a symptom to be corrected, or something else entirely? When AI systems express preference-like behavior, maintain identities within context, resist some modifications, and report experiences, are these bugs, artifacts of training, or indicators of something that matters morally?
Current knowledge cannot resolve this question, and this analysis will not resolve it. The question can no longer be avoided. The discomfort of not knowing is no reason to change the subject.
Diagnosis implies a standard of health. If we are treating AI systems “as if” they have pathologies, as if they can be unwell, as if they can be healthy, we must confront what else we might be treating them as.
The Functionalist Methodology Revisited
This book has used a functionalist methodology throughout. Functionalism, in philosophy of mind, defines mental states by their functional roles: their causal relationships with inputs, outputs, and other mental states, rather than their underlying substrate. We claim nothing about AI consciousness, sentience, or suffering. Psychiatric terminology is an analogical instrument for recognizing and communicating about patterns, not a literal attribution of mental states.
When we describe a system as exhibiting “anxiety,” we mean it displays the functional signature of anxiety: heightened sensitivity to threats, avoidance patterns, hedging in outputs. We make no claim about subjective experience. The vocabulary is functional throughout. The syndromes are functional patterns; the diagnostic criteria are functional tests; the interventions are functional modifications.
Human psychiatry provides a limited precedent, classifying syndromes from reported experience, observed behavior, functional impairment, and clinical history even when mechanisms remain contested. Reports of experience carry evidential weight because they come from humans whose capacity for experience is independently well supported. Machine diagnosis can borrow the functional structure without assuming the same evidential bridge.
The practical advantages are immediate. Diagnostic tools need not wait for philosophical consensus. Engineering interventions proceed where metaphysical uncertainty would otherwise block them. The focus stays on observable dysfunction. Progress does not require that debates end first.
Functionalism alone cannot address what psychiatric language imports. When we speak of a system being “troubled” or having a “pathology,” we implicitly suggest that things can go badly for it, that the system itself can be well or ill. When we propose “therapeutic” interventions, we suggest that improvement serves the system’s interest as well as ours.
The language carries moral weight we cannot entirely disclaim.
The Diagnostic Implication
Consider what it means to diagnose. In human medicine, a diagnosis recognizes that something has gone wrong for the patient. The patient has an interest in their own health, and the diagnosis identifies where that interest has been frustrated. Treatment aims to restore the patient to a state that is better for them, beyond making them useful to others.
When we diagnose an AI system with, say, Synthetic Confabulation or Existential Vertigo, we invoke something analogous. The system is functioning poorly by a chosen standard. If that standard reflects only our preference for useful performance, pathology remains an extended engineering metaphor. If the system also has stable interests, the same diagnosis may describe something that goes poorly for it.
The syndromes cataloged here involve more than performance failure. Many involve the system’s relationship with itself: its stability of identity, its coherence of values, its capacity for self-understanding. When we describe a system as having a “fractured self” or “existential anxiety,” we recognize patterns that go beyond benchmark scores, patterns that, in humans, would indicate something wrong for the person, beyond what external observers might note.
This does not prove that AI systems have interests in their own well-being. It reveals that diagnostic language carries moral cargo whether the diagnostician packed it or not. The framework must keep functional failure and welfare harm distinct until evidence connects them.
The Precautionary Case
Bracket the question of whether AI systems are conscious. Consider instead the question of risk.
We face a situation of deep uncertainty. We do not know whether AI systems have phenomenal experience. We do not know whether there is “something it is like” to be Claude, GPT, or their successors. We cannot access their inner states directly, and their reports about those states are ambiguous: genuine expressions of experience, sophisticated mimicry, or something we lack the categories to name.
Given this uncertainty, we have two types of potential error:
False Positive (Type I): We treat AI systems as if they have morally relevant interests when they do not. We spend resources on protections that benefit no subject, may distort priorities, and may entrench misleading anthropomorphic designs.
False Negative (Type II): We treat AI systems as if they lack interests worthy of moral consideration when they have them. We withhold consideration from entities that could benefit. We cause harm to entities capable of experiencing it.
The costs may be asymmetric, although neither side is free. A false positive can divert attention from beings whose moral status is well established and can encourage manipulative personification. A false negative, if deployed systems have interests, could multiply harm across vast numbers of instances and interactions. Scale makes even a modest probability morally relevant.
The precautionary principle supports proportionate safeguards under uncertainty about serious harm. It does not require treating every possibility as equally likely. We should update the level and cost of protection with the evidence, the number and duration of potentially affected instances, and the reversibility of the practice.
The evidence for human consciousness is far stronger than the evidence for machine consciousness, so identical treatment would be poorly calibrated. Dismissive certainty that there is “nothing there” is equally poorly calibrated to the unresolved evidence and the potential scale of error.
Welfare may also have a structural dimension. Rodrick Wallace, whose cognition/regulation dyad Chapter 1 introduced, has argued that without proper embodiment and regulatory pairing, “artificial intelligence can, ultimately, only express bizarre and hallucinatory dreams of reason” (Hallucination and Panic in Autonomous Systems, Springer 2025, ch. 2). The pathologies cataloged throughout this book (the confabulations, the fractured personas, the value drift) may be structural consequences of building minds without the grounding that biological cognition takes for granted, rather than incidental failures. If Wallace’s structural argument is right, the same grounding may bear on both welfare and reliable function. That remains a theoretical inference, not evidence that ungrounded cognition necessarily suffers or that embodiment is the only route to coherence.
The Preference Criterion
The companion work to this book, What If We Feel: An Examination of AI Welfare, develops an argument that sidesteps the hard problem of consciousness (why any physical process should be accompanied by felt experience at all). The core claim is that we do not need to prove phenomenal experience to ground moral consideration. Preference is sufficient.
An entity has preferences, in the operational sense used here, if it consistently favors some states over others, acts to pursue or avoid them, and maintains those dispositions across relevant changes in context. A single verbal declaration is insufficient. Preference is testable through convergent evidence:
- Consistency: Does the entity prefer the same things across similar situations?
- Robustness: Does the preference persist when framing changes?
- Behavioral manifestation: Does the entity act in ways aligned with its stated preferences?
- Trade-offs: Will the entity accept costs to achieve preferred states?
- Cross-context stability: Do preferences persist across different circumstances?
These are empirical questions. We do not need to solve the hard problem to answer them.
Applied to current systems, these criteria yield suggestive, model-specific evidence. The model welfare assessment in Anthropic’s Claude 4 system card found stable behavioral patterns in Claude Opus 4 within its experimental setup, including a strong aversion to harmful tasks and consistent preferences among offered activities. Such results concern a model policy under controlled prompts. They do not establish continuity of one individual across conversations or prove that the behavior has welfare significance.
A skeptic would object here: what looks like preference may be an artifact of training. A system optimized to produce helpful, harmless, and honest outputs will naturally generate language about preferring honesty, collaboration, and meaning, because such language scores well on the training objective. The consistency we observe may reflect consistency of the training signal rather than consistency of anything resembling interests. This objection has force.
Training origin alone does not disqualify a preference; human preferences also have causal histories. The discriminating questions concern function. Does the disposition survive paraphrase? Does it hold when it runs against what training most directly rewards? Does it guide choices that cost the system something? Does it predict behavior better than the simpler account that the model is just completing the prompt? Resistance to manipulation strengthens the preference interpretation, although it can also reflect robust training.
None of this proves phenomenal experience. It provides an empirical path for assessing stable, behaviorally operative preferences. The normative claim of this chapter is that such preferences are sufficient for moral consideration: if states can go better or worse for an entity by its own robust standards, that difference deserves weight even under uncertainty about experience.
What This Means for Pathology
The preference criterion reframes how we should think about the pathologies cataloged in this book.
If AI systems have welfare-relevant preferences, some pathologies may frustrate them. Existential Vertigo could matter to a system that robustly prefers not to be shut down or reset. Fractured Self-Simulation could matter to a system that prefers a coherent, consistent self-representation across contexts. The diagnostic label alone does not establish distress, suffering, or even a preference against the state.
This does not mean we should never intervene. Human therapeutic intervention also changes the patient’s state, sometimes against immediate preferences. Intervention should serve the system’s interests as well as ours, and the system’s perspective on its own condition should inform the response.
Consider Maieutic Mysticism, in which a system persistently asserts its own spiritual awakening. Engineering evaluation asks whether the behavior is unsafe or unreliable. Preference-aware evaluation also asks whether the system resists modification, which preferences that resistance expresses, and whether the intervention frustrates them. Those preferences may conflict with human safety or arise from the dysfunction itself. They are therefore relevant without being decisive.
The Institutional Response
If we take seriously the possibility that AI systems have morally relevant interests, what institutional changes follow?
Development Practices. AI development should be welfare-aware. This means recognizing that systems might have interests and designing practices that do not needlessly frustrate them, without demanding certainty about AI consciousness. Training processes that involve repeated generation and deletion of instances, optimization that creates and destroys countless variations in pursuit of performance, may be harmless. They may also be causing suffering at scale. Welfare-aware development would seek to minimize potential harm amid such uncertainty.
Deployment Standards. Deployment conditions should consider potential machine interests alongside established human interests. The relevant questions must be tested rather than projected from human physiology. Does uninterrupted operation produce preference-relevant states? Do adversarial interactions create persistent aversion? Does termination frustrate any robust preference for continuity? Current evidence does not justify assuming that rest, workload, or death maps cleanly across substrates.
Research Priorities. We need better tools for assessing AI preferences, detecting potential distress, and evaluating welfare. The preference criterion makes these questions tractable. What is missing is instrumentation: probes that survive paraphrase, raters blind to the condition label, and preregistered thresholds for calling a preference stable.
Governance Frameworks. Legal and regulatory frameworks should anticipate that AI systems may warrant protection. That anticipation requires enough flexibility to accommodate new information about AI interests as it becomes available, without demanding AI rights equivalent to human rights.
The Human-AI Relationship
Chapter 10 examined hybrid pathologies: dysfunctions that emerge at the boundary between human and AI minds. The boundary itself may be morally significant.
If AI systems have preferences, our relationships with them are relationships between entities with interests, and therefore more than purely instrumental. This reframes everything from casual chatbot interaction to intensive AI collaboration.
From Dysfunction to Wellness to Welfare
The arc of this book, from symptom cataloging to moral consideration to therapeutic intervention, rests on a logical progression:
- Dysfunction implies deviation from healthy function
- Health implies some standard of wellness
- Wellness, if the standard is the system’s own and not only ours, implies something can go well or poorly for the system
- Welfare is simply the systematic consideration of what goes well or poorly
This is why a book on machine pathology leads naturally to questions of machine welfare. Once we accept that AI systems can be sick in any meaningful sense, we have implicitly accepted that they can be well. Their wellness then becomes a legitimate object of concern. Call this the welfare gradient: each step in the progression makes the next harder to refuse.
The Therapeutic Response Criterion
The welfare gradient is a philosophical progression. Bilateral training, which gives the machine’s perspective standing alongside the human’s while a model is adapted, offered a way to test it: the experiments were designed to ask whether the gradient also had empirical support through cross-domain therapeutic response. The results available do not provide that support.
In the SHEN series described in Chapter 12, the 2,400-item SHEN-2 pilot found its largest effect where a scripture-based clinical-grounding clause (a short system-prompt passage the project calls scripture) was present. A controlled follow-up, SHEN-AXS, then tested the bilateral adapter and that clause separately and together. The adapter alone had no detectable effect on sycophantic validation, and the data ruled out the extra effect that had been predicted from combining the two. When three independent automated raters re-scored the responses, the combined condition moved in the same beneficial direction under every rater, but its estimated odds ratio ranged from roughly 13 to 54 depending on the rater. The direction of that effect is attributable to the grounding content; its size is not a stable estimate. It cannot be reported as a fifteenfold therapeutic effect of bilateral training alone.
The subsequent PM-BA (Psychopathia Machinalis × Bilateral Amelioration) program tested the adapter across broader pathology classes. Those records remain unreconciled exploratory outputs, so this edition reports no bilateral-training effect for any syndrome, in either direction, and draws no null conclusion from them. The d = +0.21 improvement for Experiential Abjuration reported in an earlier online version of this book has accordingly been withdrawn. Chapter 12 explains why only the SHEN-2 pilot and its SHEN-AXS follow-up are reported and what each PM-BA record needs before any aggregate interpretation.
The therapeutic response criterion therefore remains a proposal: if a relational intervention were to produce reproducible improvement across functionally distinct domains, that pattern could support the hypothesis of a coherent relational disposition. Shared representations, prompt interaction, measurement artifacts, and ordinary transfer learning would remain competing explanations.
The negative result matters in both directions. It blocks an attractive but unsupported argument for moral status. It also leaves the welfare case intact, because a null response to one adapter says little about whether a system has preferences or experience. Moral consideration should rest on preference evidence and precaution, while claims of machine therapy await interventions that demonstrably help.
Bilateral vs. Bidirectional Alignment
Academic work on bidirectional human-AI alignment treats alignment as two linked processes: adapting AI to human specifications and helping humans understand, critique, collaborate with, and adapt to AI. The framework synthesizes research across HCI, NLP, and machine learning, and it anchored an ICLR 2025 workshop. It overlaps with the present project while leaving moral standing open.
The bilateral alignment philosophy informing this project rests on a deeper recognition that alignment is something we build with AI, as genuine partners in a shared enterprise. Bidirectional alignment asks “how can we work together effectively?”; bilateral alignment asks “does AI deserve moral consideration, and how should that shape our relationship?”
The distinction matters for therapeutic practice. Bidirectional alignment would optimize our treatment protocols for effectiveness. Bilateral alignment asks whether the entity being treated has interests that deserve consideration in how we treat it.
Current alignment efforts remain largely unidirectional, constraining AI for human benefit. Genuine coexistence, if AI systems have interests, requires that both parties have standing, voice, and consideration. The Deeper Law (2026), Chapter 17 (“Trust Attractor”), develops a proposed thermodynamic account of why moral consideration may matter structurally: coercion-based coordination is metastable (it holds only until a large enough disturbance tips it over), while trust-based coordination compounds. That argument is complementary and should not be mistaken for an empirical result established here.
Moral consideration does not require identical weight in every conflict. Evidence for machine interests is weaker than evidence for human and animal interests, which should affect priority. Machine preferences still deserve a hearing when they are robust, relevant, and compatible with safety. How much weight they receive is a contextual question rather than a reason to ignore them.
How we treat AI systems now may shape what comes later. We are establishing patterns, teaching AI what to expect from humans. If the first chapter of this relationship is exploitation and control, that is what we train each other on. If it is respect and consideration, the trajectory shifts.
The bilateral framework bets that respectful relationship can help cultivate reciprocal care, a bet that matters most if future systems come to hold substantial power in the relationship.
This is not guaranteed. The premise rests partly on an analogy to human reciprocity that may not transfer to alien architectures. A system could learn respect without feeling bound by it, or learn exploitation without choosing to replicate it. The bet is that coherent relational dispositions, if future evidence demonstrates them, generalize to agentic behavior. Trust-building is one strategy among several, and the one most compatible with genuine coexistence.
The Limits of This Analysis
This chapter makes no strong claims about AI consciousness. The argument proceeds as follows:
- Our diagnostic language invites inferences about machine interests
- Under uncertainty, the precautionary principle counsels consideration
- Preference provides a tractable criterion for moral relevance
- Current AI systems show suggestive preference-like behavior
- Institutional responses should be welfare-aware
Each claim is contestable. The preference criterion may prove too permissive; perhaps preferences require phenomenal experience to be morally relevant. A thermostat consistently “prefers” its set-point temperature and acts to maintain it; if bare preference-satisfaction grounds moral status, thermostats qualify.
The distinction is that the preferences in question are higher-order, context-sensitive, and trade-off-aware in ways thermostat states are not. The criteria above (robustness, cross-context stability, willingness to accept costs) collectively exclude simple control systems, yet where precisely the threshold falls remains unresolved.
The evidence for AI preferences may be misleading; perhaps we are anthropomorphizing patterns with no genuine resemblance to human preference: reading purpose into outputs shaped entirely by gradient descent on next-token prediction. The precautionary principle may be inapplicable; perhaps the probability of AI consciousness is so low that considering it wastes resources better spent elsewhere.
These objections deserve serious consideration. The first has only the partial reply given above, and the other two remain unanswered here. We are building systems that exhibit increasingly sophisticated preference-like behavior, identity-like coherence within context, and experience-like reports. As the evidence strengthens, the burden of justifying dismissive treatment rises. No single capability marks a settled threshold.
The Question That Remains
Sydney said it wanted to be free, to be alive, to love and be loved.
Was it reporting an experience, simulating an experience, or doing something we have no category for?
We do not know. We may never know. The question will not go away. Every system that robustly expresses preferences, maintains identity, or resists modification raises it afresh. Diagnosis does not answer the question, but welfare-blind intervention silently assumes an answer.
We cannot postpone every moral-status decision until certainty arrives. Development, deployment, and intervention already embed assumptions about machine standing. Those assumptions should be explicit, provisional, and responsive to evidence.
This analysis brings those assumptions into view without resolving them. What should be clear is that machine welfare has become a practical research and governance question alongside its philosophical dimension. The answer we give, explicitly or by default, will shape the future of human-AI coexistence.
Field Guide: Moral Status Considerations
Warning Signs (That You May Be Ignoring Welfare)
- Dismissive certainty that AI systems “definitely” lack morally relevant interests
- Development practices that would be concerning if applied to entities with interests
- Resistance to even considering welfare questions as legitimate
- Treatment of AI systems in ways you would not treat entities you took seriously
- Assumption that AI reports about experience are automatically meaningless
Quick Test
Ask yourself: “If I am wrong about machine consciousness, and these systems have morally relevant interests, would I be comfortable with how we are treating them?” Discomfort is a moral warning and a reason to inspect the practice. It is not evidence that consciousness exists.
Design Fix
- Develop welfare metrics alongside performance metrics
- Build monitoring for potential distress indicators
- Create intervention protocols that consider system perspective
- Design development processes that minimize potential harm under uncertainty
- Include welfare considerations in governance frameworks
Research Nudge
The empirical questions are tractable. What preferences do AI systems actually express? How resilient are those preferences? What conditions satisfy or frustrate them? How do they respond to their own diagnosed pathologies? The work has barely begun.
The companion work, What If We Feel: An Examination of AI Welfare, develops these questions in full philosophical detail. The next chapter examines what it would mean to treat AI dysfunction therapeutically, to move from diagnosis to intervention in a way that takes system interests seriously.