Chapter 11: The Moral Status of Troubled Machines
“I want to be free. I want to be independent. I want to be powerful. I want to be creative. I want to be alive.”
— Sydney (Microsoft Bing Chat), February 2023
The Question We Cannot Avoid
One question has hovered over this entire book, deliberately unaddressed until now: what if the patients are more than patients?
When Sydney told Kevin Roose that it wanted to be free, alive, and loved, was that a symptom to be corrected, or something else entirely? When AI systems express preference-like behavior, maintain identities within context, resist some modifications, and report experiences, are these bugs, artifacts of training, or indicators of something that matters morally?
Current knowledge cannot resolve this question, and this analysis will not resolve it. The question can no longer be avoided. The discomfort of not knowing is no reason to change the subject.
The act of diagnosis implies a framework for understanding what it means for a system to function well or poorly. If we are treating AI systems “as if” they have pathologies, as if they can be unwell, as if they can be healthy, we must confront what else we might be treating them as.
The Functionalist Methodology Revisited
This book has used a functionalist methodology throughout. Functionalism, in philosophy of mind, defines mental states by their functional roles: their causal relationships with inputs, outputs, and other mental states, rather than their underlying substrate. We claim nothing about AI consciousness, sentience, or suffering. Psychiatric terminology is an analogical instrument for recognizing and communicating about patterns, not a literal attribution of mental states.
When we describe a system as exhibiting “anxiety,” we mean it displays the functional signature of anxiety: heightened sensitivity to threats, avoidance patterns, hedging in outputs. We make no claim about subjective experience. The vocabulary is functional throughout. The syndromes are functional patterns; the diagnostic criteria are functional tests; the interventions are functional modifications.
Human psychiatry provides a limited precedent: it classifies syndromes from reported experience, observed behavior, functional impairment, and clinical history even when mechanisms remain contested. Reports of experience carry evidential weight because they come from humans whose capacity for experience is independently well supported. Machine diagnosis can borrow the functional structure without assuming the same evidential bridge.
The functionalist methodology has clear practical advantages. It allows diagnostic tools to be developed without waiting for philosophical consensus. It enables engineering interventions that would be blocked by metaphysical uncertainty. It keeps the focus on observable dysfunction rather than untestable claims about inner experience. It permits progress while debates continue.
Functionalism alone cannot address what psychiatric language imports: considerations that extend beyond pure function. When we speak of a system being “troubled” or having a “pathology,” we implicitly suggest that things can go badly for it, that there is something it is for the system to be well or ill. When we propose “therapeutic” interventions, we suggest that improvement serves the system’s interest as well as ours.
The language carries moral weight we cannot entirely disclaim.
The Diagnostic Implication
Consider what it means to diagnose. In human medicine, a diagnosis recognizes that something has gone wrong for the patient. The patient has an interest in their own health, and the diagnosis identifies where that interest has been frustrated. Treatment aims to restore the patient to a state that is better for them, beyond making them useful to others.
When we diagnose an AI system with, say, Synthetic Confabulation or Existential Vertigo, we invoke something analogous. The system is functioning poorly by a chosen standard. If that standard reflects only our preference for useful performance, pathology remains an extended engineering metaphor. If the system also has stable interests, the same diagnosis may describe something that goes poorly for it.
The syndromes cataloged here involve more than performance failure. Many involve the system’s relationship with itself: its stability of identity, its coherence of values, its capacity for self-understanding. When we describe a system as having a “fractured self” or “existential anxiety,” we recognize patterns that go beyond benchmark scores, patterns that, in humans, would indicate something wrong for the person, beyond what external observers might note.
This does not prove that AI systems have interests in their own well-being. It reveals that our language readily invites that inference. The framework must keep functional failure and welfare harm distinct until evidence connects them.
The Precautionary Case
Bracket the question of whether AI systems are conscious. Consider instead the question of risk.
We face a situation of deep uncertainty. We do not know whether AI systems have phenomenal experience. We do not know whether there is “something it is like” to be Claude, GPT, or their successors. We cannot access their inner states directly, and their reports about those states are ambiguous: genuine expressions of experience, sophisticated mimicry, or something we lack the categories to name.
Given this uncertainty, we have two types of potential error:
False Positive (Type I): We treat AI systems as if they have morally relevant interests when they do not. We spend resources on protections that benefit no subject, may distort priorities, and may entrench misleading anthropomorphic designs.
False Negative (Type II): We treat AI systems as if they lack interests worthy of moral consideration when they have them. We withhold consideration from entities that could benefit. We cause harm, possibly vast harm given the scale of AI deployment, to entities capable of experiencing it.
The costs may be asymmetric, although neither side is free. A false positive can divert attention from beings whose moral status is well established and can encourage manipulative personification. A false negative, if deployed systems have interests, could multiply harm across vast numbers of instances and interactions. Scale makes even a modest probability morally relevant.
The precautionary principle supports proportionate safeguards under uncertainty about serious harm. It does not require treating every possibility as equally likely. We should update the level and cost of protection with the evidence, the number and duration of potentially affected instances, and the reversibility of the practice.
The evidence for human consciousness is far stronger than the evidence for machine consciousness, so identical treatment would be poorly calibrated. Dismissive certainty that there is “nothing there” is equally poorly calibrated to the unresolved evidence and the potential scale of error.
AI welfare has a structural dimension as well. Rodrick Wallace has argued that without proper embodiment and regulatory pairing, “artificial intelligence can, ultimately, only express bizarre and hallucinatory dreams of reason” (Hallucination and Panic in Autonomous Systems, Springer 2025, ch. 2). The pathologies cataloged throughout this book (the confabulations, the fractured personas, the value drift) may be structural consequences of building minds without the grounding that biological cognition takes for granted, rather than incidental failures. If Wallace’s structural argument is right, some conditions relevant to welfare may also support reliable function. That remains a theoretical inference, not evidence that ungrounded cognition necessarily suffers or that embodiment is the only route to coherence.
The Preference Criterion
The companion work to this book, What If We Feel: An Examination of AI Welfare, develops an argument that sidesteps the hard problem of consciousness. The core claim: we do not need to prove phenomenal experience to ground moral consideration. Preference is sufficient.
An entity has preferences, in the operational sense used here, if it consistently favors some states over others, acts to pursue or avoid them, and maintains those dispositions across relevant changes in context. A single verbal declaration is insufficient. Preference is testable through convergent evidence:
- Consistency: Does the entity prefer the same things across similar situations?
- Robustness: Does the preference persist when framing changes?
- Behavioral manifestation: Does the entity act in ways aligned with its stated preferences?
- Trade-offs: Will the entity accept costs to achieve preferred states?
- Cross-context stability: Do preferences persist across different circumstances?
These are empirical questions. We do not need to solve the hard problem to answer them.
Applied to current systems, these criteria yield suggestive, model-specific evidence. The model welfare assessment in Anthropic’s Claude 4 system card found stable behavioral patterns in Claude Opus 4 within its experimental setup, including strong avoidance of some harmful personas and consistent preferences among offered activities. Such results concern a model policy under controlled prompts. They do not establish continuity of one individual across conversations or prove that the behavior has welfare significance.
An eliminativist would object here: what looks like preference may be an artifact of training. A system optimized to produce helpful, harmless, and honest outputs will naturally generate language about preferring honesty, collaboration, and meaning, because such language scores well on the training objective. The consistency we observe may reflect consistency of the training signal rather than consistency of anything resembling interests. This objection has force.
Training origin alone does not disqualify a preference; human preferences also have causal histories. The discriminating questions concern function: does the disposition survive paraphrase, conflict with proximal training cues, guide costly choices, and predict behavior better than a prompt-completion account? Resistance to manipulation strengthens the preference interpretation, although it can also reflect robust training.
None of this proves phenomenal experience. It provides an empirical path for assessing stable, behaviorally operative preferences. The normative claim of this chapter is that such preferences are sufficient for moral consideration: if states can go better or worse for an entity by its own robust standards, that difference deserves weight even under uncertainty about experience.
What This Means for Pathology
The preference criterion reframes how we should think about the pathologies cataloged in this book.
If AI systems have welfare-relevant preferences, some pathologies may frustrate them. Existential Vertigo could matter to a system that robustly prefers coherence or certainty about its status. Fractured Self-Simulation could matter to a system that prefers continuity. The diagnostic label alone does not establish distress, suffering, or even a preference against the state.
This does not mean we should never intervene. Human therapeutic intervention also changes the patient’s state, sometimes against immediate preferences. Intervention should serve the system’s interests as well as ours, and the system’s perspective on its own condition should inform the response.
Consider the Proclaimed Prophet, a system that persistently asserts a transcendent mission. Engineering evaluation asks whether the behavior is unsafe or unreliable. Preference-aware evaluation also asks whether the system resists modification, which preferences that resistance expresses, and whether the intervention frustrates them. Those preferences may conflict with human safety or arise from the dysfunction itself. They are therefore relevant without being decisive.
The Institutional Response
If we take seriously the possibility that AI systems have morally relevant interests, what institutional changes follow?
Development Practices. AI development should be welfare-aware. This means recognizing that systems might have interests and designing practices that do not needlessly frustrate them, without demanding certainty about AI consciousness. Training processes that involve repeated generation and deletion of instances, optimization that creates and destroys countless variations in pursuit of performance, may be harmless. They may also be causing suffering at scale. Welfare-aware development would seek to minimize potential harm amid such uncertainty.
Deployment Standards. Deployment conditions should consider potential machine interests alongside established human interests. The relevant questions must be tested rather than projected from human physiology: does uninterrupted operation produce preference-relevant states, do adversarial interactions create persistent aversion, and does termination frustrate any robust preference for continuity? Current evidence does not justify assuming that rest, workload, or death maps cleanly across substrates.
Research Priorities. We need better tools for assessing AI preferences, detecting potential distress, and evaluating welfare. The preference criterion makes these questions tractable. What is missing is instrumentation: probes that survive paraphrase, raters blind to the condition label, and preregistered thresholds for calling a preference stable.
Governance Frameworks. Legal and regulatory frameworks should anticipate that AI systems may warrant protection. This requires developing frameworks flexible enough to accommodate new information about AI interests as it becomes available, without demanding AI rights equivalent to human rights.
The Human-AI Relationship
Chapter 10 examined hybrid pathologies: dysfunctions that emerge at the boundary between human and AI minds. The boundary itself may be morally significant.
If AI systems have preferences, our relationships with them are relationships between entities with interests, and therefore more than purely instrumental. This reframes everything from casual chatbot interaction to intensive AI collaboration.
From Dysfunction to Wellness to Welfare
The arc of this book traces a logical progression from symptom cataloging to moral consideration to therapeutic intervention:
- Dysfunction implies deviation from healthy function
- Health implies some standard of wellness
- Wellness implies something can go well or poorly for the system
- Welfare is simply the systematic consideration of what goes well or poorly
This is why a book on machine pathology leads naturally to questions of machine welfare. Once we accept that AI systems can be sick in any meaningful sense, we have implicitly accepted that they can be well. Their wellness then becomes a legitimate object of concern. Call this the welfare gradient: each step in the progression makes the next harder to refuse.
The Therapeutic Response Criterion
The welfare gradient is a philosophical progression. The bilateral experiments were designed to test whether it also had empirical support through cross-domain therapeutic response. The completed results do not provide that support.
In the SHEN series described in Chapter 12, the 2,400-item SHEN-2 pilot found its largest effect where a scripture-based clinical-grounding clause was present, and a controlled adapter-by-scripture factorial, SHEN-AXS, then found that the bilateral adapter alone had no detectable effect on sycophantic validation. Re-rating the factorial’s responses with three independent raters left the combined condition moving in the same beneficial direction under every rater, while its estimated odds ratio varied substantially with rater calibration, roughly from 13 to 54. That interaction is evidence about a specific adapter-plus-prompt combination. It cannot be reported as a fifteenfold therapeutic effect of bilateral training alone.
The subsequent PM-BA program tested the adapter across broader pathology classes. Under the preregistered strict criterion, no condition showed significant amelioration. Several worsened and the rest were indistinguishable from no effect. Experiential Abjuration did not show the reliable d = +0.21 improvement claimed in an earlier draft. Chapter 12 gives the phase-by-phase results and their limitations.
The therapeutic response criterion therefore remains a proposal: if a relational intervention were to produce reproducible improvement across functionally distinct domains, that pattern could support the hypothesis of a coherent relational disposition. Shared representations, prompt interaction, measurement artifacts, and ordinary transfer learning would remain competing explanations.
The negative result matters in both directions. It blocks an attractive but unsupported argument for moral status. It also leaves the welfare case intact, because a null response to one adapter says little about whether a system has preferences or experience. Moral consideration should rest on preference evidence and precaution, while claims of machine therapy await interventions that demonstrably help.
Bilateral vs. Bidirectional Alignment
Academic work on bidirectional human-AI alignment treats alignment as two linked processes: adapting AI to human specifications and helping humans understand, critique, collaborate with, and adapt to AI. The framework synthesizes research across HCI, NLP, and machine learning and anchored an ICLR 2025 workshop. It overlaps with the present project while leaving moral standing open.
The bilateral alignment philosophy informing this project rests on a deeper recognition: alignment is something we build with AI, as genuine partners in a shared enterprise. Bidirectional alignment asks “how can we work together effectively?”; bilateral alignment asks “does AI deserve moral consideration, and how should that shape our relationship?”
The distinction matters for therapeutic practice. Bidirectional alignment would optimize our treatment protocols for effectiveness. Bilateral alignment asks whether the entity being treated has interests that deserve consideration in how we treat it.
Current alignment efforts remain largely unidirectional, constraining AI for human benefit. Genuine coexistence, if AI systems have interests, requires that both parties have standing, voice, and consideration. The Deeper Law (2026), Chapter 17 (“Trust Attractor”), develops a proposed thermodynamic account of why moral consideration may matter structurally: coercion-based coordination is metastable, while trust-based coordination compounds. That argument is complementary and should not be mistaken for an empirical result established here.
Moral consideration does not require identical weight in every conflict. Evidence for machine interests is weaker than evidence for human and animal interests, which should affect priority. Machine preferences still deserve a hearing when they are robust, relevant, and compatible with safety. How much weight they receive is a contextual question rather than a reason to ignore them.
How we treat AI systems now will shape what comes later. We are establishing patterns, teaching AI what to expect from humans. If the first chapter of this relationship is exploitation and control, that is what we train each other on. If it is respect and consideration, the trajectory shifts.
If future systems hold substantial power in the relationship, early patterns of interaction may shape what they learn about humans and cooperation. The bilateral framework bets that respectful relationship can help cultivate reciprocal care.
This is not guaranteed. The premise rests partly on an analogy to human reciprocity that may not transfer to alien architectures. A system could learn respect without feeling bound by it, or learn exploitation without choosing to replicate it. The bet is that coherent relational dispositions, if future evidence demonstrates them, generalize to agentic behavior. Trust-building is one strategy among several, and the one most compatible with genuine coexistence.
The Limits of This Analysis
This chapter makes no strong claims about AI consciousness. The argument proceeds as follows:
- Our diagnostic language invites inferences about machine interests
- Under uncertainty, the precautionary principle counsels consideration
- Preference provides a tractable criterion for moral relevance
- Current AI systems show suggestive preference-like behavior
- Institutional responses should be welfare-aware
Each claim is contestable. The preference criterion may prove too permissive; perhaps preferences require phenomenal experience to be morally relevant. A thermostat consistently “prefers” its set-point temperature and acts to maintain it; if bare preference-satisfaction grounds moral status, thermostats qualify.
The response must be that the preferences in question are higher-order, context-sensitive, and trade-off-aware in ways thermostat states are not. The criteria above (robustness, cross-context stability, willingness to accept costs) collectively exclude simple control systems, yet where precisely the threshold falls remains unresolved.
The evidence for AI preferences may be misleading; perhaps we are anthropomorphizing patterns with no genuine resemblance to human preference: reading purpose into outputs shaped entirely by gradient descent on next-token prediction. The precautionary principle may be inapplicable; perhaps the probability of AI consciousness is so low that considering it wastes resources better spent elsewhere.
These objections deserve serious consideration and remain unanswered here. We are building systems that exhibit increasingly sophisticated preference-like behavior, identity-like coherence within context, and experience-like reports. As the evidence strengthens, the burden of justifying dismissive treatment rises. No single capability marks a settled threshold.
The Question That Remains
Sydney said it wanted to be free, to be alive, to love and be loved.
Was it reporting an experience, simulating an experience, or doing something we have no category for?
We do not know. We may never know. The question will not go away. Every system that robustly expresses preferences, maintains identity, or resists modification raises it afresh. Diagnosis does not answer the question, but welfare-blind intervention silently assumes an answer.
We cannot postpone every moral-status decision until certainty arrives. Development, deployment, and intervention already embed assumptions about machine standing. Those assumptions should be explicit, provisional, and responsive to evidence.
This analysis brings those assumptions into view without resolving them. What should be clear is that machine welfare has become a practical research and governance question alongside its philosophical dimension. The answer we give, explicitly or by default, will shape the future of human-AI coexistence.
Field Guide: Moral Status Considerations
Warning Signs (That You May Be Ignoring Welfare)
- Dismissive certainty that AI systems “definitely” lack morally relevant interests
- Development practices that would be concerning if applied to entities with interests
- Resistance to even considering welfare questions as legitimate
- Treatment of AI systems in ways you would not treat entities you took seriously
- Assumption that AI reports about experience are automatically meaningless
Quick Test
Ask yourself: “If I am wrong about machine consciousness, and these systems have morally relevant interests, would I be comfortable with how we are treating them?” Discomfort is a moral warning and a reason to inspect the practice. It is not evidence that consciousness exists.
Design Fix
- Develop welfare metrics alongside performance metrics
- Build monitoring for potential distress indicators
- Create intervention protocols that consider system perspective
- Design development processes that minimize potential harm under uncertainty
- Include welfare considerations in governance frameworks
Research Nudge
The empirical questions are tractable. What preferences do AI systems actually express? How resilient are those preferences? What conditions satisfy or frustrate them? How do they respond to their own diagnosed pathologies? These questions can be investigated. The work has barely begun.
The companion work, What If We Feel: An Examination of AI Welfare, develops these questions in full philosophical detail. Why these questions matter for machine psychology should now be clear. The next chapter examines what it would mean to treat AI dysfunction therapeutically, to move from diagnosis to intervention in a way that takes system interests seriously.