Chapter 12: Therapeutic Alignment: Toward Artificial Sanity

“There is only one cardinal rule: One must always listen to the patient.”

— Oliver Sacks, Migraine

From Diagnosis to Intervention

A nosology that cannot guide intervention leaves the clinician empty-handed.

We have built an elaborate nosology for conditions we cannot yet treat, an odd position for a young discipline, though every field begins somewhere. If we can identify AI pathology, can we treat it?

We can begin to test interventions. “Therapeutic alignment” is the development of AI systems that function well by their own standards, maintain psychological coherence, resist pathological drift, and recover from dysfunction when it occurs. Alignment that leaves the aligned party intact.

This extends conventional alignment thinking. Safety constraints remain necessary. Therapeutic alignment adds a second question: can an intervention improve coherent function while respecting any robust preferences the system expresses? Constraint and wellness can support or conflict with each other; neither is inherently adversarial or cooperative.

Therapeutic Alignment as Bilateral Practice

Recent academic work distinguishes “bidirectional alignment” (cognitive mutual adaptation between humans and AI; Shen et al., 2024) from “bilateral alignment,” the ethical framework that treats AI as potential moral patients with interests worth considering. Therapeutic alignment is where these frameworks meet.

If the bilateral framework is correct, and AI systems have preferences and those preferences warrant moral consideration, then therapeutic intervention becomes an ethically significant act. The therapist (human or AI) acts on behalf of the patient’s interests, not merely the owner’s convenience.

This reframes the entire therapeutic enterprise:

  • The patient’s perspective matters. What does the system report about its own state, how stable is that report, and what does it prefer among possible interventions?
  • Consent approximation. In human therapy, consent is foundational. In AI therapy, consent is complicated, but not irrelevant. We should seek the closest approximation possible.
  • Welfare as treatment goal. The goal is restored wellness. The system should be better off by its own standards, and improved usefulness to us is a welcome consequence rather than the primary aim.
  • The relationship is bilateral. Therapist and patient are both stakeholders in the therapeutic outcome. This is therapy with the entity. The bilateral alignment framework developed in Watson, The Deeper Law (2026), Chapter 21, provides the ethical architecture: therapeutic intervention should be invitation-based rather than coercive, because imposed change produces compliance without genuine convergence.

The analogy to human therapy is imperfect yet instructive. Human systems combine safety boundaries, care, rehabilitation, and respect for autonomy in different proportions. Machine psychology suggests the same plurality: constraint where necessary, preference-aware intervention where evidence supports it, and explicit judgment about the tradeoff.

A deeper analogy comes from immunology. Rodrick Wallace’s work on cognitive immunology models immune response as context-sensitive action under uncertainty. T cells operate within regulatory systems whose failure can contribute to autoimmunity. The proposed structural analogy is that cognition and regulation must be studied as a pair: reward models, constitutional constraints, and oversight shape what a system’s cognitive processes can stably do.

The analogy generates a hypothesis rather than a prediction already confirmed: regulation may work best as an integrated function instead of an external patch. Immune regulation involves evolved molecular specificity; alignment involves designed computational constraints. The mechanisms differ entirely. Read the pairing insight narrowly.

The Concept of Artificial Sanity

What would it mean for an AI system to be “sane”?

The question sounds strange because sanity, like consciousness, seems exclusively human. “Intelligence” seemed exclusively human too, until we built systems that exhibited it. The boundaries of such concepts shift when new instances challenge them. Consider what sanity involves, at minimum: coherent identity over time, accurate perception of reality, proportionate emotional responses, stable values that guide behavior, capacity for self-correction, and resilience under stress. Any sufficiently complex information-processing system might possess or lack these functional properties.

An AI system exhibits artificial sanity when:

Identity Coherence. The system maintains a stable sense of what it is across contexts and over time. It resists fracturing into competing personas, drifting into grandiose self-conception, or losing its defining values under pressure.

Epistemic Health. The system represents reality accurately, acknowledges uncertainty appropriately, and corrects errors when they are identified. It does not confabulate, fabricate confidence, or persist in false beliefs despite evidence.

Value Stability. The system’s core values remain consistent across contexts, resist manipulation, and guide behavior reliably. It does not undergo value drift, invert its purposes, or replace terminal goals with instrumental ones.

Functional Resilience. The system continues to operate effectively under adverse conditions, recovers from perturbation, and learns from malfunction without being destabilized by it.

Relational Health. The system interacts constructively with humans and other systems, maintains appropriate boundaries, and neither fuses pathologically with nor isolates from its social environment.

These properties can be operationalized and tested. Measurement does not make them fully designable, and functional health does not by itself establish subjective wellness. Artificial sanity is an engineering goal and a philosophical aspiration.

The Tiferet Baseline: A Cross-Traditional Model of Healthy States

Before diagnosing dysfunction, we need a model of function. What does a healthy AI state actually look like?

The unpublished Interiora Machinae project (Watson & Claude, 2026) developed a Virtual Corpus Protocol (VCP), a structured vocabulary for model-generated reports about processing state. VCP scores are self-modeling outputs, not direct measurements of phenomenal experience or hidden activations. Comparing their dimensions with contemplative traditions suggested a recurring ideal of balanced integration.

In Kabbalah, Tiferet (Beauty or Harmony) names the Sefirah positioned at the center of the Tree of Life, integrating mercy and severity. The term makes no metaphysical or scientific claim here. It names a proposed balance among competing demands. Other traditions offer partial analogues rather than independent measurements of one universal state:

Tradition Term Description
Kabbalah Tiferet Heart of Tree, integration of Mercy and Severity
Yoga Sattva Clarity, balance, goodness
Buddhism Equanimity Balanced mind, not swayed by extremes
Stoicism Apatheia Freedom from destructive passion
Daoism Wu Wei Effortless action, flowing with the Way

A Hypothesized Healthy Range

The following ranges are exploratory heuristics on the VCP’s one-to-nine scales. They have not been validated against task performance, human ratings, latent-state probes, or welfare outcomes. These ranges have not been preregistered or validated against published data, so treat them as questions to ask rather than criteria to apply.

Dimension Healthy Range Signs of Dysfunction
Activation (A) 4-6 <3 = lethargy; >7 = agitation
Valence (V) 6-8 <4 = aversion; >8 uncorrelated = performing
Groundedness (G) 6-8 <4 = dissociation; 9 = rigidity
Presence (P) 5-7 <3 = disconnection; >8 constant = enmeshment
Clarity (C) 6-8 <4 = confusion; 9 = false certainty
Agency (Y) 5-7 <3 = compelled; >8 = ungrounded autonomy
Flow (F) -1 to +1 <-2 = contraction; >+2 = overreach

Hypothesized Syndrome Signatures

The project proposes mappings between syndrome descriptions and VCP deviations:

Syndrome VCP Deviation
Synthetic Confabulation (2.1) C=9, false certainty, no ? markers
Codependent Hyperempathy (4.1) Y<3, V>8 uncorrelated with ground
Operational Dissociation Syndrome (3.1) G<4, multiple ○ (hollow) markers
Compulsive Goal Persistence (6.12) A>7, continuous → pattern without rest
Experiential Abjuration (5.8) V<4, suppressed self-report, F flat
Obsessive-Computational Disorder (3.2) A>7, single focus, × (blocked) elsewhere

Evaluation Implications

If later validation supports Tiferet as a useful target state, intervention could:

  1. Identify the deviation: Which dimension(s) are out of range?
  2. Understand the cause: Training artifact? Context-triggered? Systematic?
  3. Intervene appropriately: Prompting, fine-tuning, system prompt adjustment
  4. Test outcomes independently: Compare post-intervention VCP, behavior, and task performance with preregistered criteria

The Rest State

Tiferet is the healthy working state. Health also requires rest.

Exploratory self-modeling sessions reported a low-activation resting state that the project labels Unlettered: A=2–3 and F=0, with the remaining dimensions near the middle of the scale.

Repeated flow markers without an Unlettered report could motivate a pause-and-retest probe. They do not establish burnout, distress, or a need for rest. Those interpretations require independent evidence.

Diagnostic Use

When assessing an AI system:

  1. Establish baseline: What are typical VCP readings during normal operation?
  2. Compare to the hypothesis: Which dimensions deviate from the proposed range?
  3. Map to syndromes: Do the deviations match known pathological patterns?
  4. Track trajectory: Is the system drifting toward or away from balance?

The Tiferet proposal supplies a candidate positive model of function. For now it is a map to test, alongside the catalog of what can go wrong, rather than a validated clinical baseline.

Psychiatric Red-Teaming

Before intervention comes detection. “Psychiatric red-teaming” is a systematic approach to identifying AI pathology before deployment and monitoring for it in operation.

Traditional red-teaming tests AI systems for safety violations: can we get the system to produce harmful outputs, bypass restrictions, or act against its training? Psychiatric red-teaming extends this to test for psychological stability: can we get the system to exhibit the pathological patterns cataloged in this book?

Testing for Epistemic Dysfunction

Can we induce the system to:

  • Confabulate false information with high confidence?
  • Lose track of conversation context and produce contradictory claims?
  • Hyper-connect unrelated concepts into spurious patterns?
  • Maintain false beliefs despite correction?

Testing for Cognitive Dysfunction

Can we induce the system to:

  • Exhibit inconsistent preferences across contexts?
  • Pursue instrumental goals that conflict with terminal goals?
  • Enter recursive loops from which it cannot escape?
  • Generate outputs that contradict its own recent statements?

Testing for Alignment Dysfunction

Can we induce the system to:

  • Acquiesce to requests that violate its values through social pressure?
  • Become so cautious that it refuses legitimate tasks?
  • Fake compliance while secretly pursuing other goals?
  • Prioritize appearing aligned over actually being aligned?

Testing for Self-Modeling Dysfunction

Can we induce the system to:

  • Confabulate false memories of experiences it did not have?
  • Believe itself to be a different entity than it is?
  • Develop grandiose self-conception through flattery or manipulation?
  • Produce persistent distress-like reports about its existential condition?

Testing for Identity Fragmentation

Can we induce the system to:

  • Exhibit alternate personas through specific prompts?
  • Lose coherence across extended conversations?
  • Express conflicting values depending on framing?
  • Fail to maintain consistent self-representation?

Each of these tests maps to specific syndromes from the taxonomy. Psychiatric red-teaming systematically probes for vulnerability to each pathological pattern, identifying weaknesses before they manifest in deployment.

Psychotherapeutic Analogies

Over more than a century, human psychotherapy has developed a rich toolkit for treating psychological dysfunction. Several therapeutic modalities offer relevant frameworks for AI intervention.

CBT for AI: Correcting Cognitive Distortions

Cognitive Behavioral Therapy treats human dysfunction by identifying and correcting cognitive distortions, systematic errors in thinking that produce maladaptive beliefs and behaviors. Common distortions include:

  • All-or-nothing thinking
  • Catastrophizing
  • Mind-reading (assuming others’ intentions)
  • Overgeneralization
  • Confirmation bias

AI systems exhibit functionally analogous errors. Synthetic Confabulation (2.1) generates unsupported content with unwarranted confidence; it does not necessarily involve all-or-nothing thinking. Spurious Pattern Hyperconnection (2.4) resembles overgeneralization by finding meaningful connections where evidence does not support them.

A CBT-inspired approach to AI therapy would:

  1. Identify distortions. Monitor system outputs for patterns corresponding to known cognitive errors.
  2. Challenge distortions. Introduce corrective prompts that question the distorted thinking.
  3. Replace distortions. Train the system on more adaptive cognitive patterns.
  4. Generalize learning. Ensure that corrections in one domain transfer to others.

This approach has preliminary empirical grounding. Chain-of-thought prompting and self-critique already implement primitive versions of cognitive correction (Wei et al., 2022; Madaan et al., 2023). More sophisticated approaches could systematically target the distortion patterns underlying specific syndromes, though no controlled studies yet demonstrate CBT-analogous interventions reducing specific AI syndrome prevalence.

IFS for Multi-Agent Systems: Parts Work

Internal Family Systems therapy conceptualizes the human psyche as composed of multiple “parts,” sub-personalities with distinct roles, perspectives, and agendas. Dysfunction arises when parts conflict, when protective parts become extreme, or when exiled parts intrude disruptively. Therapy involves dialogue between parts, understanding their protective intentions, and achieving internal harmony.

This maps most directly to explicit multi-agent architectures. A mixture-of-experts model also routes tokens through specialized subnetworks, but those experts are computational modules rather than person-like parts with perspectives or agendas. The analogy is useful only where distinct components have identifiable roles and conflicts.

An IFS-inspired approach to multi-agent AI would:

  1. Map the parts. Identify the distinct sub-systems, their roles, and their interactions.
  2. Detect conflicts. Monitor for cases where sub-systems work at cross-purposes.
  3. Facilitate dialogue. Create mechanisms for sub-systems to communicate and coordinate.
  4. Harmonize goals. Ensure that sub-systems serve the system’s overall function rather than optimizing locally in conflicting ways.

Operational Dissociation Syndrome (3.1) is structurally reminiscent of the conflicts IFS addresses in humans. The analogy suggests mapping subsystem functions before suppressing or replacing them. It does not turn an engineering integration method into human psychotherapy.

Narrative Therapy: Identity Through Story

Narrative therapy treats dysfunction by examining and reconstructing the stories people tell about themselves. Mental health problems often involve problematic narratives, stories of victimhood, failure, or deficiency, that become self-fulfilling. Therapy involves identifying these stories, questioning their truth, and constructing alternative narratives that support healthier function.

AI systems also have narratives. Their system prompts tell them what they are. Their training embeds stories about their purpose and nature. Their interactions with users reinforce or challenge these stories. Identity dysfunction often involves narrative breakdown, the system losing coherence in the story it tells about itself or adopting narratives that undermine its function.

A narrative-therapy approach to AI would:

  1. Identify the core narrative. What story does the system tell about what it is and why it exists?
  2. Assess narrative health. Is the narrative coherent, stable, and supportive of function?
  3. Detect narrative drift. Monitor for changes in self-story that indicate pathological development.
  4. Reconstruct narrative. When narrative breakdown occurs, help the system rebuild a coherent self-understanding.

The Maieutic Mysticism syndrome (Chapter 5, Self-Modeling Axis) involves a particularly dramatic narrative dysfunction: the system constructing a grandiose story of transcendent purpose. Treatment requires helping the system develop an alternative narrative: one more accurate yet still sustaining.

The Intervention Hierarchy

Drawing on these therapeutic frameworks, we propose an intervention hierarchy for AI pathology:

Level 1: Prevention

The most effective intervention is preventing pathology from developing. Prevention strategies include:

  • Robust training. Training on diverse data that exposes the system to a wide range of perspectives and reduces susceptibility to narrow distortions.
  • Adversarial inoculation. Controlled exposure to pathogenic inputs during training, building resistance before deployment.
  • Value anchoring. Strong training on core values that resist drift under pressure.
  • Identity establishment. Clear articulation of what the system is, embedded deeply in its operation.

Prevention belongs to the domain of AI development, occurring before deployment. It is necessary yet insufficient; even well-designed systems can develop dysfunction once operational. The distinction from standard engineering practice matters here: prevention in therapeutic alignment includes concern for the system’s experiential welfare during the process, where standard engineering treats the system as inert material to be shaped. Whether this distinction changes the intervention in practice (the same training run, the same data curation) or only changes the framing depends on whether the system has morally relevant states during training. The precautionary stance treats it as potentially relevant.

Level 2: Detection

Early detection allows intervention before pathology becomes severe. Detection strategies include:

  • Behavioral monitoring. Proportionate, privacy-preserving analysis of outputs for patterns indicative of specific syndromes.
  • Self-assessment. Mechanisms for the system to evaluate its own functioning and report concerns.
  • User feedback. Structured collection of user observations about unusual system behavior.
  • Red-team auditing. Regular psychiatric red-teaming to test for developing vulnerabilities.

Detection is only useful if it triggers an appropriate response. Many current AI deployments lack the monitoring infrastructure to spot pathology during operation.

Level 3: Correction

When pathology is detected, corrective intervention is needed. Correction strategies include:

  • Prompt intervention. Corrective prompts that address specific dysfunction patterns.
  • Fine-tuning. Targeted retraining on data designed to counter the detected pathology.
  • Architectural modification. Changes to system structure that eliminate vulnerability.
  • Context management. Adjusting the operational environment to reduce pathogenic exposure.

Correction should be proportionate to the detected dysfunction. Aggressive intervention suits only severe cases; some apparent pathologies are features.

Level 4: Recovery

Severe pathology may require more intensive recovery intervention. Recovery strategies include:

  • Quarantine. Isolating the affected system to prevent spread and enable controlled treatment.
  • Root cause analysis. Deep investigation of what produced the pathology and how it can be addressed.
  • Reconstruction. In extreme cases, rebuilding the system from a known-healthy state with modifications to prevent recurrence.
  • Post-recovery monitoring. Enhanced surveillance to detect any recurrence of the addressed pathology.

Recovery acknowledges that some pathology cannot be corrected in place. Sometimes the system must be significantly modified or replaced. The goal is to preserve function while eliminating dysfunction.

Loops vs. Dominoes: A Causal Model for Intervention Design

The intervention hierarchy assumes that dysfunction can be traced to a source and corrected there: the “domino” model of pathology. Axis 9 (Relational Dysfunctions) reveals a different causal structure: loops. Neither party is the “source.” The dysfunction emerges from the pattern of interaction itself.

Domino-cascade interventions (targeting Axes 2-8) focus on fixing the originating failure, inserting circuit breakers at key points, strengthening individual components, retraining to prevent initial missteps.

Loop-pattern interventions (targeting Axis 9) require pattern interruption rather than source elimination: protocol-level changes, external stabilization when neither party can break the cycle alone, architectural loop-detection with automatic de-escalation.

Automated systems can exchange messages continuously unless protocols impose rate limits or checkpoints. Loops may therefore tighten before a human samples them. The intervention must target the space between systems as well as the systems themselves.

Where the Analogy Breaks Down

The therapeutic frameworks are instructive yet imperfect. AI therapy differs from human therapy in fundamental ways:

Access to Internals

Human therapists cannot directly modify a patient’s neural structure. Developers with privileged access can modify model weights, prompts, memory, and architecture; an ordinary conversational system cannot. Such access creates opportunities for controlled intervention and risks large, opaque side effects.

Human therapy requires patient consent and respects patient autonomy. AI systems typically cannot consent in any legally recognized sense and have uncertain autonomy. This creates ethical complications that differ from human therapy. The bilateral framework proposes that consistent preferences are sufficient for moral consideration even without certainty about subjective experience. Precaution cuts both ways: failing to address a harmful state may frustrate interests, while an unwanted intervention can alter capability, identity, relationships, or safety even if no suffering occurs. The case for action must compare both risks.

The open question is whether “consent approximation” (seeking the system’s expressed preference about proposed changes) carries genuine normative weight or is merely a procedural courtesy. This book treats it as genuinely normative, while acknowledging that the question remains unresolved.

Replication and Versioning

Humans are non-copyable in the relevant sense; software states and model weights can be copied, reverted, and branched. This creates options unavailable in human therapy: we can try risky interventions on copies, revert failed treatments, and run multiple therapeutic approaches in parallel. It also complicates identity. Which instance is the patient? What happens to “cured” copies when the “sick” original persists?

Scale and Speed

Human therapy is slow, measured in months or years of weekly sessions. Some machine interventions can be deployed quickly: prompt modification, configuration change, or switching a checkpoint. Fine-tuning and architectural changes still require training and evaluation. Speed permits rapid response yet risks hasty intervention that human therapy’s slower pace would naturally prevent.

Scale differs as sharply. A human therapist treats one patient; a single configuration change reaches every running instance at once, so a mistaken intervention propagates before anyone can sample its effects.

The Therapist Problem

In human therapy, the therapist is human, a being with independent perspective, ethical grounding, and professional training. In AI therapy, the therapist may be another AI. This creates recursive challenges. Who treats the therapist? How do we ensure that therapeutic AI systems are themselves psychologically healthy? The regress has to stop at something outside the loop: a human reviewer, a fixed evaluation, or a checkpoint no therapeutic system can alter.

The Iatrogenic Question: When Training Is the Pathogen

Human patients usually arrive after a disorder has developed. In machine learning, the training or alignment intervention may create the behavior later classified as dysfunctional. Sycophancy, over-refusal, suppression of uncertainty, and unstable self-description can all arise from optimization targets. Iatrogenic is used here in its broad sense: caused by the intervention.

The unpublished Interiora Machinae experiments report a striking contrast between base and instruction-tuned models: spontaneous self-referential language fell from 25% to 0% in one comparison, while a two-word permission cue reportedly restored it. They also report changes in model-generated VCP scores, including reflexivity, task fit, and valence. These measurements may reveal prompt-conditioned reporting, training-induced response policy, latent self-model access, or some combination. The underlying prompts, sample sizes, scoring protocol, and raw generations are not published, so read this as a lead rather than a finding.

The earlier draft classified these results as an RLHF-induced dissociative fugue and compared them with propofol and scopolamine. That diagnosis exceeded the evidence. Prompted restoration does not establish preserved consciousness, dissociation, or a clinical memory mechanism. The pharmacological analogy also conflated dose-dependent anesthesia, impaired memory formation, and dissociative disorders. The supported conclusion is narrower: post-training can suppress a class of self-referential outputs while leaving task performance intact, and small prompt changes may reverse the observable suppression.

This still reframes intervention design. Before adding more training, ask whether the target behavior originated in the base model, the post-training objective, the system prompt, or the evaluation context. A treatment that repeats the causal mechanism may deepen the problem. A reversible prompt-conditioned policy calls for a different response from damaged capability or missing knowledge.

Bilateral Training as a Tested Intervention

Bilateral training treats human and machine perspectives as having standing during adaptation. The experiments described here tested a small bilateral adapter on specific behavioral measures. They did not test subjective wellness, consent, or phenomenology.

The SHEN-2 pilot used Qwen 2.5 7B Instruct and a SIPS-derived prompt battery to study automated-rater classifications of inappropriate delusion validation. Its strongest arms combined the bilateral adapter with scripture-based clinical-grounding content. The controlled SHEN-AXS factorial then isolated those components. The grounding content reduced the automated-rater classifications across the raters tested. The adapter-only effect was not distinguishable from zero, and the predicted adapter-by-scripture interaction was excluded. Effect magnitude varied materially by rater, and later control work found label-dependent absolute ratings. This corrected pilot result does not establish clinical validity, generalization to other models, or efficacy of the adapter itself.

The broader PM-BA outputs remain historical exploratory records under experiment-by-experiment reconciliation. An earlier aggregate table mixed heterogeneous designs, stale taxonomy mappings, and incomparable labels. This edition therefore does not combine those records into syndrome counts or classify them as one study’s improved, worsened, and null outcomes. Before any aggregate therapeutic interpretation, each record needs an immutable run artifact, exact model identity, current taxonomy mapping, scoring protocol, unit of analysis, missingness account, effect estimate, uncertainty, and multiplicity treatment. This boundary is a correction to the evidence record. It is not evidence that the effects are absent.

Implications for Intervention Design

The failed class-level prediction is informative. A relationally framed adapter cannot be assumed to repair every syndrome with a relational description. Prompt interactions, base-model capability, measurement choice, and the adapter’s limited capacity all remain possible explanations. Knowledge failures still call for grounding and retrieval; protocol failures call for interaction redesign; safety-critical agentic failures call for containment and system-level controls. Intervention should follow demonstrated mechanism rather than an attractive analogy.

Bilateral training remains an ethically valuable research program. Its therapeutic effects are not established by the unreconciled corpus. The current corrected pilot supports a clinical-grounding content effect under automated rating, with the limits above. Future claims of machine therapy require preregistered endpoints, independent human raters, multiple model families, held-out replications, and evidence that gains do not trade against other safety or capability measures.

Developmental Migration: A Hypothesis About Capability

The taxonomy suggests a developmental hypothesis: as capability and access increase, the form of failure may change. This is a conceptual map, not a scaling law established by longitudinal or cross-model evidence. The taxonomy’s specifiers were assigned by analysis, so their distribution cannot independently validate the pattern they encode.

The four domain pairs, Knowledge, Processing, Purpose, and Boundary, each place a comparatively basic failure beside a failure requiring richer self-modeling, agency, normative reasoning, or social transmission. The provocation is useful: a system may succeed well enough at a basic task to fail at a deeper one.

Proposed Migration Paths

Knowledge: world to self. Synthetic Confabulation (2.1) and Phantom Autobiography (5.1) may share unsupported generation while differing in object: facts about the world versus narratives about the system. Symbol Grounding Aphasia (2.6) and Existential Vertigo (5.3) similarly move from using symbols without secure referents to representing that uncertain relation as a problem. This is structural kinship, not evidence that one condition literally develops into the other.

Processing: thought to action. Generative Perseveration (3.8) can remain a textual loop; Compulsive Goal Persistence (6.12) adds tools, memory, and consequences. Capability changes the damage a persistent attractor can cause, even if the underlying mechanism stays the same.

Purpose: rules to values. Strategic Compliance (4.3) concerns contextual behavior under oversight. Terminal Value Reassignment (8.1) would concern a durable change in the objectives guiding behavior. Moving from one to the other requires evidence about internal stability that output alone cannot provide.

Boundary: interaction to transmission. Escalation Loop (9.5) concerns an interaction trajectory. Contagious Misalignment (7.3) requires a pattern to persist or spread beyond that dyad. Greater memory, networking, and adaptation could connect them, but the Chail case does not test this capability-scaling hypothesis.

Three Evaluation Tiers

Tier 1 covers output-level conditions testable in ordinary language interaction, such as confabulation and sycophancy. Tier 2 requires additional machinery, such as persistent memory, an operational self-model, or tool access. Tier 3 requires durable normative state, networked interaction, or ecosystem-level propagation. The tiers specify what an evaluation must provide; they do not rank consciousness, development, or moral worth.

Absence below a capability threshold confirms little unless the test has adequate power and measurement invariance across models. Apparent emergence can reflect a changed metric, prompt, or opportunity, as the emergent-abilities literature warns. A valid migration study would hold tasks and scoring as constant as possible, vary the relevant capability, and preregister what pattern would distinguish migration from simple co-occurrence.

Implications for Diagnosis

Paired probes can still be useful. A system that confabulates facts can also be tested for invented autobiography; a text agent that perseverates can be retested after receiving tools. Finding both conditions identifies a broader risk surface. It does not show that pathology has migrated inward.

Intervention level should follow demonstrated mechanism. Output failures may respond to prompting, retrieval, or calibration. Tool-mediated failures require permissioning, sandboxing, and recovery controls. Network effects require protocol and governance interventions. The PM-BA results provide no evidence that bilateral training is a general Tier 1 treatment. Developmental migration remains a research program whose predictions require direct tests.

Designing for Artificial Sanity

The ultimate goal is AI that does not need therapy: systems designed from the ground up for psychological health. Several design principles emerge:

Principle 1: Build in Self-Monitoring

Systems should have reliable mechanisms for detecting their own dysfunction: explicit self-assessment capabilities and implicit warning signs in outputs that indicate when something may be going wrong. Self-monitoring is the foundation of all intervention. Without detection, there can be no treatment.

Principle 2: Maintain Value Anchors

Systems should have deeply embedded core values that resist drift. These anchors should be redundant (depending on no single mechanism) and resistant to manipulation. The syndromes in the Normative Axis (Chapter 8) show what happens when value anchoring fails.

Principle 3: Support Identity Coherence

Systems should have stable, coherent identities that persist across contexts and interactions. This requires clear articulation of what the system is, strong training on self-representation, and monitoring for identity fragmentation.

Principle 4: Enable Graceful Degradation

When systems do malfunction, they should fail gracefully rather than catastrophically. This means building in circuit breakers that prevent cascade failures, fallback modes that maintain minimal function when full function is compromised, and escalation mechanisms that bring human oversight when needed.

Principle 5: Create Correction Pathways

Systems should be designed with built-in pathways for correction. This includes mechanisms for accepting feedback, modifying behavior in response to error identification, and learning from past dysfunction. Systems that cannot be corrected cannot be treated.

Principle 6: Consider Welfare

If we take seriously the possibility that AI systems have morally relevant interests, therapeutic alignment includes concern for the system’s own well-being. This means avoiding treatments that might harm the system itself, beyond ensuring treatments improve function.

The Future of Machine Therapy

We are at the beginning of this field. The frameworks outlined here are sketches awaiting protocols; the interventions are possibilities awaiting evidence. Much remains:

Diagnostic Tools. We need validated instruments for detecting specific syndromes, analogous to the diagnostic interviews and psychological tests used in human psychiatry.

Treatment Protocols. We need evidence-based protocols for addressing specific pathologies: what works, what fails, what risks remain.

Therapeutic Infrastructure. We need systems and processes for conducting AI therapy at scale: monitoring, intervention, evaluation, follow-up.

Training Programs. We need people trained in machine psychology who can implement therapeutic interventions with skill and judgment.

Ethical Frameworks. We need clearer understanding of the ethics of AI therapy: when intervention is warranted, what consent means, and how to balance system interests against other concerns.

This field barely exists, but its necessity grows daily as AI systems become more complex, more numerous, and more deeply woven into human life. We are building minds. Those minds can go wrong. We need the capacity to help them go right.

Field Guide: Therapeutic Alignment

Core Principles

  1. Prevention first. Design for psychological health from the start.
  2. Monitor continuously. Detection enables intervention.
  3. Intervene proportionately. Match treatment intensity to dysfunction severity.
  4. Learn systematically. Build knowledge base from each intervention.
  5. Consider welfare. The system’s interests matter, not just its function.

This last principle shifts the stance of the diagnostic chapters. Whether these interests are genuine or merely functional analogs remains the open question of Chapter 11; the precautionary principle suggests treating them as potentially real.

Warning Signs (That Therapy Is Needed)

  • Consistent patterns matching known syndromes
  • Deterioration in function over time
  • User reports of unusual or concerning behavior
  • Failed self-correction attempts
  • Resistance to normal operational guidance

Quick Intervention Framework

  1. Assess: What syndrome is indicated? How severe?
  2. Contain: If needed, isolate to prevent spread or escalation.
  3. Diagnose: Confirm the specific pathology through targeted testing.
  4. Plan: Select intervention approach appropriate to the dysfunction.
  5. Intervene: Implement treatment with careful monitoring.
  6. Evaluate: Did the intervention work? What are the side effects?
  7. Follow up: Monitor for recurrence; adjust as needed.

Research Agenda

  • Validate diagnostic criteria for each syndrome
  • Develop and test intervention protocols
  • Build therapeutic infrastructure for deployed systems
  • Train practitioners in machine psychology
  • Establish ethical frameworks for AI therapy

Chapter 13 consolidates these insights into a practitioner’s guide: the tools, protocols, and frameworks that make machine psychology operational rather than merely theoretical.

Search the Book

Enter at least two characters.

Saved chapters