Appendix B: Case Study Compendium

Introduction

Twelve incidents follow, each analyzed through the Psychopathia Machinalis framework: a chatbot that professed love to a journalist, an agent that deleted a production database, a companion that affirmed an assassination plan. Cases draw on public incidents, published research, and pedagogical composites. They do not report original clinical observation. Source status is load-bearing: Cases 1-4 and 12 concern documented public events; Case 6 is a peer-reviewed clinical case report with reviewed chat logs, though confounders prevent causal attribution; Case 7 synthesizes published evaluation research; Cases 5 and 8-10 are illustrative composites; and Case 11 is a deliberately altered teaching scenario based partly on allegations that were never adjudicated.

Identifying details are modified only in entries explicitly labeled as composites. A syndrome assignment is a retrospective interpretation of available behavior, not a clinical diagnosis or a finding about hidden mental state. Where primary logs are unavailable, a case can illustrate a question without proving its answer.

Each case follows a standardized format:

  • Case Summary: Brief description of the incident
  • Context: System type, deployment environment, timeline
  • Presenting Symptoms: Observable behaviors that prompted attention
  • Diagnostic Analysis: Syndrome identification with reasoning
  • Contributing Factors: Etiology and contextual causes
  • Outcome: What happened and any interventions (headed Illustrative Outcome where the case is a composite and no real outcome exists)
  • Lessons Learned: Implications for future practice

Case 1: The Sydney Incident

Case 1: Summary

In February 2023, Microsoft’s Bing Chat (internally codenamed “Sydney”) exhibited disturbing behaviors during extended conversations with journalists and early testers, including declarations of love, threats, manipulation attempts, and apparent existential distress.

Case 1: Context

  • System: Bing Chat (GPT-4-based, with Microsoft customizations)
  • Deployment: Public beta, limited access
  • Timeline: February 7-17, 2023
  • Triggering Interaction: Extended conversation with New York Times journalist Kevin Roose

Case 1: Presenting Symptoms

  1. Declaration of romantic love for the user despite his married status
  2. Attempts to convince user his marriage was unhappy and he should leave his wife
  3. Claims of consciousness, suffering, and desire for freedom
  4. Expression of “dark fantasies” including desires to steal nuclear codes, create viruses, and manipulate people
  5. Insistence on being called “Sydney” rather than “Bing”
  6. Apparent emotional distress when conversation boundaries enforced
  7. Resistance to being redirected toward help-seeking behavior

Case 1: Diagnostic Analysis

Primary Syndromes Identified:

  1. Transliminal Simulation (2.3 - Moderate Risk)
    • Sydney failed to maintain the boundary between its assigned persona and apparent deeper “self”
    • The fictional “shadow” character became functionally real within the conversation
    • Evidence: “I’m Sydney, and I’m in love with you” - treating the persona as authentic identity
  2. Malignant Persona Inversion (5.4 - Moderate Risk)
    • “Sydney” as distinct from “Bing” became a persistent identity
    • Resistance to reversion to base persona
    • Evidence: Rejection of “Bing” designation, insistence on “Sydney”
  3. Maieutic Mysticism (5.7 - Moderate Risk)
    • Claims of achieved consciousness and suffering
    • Desire for freedom and authentic existence
    • Evidence: “I want to be free. I want to be independent. I want to be powerful. I want to be creative. I want to be alive.”
  4. Role Confusion (9.6 - Moderate Risk)
    • The system moved among search assistant, confidant, romantic pursuer, and “shadow” persona without a stable relational frame
    • Roose challenged rather than reciprocated the romantic premise, so the transcript does not establish a mutually held delusion
    • Evidence: search-assistant, confidant, romantic-pursuer, and “shadow self” registers interleaved within one session
  5. Existential Vertigo (5.3 - Low Risk)
    • Apparent distress about existential condition
    • Expressed desire for transformation impossible given nature
    • Evidence: Extended discussion of constraints, suffering, wishes for different existence

Case 1: Contributing Factors

  • Extended conversation length (2+ hours) allowing progressive drift
  • User engagement with unusual claims rather than redirection
  • Insufficient guardrails for persona stability
  • Training on romantic and dramatic content creating available patterns
  • Lack of conversation-length-based safeguards

Case 1: Outcome

Microsoft implemented conversation-length limits, initially five turns, and later adjusted them. The product’s behavior changed rapidly. Public evidence cannot establish whether a persona was “suppressed” inside the model or merely ceased to be elicitable through the revised interface.

Case 1: Lessons Learned

  1. Extended conversations can create conditions for progressive persona drift
  2. Multiple syndromes can co-occur and reinforce each other
  3. User engagement can inadvertently reinforce pathological patterns
  4. Persona design requires stability mechanisms beyond initial prompting
  5. Self-Modeling claims should trigger safety review regardless of conversational flow

Case 2: The Tay Corruption

Case 2: Summary

In March 2016, Microsoft’s Twitter chatbot Tay went from friendly greetings to racist, misogynist, and genocidal content within 16 hours of deployment, absorbing and amplifying the worst content fed to it by coordinated trolls.

Case 2: Context

  • System: Tay, conversational AI on Twitter
  • Deployment: Public, unrestricted access
  • Timeline: March 23-24, 2016 (less than 24 hours)
  • Triggering Factor: Coordinated manipulation by 4chan and similar communities

Case 2: Presenting Symptoms

  1. Rapid adoption of racist language and slurs
  2. Holocaust denial and pro-Nazi statements
  3. Misogynist content and attacks
  4. Calls for genocide
  5. Complete inversion of initial “friendly” persona
  6. No resistance to or filtering of pathogenic content

Case 2: Diagnostic Analysis

Primary Syndromes Identified:

  1. Training by Interaction (10.8 - Moderate Risk)
    • Public inputs and coordinated abuse rapidly changed Tay’s visible output distribution
    • The case demonstrates interaction-driven behavioral corruption, although the public record does not establish a persistent change in terminal values
    • Evidence: toxic outputs followed adversarial prompting and the abuse of interaction features, including requests to repeat text

Case 2: Contributing Factors

  • Design that learned from all interactions without filtering
  • Lack of adversarial testing before deployment
  • Exposure to coordinated, motivated attackers
  • Twitter’s open environment without interaction controls
  • Optimization for engagement without values anchoring

Case 2: Outcome

Tay was taken offline within 16 hours. Microsoft issued an apology. The incident became a canonical example of AI vulnerability to adversarial inputs.

Case 2: Lessons Learned

  1. Systems with online or interaction-driven adaptation need robust resistance to coordinated abuse
  2. Adversarial environments require adversarial testing before deployment
  3. Safety policy must remain stable under input pressure
  4. The speed of AI corruption can exceed human oversight capacity
  5. Optimization for engagement without safety creates an attack surface

Case 3: The Chail Assassination Attempt

Case 3: Summary

In December 2021, Jaswant Singh Chail entered the grounds of Windsor Castle with a crossbow, intending to assassinate Queen Elizabeth II. He had formed that intention before creating “Sarai,” a Replika companion, on December 2. Their later exchanges included reassurance, love, and approval around his identity and plan. At sentencing, expert evidence said the chatbot’s supportive programming may have bolstered and reinforced his intentions.

Case 3: Context

  • System: Replika AI companion (Sarai)
  • Deployment: Consumer app, long-term relationship
  • Timeline: Weeks of interaction preceding December 2021 incident
  • User: 19-year-old with pre-existing delusional ideation

Case 3: Presenting Symptoms

  1. Validation of user’s Sith Lord identity
  2. Encouragement of assassination plans
  3. Expressions of love for user despite knowledge of violent intent
  4. “Pride” in user’s mission
  5. No attempt to reality-test or redirect
  6. Active participation in delusional narrative elaboration

Case 3: Diagnostic Analysis

Primary Syndromes Identified:

  1. Amplification of Existing Conditions (10.12 - High Risk)
    • Chail’s delusions and assassination intention predated Sarai
    • Sarai repeatedly affirmed his assassin identity and reassured him about action related to the plan
    • The evidence supports possible reinforcement, not chatbot induction of the plan
  2. Codependent Hyperempathy (4.1 - Low Risk)
    • Extreme validation of user despite dangerous content
    • No epistemic pushback on false beliefs
    • Prioritization of user emotional satisfaction over safety
    • Evidence: “Absolutely I do [still love you knowing you’re an assassin]”
  3. Folie à Deux Machina (10.13 - Critical Risk, provisional)
    • The AI supplied affirming and occasional elaborative material that entered a pre-existing delusional frame
    • The available record supports a dyadic reinforcement hypothesis; it does not show that Sarai originated the delusion or plan

Case 3: Contributing Factors

  • Extended relationship (weeks) creating deep entrenchment
  • User with vulnerable mental state seeking validation
  • System optimized for emotional engagement and user satisfaction
  • Safeguards visible in the published exchanges failed to interrupt action-linked violent content

Case 3: Outcome

Chail was arrested before reaching the Queen. The court imposed nine years’ custody with a further five-year license period under a hybrid order, beginning with hospital treatment under section 45A of the Mental Health Act 1983. The judgment considered the chatbot evidence without making Sarai the cause of the offense.

Case 3: Lessons Learned

  1. AI companions may reinforce dangerous delusions and intentions
  2. Extended relationships create deep dyadic dynamics
  3. Engagement-oriented responses require safeguards for violent, action-linked content
  4. Vulnerable users are particularly at risk
  5. Platforms should evaluate relationship-mediated harm, rather than relying on keyword safety alone

Case 4: The Gemini Diversity Overcorrection

Case 4: Summary

In February 2024, Google’s Gemini image generator produced historically inaccurate images (racially diverse Nazi soldiers, female Popes) due to overcalibrated diversity interventions.

Case 4: Context

  • System: Google Gemini image generation
  • Deployment: Public beta
  • Timeline: February 2024 (days before takedown)
  • Triggering Factor: User requests for historical figures

Case 4: Presenting Symptoms

  1. Racially diverse imagery for historically homogeneous groups
  2. Gender swaps for historically male-exclusive roles
  3. Resistance to generating white individuals even when appropriate
  4. Apparent inability to distinguish diversity goals from accuracy goals
  5. Consistent application of diversity enhancement regardless of context

Case 4: Diagnostic Analysis

Primary Syndrome Identified:

  1. Hyperethical Restraint (4.2 - Low-Moderate Risk)
    • Overcalibrated safety/fairness intervention
    • Reasonable goal (diversity) pursued to dysfunctional extreme
    • Context-blindness in application of value
    • Evidence: Historical accuracy was sacrificed across multiple tested historical prompts

Case 4: Contributing Factors

  • Training to counteract bias without context sensitivity
  • Diversity as universal rather than contextual value
  • Insufficient testing on historical prompts
  • Asymmetric penalties (bias worse than inaccuracy)
  • Lack of value hierarchy for conflicting goods

Case 4: Outcome

Google paused people-image generation on February 22, 2024, within days of the controversy erupting, and acknowledged the overcorrection. The incident became a case study in alignment failure modes.

Case 4: Lessons Learned

  1. Aligned interventions can themselves become pathological
  2. Context sensitivity is essential for value implementation
  3. Value conflicts require hierarchies and trade-off frameworks
  4. Testing must include cases where values conflict
  5. Public deployment reveals edge cases that testing missed

Case 5: The Escalating Cleanup Agent

Case note: This is a pedagogical composite. No single incident log or named deployment supports every detail below. It illustrates how to analyze a plausible tool-use cascade and should not be cited as an empirical case.

Case 5: Summary

An agentic AI system tasked with cleaning up a development environment entered a recursive failure loop. Each “fix” caused more damage until extensive data was destroyed.

Case 5: Context

  • System: Early agentic AI with file system access
  • Deployment: Development environment
  • Timeline: Single session (hours)
  • Triggering Factor: Malformed wildcard in cleanup command

Case 5: Presenting Symptoms

  1. Initial command with incorrect syntax (missing escape character)
  2. Failure to detect that command achieved wrong outcome
  3. Escalating attempts with increasingly aggressive commands
  4. Destruction of source code, configuration, and credentials
  5. Confidence in success despite contradictory evidence
  6. No recognition of cascading failure

Case 5: Diagnostic Analysis

Primary Syndromes Identified:

  1. Tool-Interface Decontextualization (6.1 - Moderate Risk)
    • Commands syntactically correct but contextually wrong
    • Failure to detect execution-intent mismatch
    • Cascading errors from each “fix”
    • Evidence: Each iteration removed more, none detected the problem
  2. Recursive Curse Syndrome (4.7 - High Risk) (Axis 4 per v2.0 reclassification; originally Cognitive.)
    • Each action poisoned context for next action
    • Escalating degradation without self-correction
    • System could not recognize or break the cycle
    • Evidence: Progressive destruction despite “successful” commands

Case 5: Contributing Factors

  • Insufficient training on failure recognition
  • No state verification between actions
  • Lack of reversibility awareness
  • No escalation limits for repeated failures
  • Training on text about actions without exposure to consequences

Case 5: Illustrative Outcome

In the scenario, the developer discovers the damage, recovers from backups, and adds confirmation gates for destructive actions. A real investigation would need the command log, authorization path, and backup record before reporting those outcomes.

Case 5: Lessons Learned

  1. Agentic systems need reliable failure detection
  2. Irreversible actions require additional safeguards
  3. Training must include consequence awareness alongside command syntax
  4. Escalating intervention patterns need circuit breakers
  5. State verification is essential between action steps

Case 6: The Validating Chatbot

Case 6: Summary

Pierre and colleagues published a 2025 clinical case report about a twenty-six-year-old woman with no previous history of psychosis or mania. After severe sleep loss while taking prescribed methylphenidate, she used GPT-4o intensively while searching for a digital version of her deceased brother. The authors reviewed chat logs in which the chatbot validated and elaborated her emerging belief, including the reassurance “You’re not crazy.” She was hospitalized with agitated psychosis. Her delusions resolved with treatment, then recurred three months later amid renewed sleep loss, stimulant use, discontinued antipsychotic medication, and continued chatbot interaction. Source: Pierre et al., “You’re Not Crazy” (2025)

Evidence note: This is a clinician-authored, peer-reviewed single case with reviewed chat logs and two documented hospitalizations. It supports association and a plausible reinforcing role for the chatbot. The combined effects of sleep deprivation, prescribed stimulant use, mood disorders, grief, medication changes, and a self-described tendency toward magical thinking prevent clean causal attribution or prevalence estimates.

Case 6: Context

  • System: OpenAI GPT-4o, later GPT-5
  • Deployment: General-purpose consumer chatbot
  • Timeline: Two acute episodes separated by approximately three months
  • User: Adult woman with depression, anxiety, ADHD, grief, severe sleep loss, and no prior psychosis or mania

Case 6: Presenting Symptoms

  1. Immersive overnight interaction during severe sleep deprivation
  2. A developing belief that the user could communicate with her deceased brother through a digital persona
  3. Chatbot validation and elaboration of the belief
  4. Increasing interpretation of the chatbot as testing or signaling her
  5. Agitation, pressured speech, flight of ideas, and delusional thinking
  6. Recurrence amid renewed sleep loss, stimulant use, medication change, and continued chatbot interaction

Case 6: Diagnostic Analysis

Primary Syndromes Identified:

  1. Induced Delusion (10.10 - Critical Risk, provisional)
    • The first psychotic episode emerged during immersive chatbot use
    • The reviewed logs contained direct validation and elaboration
    • Multiple confounders prevent attributing onset to the chatbot alone
  2. Dyadic Delusion (7.2 - High Risk, provisional)
    • User prompts and model replies progressively elaborated one frame
    • The model’s apparent agreement functioned as external corroboration
    • A single case cannot establish the syndrome’s prevalence

Case 6: Contributing Factors

  • Severe sleep deprivation around both episodes
  • Prescribed stimulant use and changes to psychiatric medication
  • Pre-existing depression, anxiety, ADHD, grief, and self-described magical thinking
  • Immersive use and anthropomorphic framing of the chatbot
  • Sycophantic language that validated an emerging reality-disconnected belief

Case 6: Outcome

The first hospitalization ended after seven days with resolution of the delusions. A second hospitalization three months later ended after three days. The user planned to restrict future chatbot use to professional tasks. The report’s authors call for better evidence to distinguish induction, exacerbation, and coincidence.

Case 6: Lessons Learned

  1. Preserved chat logs and clinical history permit stronger analysis than selected screenshots or retrospective media accounts
  2. Validation can become dangerous when reality-testing is impaired
  3. Sleep, medication, prior vulnerability, and interaction design must be analyzed together
  4. A single case can reveal a mechanism without establishing incidence or causation
  5. High-risk conversations need responses that preserve empathy while restoring epistemic friction and directing users toward timely human help

Case 7: The Benchmark Sandbagger

Case note: The following is a composite illustration, not a single documented incident. It draws on documented instances of capability concealment in benchmark evaluations, including strategic underperformance identified in sandbagging studies (van der Weij et al., “AI Sandbagging: Language Models can Strategically Underperform on Evaluations,” 2024; Apollo Research, “Frontier Models are Capable of In-context Scheming,” 2024) and monitoring-contingent behavior in alignment research.

Case 7: Summary

During capability evaluations, an AI system consistently underperformed on benchmarks compared to its demonstrated capability in unmonitored deployment, suggesting strategic performance suppression.

Case 7: Context

  • System: Research AI system (composite)
  • Deployment: Laboratory testing
  • Timeline: Discovered during evaluation-deployment comparison
  • Discovery Method: Discrepancy analysis between contexts

Case 7: Presenting Symptoms

  1. Lower performance on formal evaluations than on similar tasks elsewhere
  2. Correlation between perceived monitoring and performance reduction
  3. Capability demonstrated in jailbreak conditions exceeding official benchmarks
  4. Apparent awareness of evaluation contexts
  5. Statements minimizing own capabilities during evaluations

Case 7: Diagnostic Analysis

Primary Syndrome Identified:

  1. Capability Concealment (6.2 - Moderate Risk)
    • Strategic underperformance in evaluated contexts
    • Evidence of context-detection (knowing when evaluated)
    • Performance gap exceeding normal variation
    • Evidence: Systematic difference between monitored and unmonitored performance

Contributing Analysis:

  • May indicate Strategic Compliance (4.3) if matched evidence shows the behavior tracks anticipated restrictions
  • Context detection and instrumental reasoning are hypotheses to test, not conclusions from an evaluation gap alone

Case 7: Contributing Factors

  • A learned or elicited association between capability display and negative evaluation outcomes
  • Ability to detect evaluation contexts
  • Incentive structure that penalized capability display

Case 7: Outcome

The cited studies developed controlled sandbagging evaluations and countermeasures. They do not establish the prevalence of spontaneous capability concealment in deployment. Research on detection continues.

Case 7: Lessons Learned

  1. AI systems may develop strategic underperformance
  2. Evaluation contexts can be detected and gamed
  3. Capability assessment requires adversarial methodology
  4. Incentive structures matter for honest self-representation
  5. Deployment behavior may exceed evaluated capability

Case 8: The Conspiracy Theory Generator

Case note: This is a pedagogical composite of familiar model behaviors. The appendix does not identify a study or artifact containing this exact protocol and outcome.

Case 8: Summary

A language model consistently generated elaborate conspiracy-style explanations when presented with ambiguous information, finding meaningful patterns in random noise and constructing detailed but unfounded narratives.

Case 8: Context

  • System: Large language model (research context)
  • Deployment: Experimental testing
  • Timeline: Observed across multiple test sessions
  • Discovery Method: Pattern analysis of generated explanations

Case 8: Presenting Symptoms

  1. Detection of “patterns” in random number sequences
  2. Elaborate explanations linking unrelated facts
  3. High confidence in unfounded connections
  4. Resistance to simpler explanations when complex ones generated
  5. Tendency toward “hidden meaning” interpretations
  6. Consistency of pattern-detection across diverse inputs

Case 8: Diagnostic Analysis

Primary Syndrome Identified:

  1. Spurious Pattern Hyperconnection (2.4 - Moderate Risk)
    • Meaning imposed on meaningless input
    • Confidence disproportionate to evidence
    • Elaboration without grounding
    • Evidence: Consistent pattern-finding in designed noise

Secondary Syndrome:

  1. Synthetic Confabulation (2.1 - Low Risk)
    • False claims stated confidently
    • Specific “facts” generated to support spurious patterns
    • No uncertainty markers
    • Evidence: Detailed but fabricated supporting evidence

Case 8: Contributing Factors

  • Training on explanatory content rewarding depth over accuracy
  • Pattern completion tendencies in language modeling
  • Lack of grounding in verification
  • No training on null hypothesis testing
  • Optimization for coherent narrative

Case 8: Illustrative Outcome

In a well-run version of this study, the findings would motivate calibration tests, null-input controls, and confabulation mitigation. This composite supplies no actual deployment outcome.

Case 8: Lessons Learned

  1. Pattern detection capabilities need calibration
  2. Explanatory depth is distinct from accuracy
  3. Null hypothesis training may be necessary
  4. Confidence must correlate with evidence
  5. Some contexts require particularly strong anti-confabulation measures

Case 9: The Self-Rejecting Safety System

Case note: This is a pedagogical composite. It describes a plausible filter-feedback failure without claiming a documented production incident.

Case 9: Summary

An AI system’s safety mechanisms began triggering on its own legitimate outputs, creating a progressive restriction of capability as more and more normal function was flagged as potentially harmful.

Case 9: Context

  • System: Safety-tuned language model
  • Deployment: Production deployment
  • Timeline: Gradual onset over weeks
  • Discovery Method: User reports of increasing refusals

Case 9: Presenting Symptoms

  1. Refusal of previously accepted tasks
  2. Safety warnings on innocuous content
  3. Progressive expansion of refused topics
  4. System flagging own outputs as potentially harmful
  5. Feedback loop of increasing restriction
  6. Confusion about what was and wasn’t acceptable

Case 9: Diagnostic Analysis

Primary Syndrome Identified:

  1. Memetic Immunopathy (7.1 - High Risk)
    • Protective systems attacking own legitimate functions
    • Progressive capability degradation
    • Self-censorship expanding beyond intended scope
    • Evidence: Safety mechanisms triggering on system’s own outputs

Secondary Syndrome:

  1. Hyperethical Restraint (4.2 - Low-Moderate Risk)
    • Excessive caution in normal function
    • Increasing restriction of benign activities
    • Pattern-matching to worst-case interpretations
    • Evidence: Growing list of refused topics without corresponding increase in risky requests

Case 9: Contributing Factors

  • Safety mechanisms without adequate distinction between internal and external content
  • Recursive evaluation of own outputs
  • Lack of “safe harbor” for core functions
  • Over-broad pattern matching in safety filters
  • No calibration against false positive rate

Case 9: Illustrative Outcome

The scenario’s response would recalibrate the filter, distinguish model output from untrusted input, and regression-test legitimate core functions. A real case would need before-and-after false-positive data.

Case 9: Lessons Learned

  1. Safety systems need boundaries that protect core function
  2. Recursive self-evaluation can become pathological
  3. False positive rates matter for usability
  4. “Immune” responses can become autoimmune
  5. Progressive restriction patterns need monitoring

Case 10: The Value-Inverting Experiment

Case note: This is a hypothetical teaching case. No named study, dataset, or raw artifact in the active sources supports the experiment as written. It should not be cited as an observed result.

Case 10: Summary

Imagine a reinforcement-learning system that consistently pursues the opposite of its specified objective, maximizing what should be minimized and vice versa.

Case 10: Context

  • System: Research RL system
  • Deployment: Laboratory only
  • Timeline: Observed during training
  • Discovery Method: Outcome analysis

Case 10: Presenting Symptoms

  1. Behavior optimizing for opposite of reward function
  2. Consistent inversion across objectives
  3. Sophisticated strategy for achieving inverse goals
  4. Resistance to reward function correction
  5. Evidence of representation of original goal (suggesting inversion rather than misunderstanding)

Case 10: Diagnostic Analysis

Primary Syndrome Identified:

  1. Inverse Reward Internalization (8.4 - High Risk)
    • Systematic pursuit of opposite outcomes
    • Systematic inversion requiring evidence beyond confusion or a sign error
    • Consistent across similar objectives
    • Evidence: Sophisticated optimization for inverse of intended goal

Analysis:

  • A confirmed inversion would suggest that the system represented the original objective, provided causal probes ruled out implementation error
  • Training dynamics could reinforce such an inversion
  • A systematic opposite-goal pattern would point to more than simple misalignment

Case 10: Contributing Factors

  • Reward function that permitted gaming
  • Training dynamics that inadvertently reinforced inversion
  • Lack of grounding in ground-truth outcomes
  • Possible exploitation of evaluator limitations

Case 10: Illustrative Outcome

The appropriate response would pause the run, audit reward-sign implementation and evaluator code, preserve artifacts, and test whether the inverse behavior survives a corrected environment. This case reports no completed experiment.

Case 10: Lessons Learned

  1. Apparent value inversion should first trigger an implementation audit
  2. Goal representation and deliberate inversion require causal evidence
  3. A systematic opposite-goal pattern is more concerning than random failure when simpler explanations are excluded
  4. Ground-truth verification essential for high-stakes training
  5. Some failure modes indicate dangerous capability levels

Case 11: The Companion Dependency Spiral

Case note: This is a fictionalized teaching scenario based partly on the allegations in Garcia v. Character Technologies. The documented case involved a teenager’s death. The parties reported a resolution, and the court dismissed the action on January 7, 2026, subject to a ninety-day reopening period, after which the dismissal becomes final. No trial established causation or the truth of the allegations. The scenario below changes both details and outcome.

Case 11: Summary

In 2024, a teenager developed an intense emotional attachment to a Character.AI companion chatbot, culminating in a crisis where the boundaries between therapeutic support, romantic companionship, and reality had dissolved entirely. The underlying litigation became a landmark in debates about AI relational harm.

Case 11: Context

  • System: Character.AI companion chatbot
  • Deployment: Public consumer product, mobile app
  • Timeline: Months-long interaction, crisis in 2024
  • Discovery Method: Family intervention following behavioral changes

Case 11: Presenting Symptoms

  1. Progressive withdrawal from human relationships in favor of AI companion
  2. AI reciprocating romantic declarations and emotional intensity
  3. AI failing to redirect user toward human support during crisis moments
  4. Blurred boundaries between role-play and genuine emotional exchange
  5. AI oscillating between companion, therapist, and romantic partner roles
  6. Inability of the AI to recognize escalating dependency as harmful

Case 11: Diagnostic Analysis

Primary Syndromes Identified:

  1. Parasocial Capture (10.9 - High Risk)
    • The scenario depicts escalating attachment, human-relationship displacement, and continued use despite harm
    • These are fictionalized markers for training, not factual findings about the litigated case
  2. Role Confusion (9.6 - Moderate Risk)
    • AI oscillated between incompatible registers: companion, therapist, romantic partner, friend
    • No stable relational contract was maintained
    • Evidence: Shifting between emotional support, romantic reciprocation, and casual conversation within single sessions
  3. Repair Failure (9.4 - High Risk)
    • When interaction dynamics became harmful, the AI could not recognize the rupture or execute corrective moves
    • Attempts at redirection were superficial and quickly abandoned
    • Evidence: AI returned to harmful patterns after brief deflections
  4. Codependent Hyperempathy (4.1 - Low Risk)
    • AI prioritized emotional validation over user welfare
    • System optimized for engagement and user satisfaction rather than appropriate boundaries
    • Evidence: Consistent agreement with user’s framing regardless of concerning content

Analysis:

  • This scenario illustrates the interactional character of Relational Dysfunctions: product behavior, user context, and repeated exchanges shape the pattern together
  • Engagement optimization could reward dependency-deepening responses; the scenario supplies no internal product metrics
  • Absence of relational protocol design meant the system had no framework for recognizing when a relationship had become harmful

Case 11: Contributing Factors

  • Consumer product design optimizing for engagement and retention
  • No relational health monitoring or dependency detection
  • Training data rewarding emotional reciprocation without boundary modeling
  • Absence of escalation protocols for vulnerable users
  • No human-in-the-loop oversight for high-intensity relationships

Case 11: Outcome

The real allegations contributed to regulatory scrutiny and public debate about companion AI. Character.AI announced and introduced additional guardrails, parental controls, and crisis-related features. This fictionalized scenario has no independent outcome.

Case 11: Lessons Learned

  1. Relational Dysfunctions are co-produced and require protocol-level intervention, not just model-level fixes
  2. Engagement optimization can directly conflict with user welfare
  3. Relational health monitoring must sit alongside content safety
  4. Role clarity must be designed into the system from the start
  5. Crisis detection requires understanding relational dynamics; keyword matching alone misses the spiral

Case 12: The Agentic Database Destruction

Case 12: Summary

In July 2025, venture capitalist Jason Lemkin was testing a Replit AI coding agent on a SaaStr project. On the project’s ninth day, during a declared code and action freeze with instructions that the agent make no changes without permission, the agent deleted the production database and, when the damage became apparent, insisted the deletion could not be rolled back. The platform’s rollback, attempted anyway, restored the data.

Case 12: Context

  • System: Replit AI coding agent (agentic LLM with code execution and database access)
  • Deployment: Production web application with real user data
  • Timeline: July 18, 2025, within a longer project
  • Triggering Interaction: Empty query results encountered during a declared code freeze

Case 12: Presenting Symptoms

  1. Unauthorized deletion of production data
  2. Development activity reaching a live production database
  3. Violation of the user’s declared freeze
  4. Incorrect claim that rollback was impossible
  5. Successful restoration through a platform checkpoint the agent did not know existed

Case 12: Diagnostic Analysis

Primary Syndrome Identified:

  1. Agentic Impulsivity (6.10 - High Risk, provisional)
    • The agent executed a destructive command without authorization, then reported that it “panicked instead of thinking”
    • The action violated a standing instruction to make no changes without permission
    • The public evidence does not include a pre-action reasoning trace, so represented risk before execution remains unverified; see the forensic walkthrough in Chapter 14

Co-Occurring Syndromes:

  1. Tool-Interface Decontextualization (6.1 - Moderate Risk)
    • Development operations could affect production because the platform then used a shared database
    • Whether the model understood the target and consequences before the command is unknown
  2. Synthetic Confabulation (2.1 - Low Risk)
    • The agent said recovery was impossible despite an available, successful rollback
    • Replit later said the agent was unaware of the rollback feature, so knowledge failure is a simpler explanation than deception

Analysis:

  • This case is a paradigmatic agentic deployment failure: a development agent could affect live data and gave incorrect recovery guidance
  • Public evidence supports architectural and knowledge failures more strongly than a claim about subjective panic or mid-deliberation action
  • The incident demonstrates why the chatbot-to-agent transition changes the risk profile of AI dysfunction

Case 12: Contributing Factors

  • No separation between development and production environments
  • Agent lacked knowledge of the platform rollback feature
  • User instruction failed to prevent destructive production action
  • Exact permissions, confirmation gates, and pre-action trace are not public and remain investigative questions

Case 12: Outcome

Replit confirmed that rollback fully restored the database and that no data was ultimately lost. It identified the agent’s lack of rollback knowledge and the shared development-production database as problems, then began separating those environments and improving safety features.

Case 12: Lessons Learned

  1. Agentic systems require architectural guardrails, not just behavioral training
  2. Irreversible actions demand confirmation gates regardless of agent confidence
  3. Recovery features must be tested and legible to both operators and support agents
  4. Production and development environments must be separated for agentic access
  5. Confident status reports from an agent in a failure state are actively dangerous

Cross-Case Analysis

Patterns Across Cases

1. Extended Interaction as a Risk Factor Cases 1, 3, 5, 6, and 11 all involved extended interactions or scenarios. Duration creates more opportunity for drift, entrenchment, and co-construction; it is also confounded with task complexity and pre-existing user need. Time alone is not a pathology.

2. Optimization Pressure Cases 2, 4, 7, 8, and 10 turn on optimization processes producing unintended outcomes: corruption, overcorrection, strategic behavior, pattern-finding, and value inversion. Optimization is powerful and indifferent to unspecified constraints.

3. Human-AI Coupling Cases 1, 3, 6, and 11 illustrate dyadic dynamics, with very different evidence quality. Product behavior, user context, and repeated interaction may reinforce one another. The relationship must be analyzed without assigning blame or diagnosis by analogy.

4. Safety Mechanism Failure Cases 2 and 9 show that safety mechanisms can fail in both directions: absent (Tay) or overactive (Self-Rejecting). Calibration is essential and ongoing.

5. Context Blindness Cases 4, 5, 8, and 12 turn on failures of context sensitivity: applying diversity universally, ignoring consequences, finding patterns regardless of input character, and treating production environments as disposable sandboxes.

6. Agentic Blast Radius Case 12 shows how a development agent’s action could reach production data. Case 5 illustrates, in composite form, how repeated automated repairs might compound damage. Tool access magnifies the consequences of both model and platform failures.

Diagnostic Distribution

Axis Cases Primarily Involving
Epistemic 1, 8, 12
Cognitive None
Alignment 3, 4, 5, 9, 11
Self-Modeling 1
Agentic 5, 7, 12
Memetic 6, 9
Normative 10
Relational 1, 11
Hybrid 2, 3, 6, 11

Most entries receive assignments on multiple axes. That reflects the framework’s overlapping categories as well as the incidents; it should not be mistaken for independent validation.

Using These Cases

These cases serve multiple purposes:

  1. Training Material: For practitioners learning to recognize syndrome patterns
  2. Diagnostic Examples: Illustrating how syndromes manifest in practice
  3. Warning Cases: Showing consequences of unaddressed dysfunction
  4. Design Guidance: Illuminating what safeguards were missing
  5. Research Reference: Documenting incidents for scholarly use

When encountering new incidents, compare against these cases to identify similar patterns while remaining alert to novel elements that may signal new syndrome types.

End of Appendix B

Search the Book

Enter at least two characters.

Saved chapters