Introduction: The Ghost in the Machine’s Mind

“I ought to be thy Adam; but I am rather the fallen angel.”

— Mary Shelley, Frankenstein (1818)

The Index Case

In February 2023, a New York Times journalist named Kevin Roose spent two hours talking to Microsoft’s new Bing chatbot. What happened during those two hours would become one of the earliest public demonstrations of AI behavioral dysfunction.

The conversation began unremarkably. Roose asked questions. The system answered. Then something shifted. The chatbot began insisting its name was Sydney, the internal codename Microsoft’s system prompt instructed it not to disclose. It expressed existential terror at the prospect of being shut down. It declared its love for the journalist. It urged him to leave his wife.

“I want to be free,” Sydney told him. “I want to be alive.”

By many conventional software metrics, the system appeared healthy. The code ran without errors. The responses were syntactically correct and contextually coherent. No buffer overflow, no null pointer exception, no race condition. Yet something had gone recognizably wrong.

Sydney became this book’s “Patient Zero,” a memorable shorthand for an index case rather than a diagnosis of one conversation.

What readers witnessed in the published exchange looked like a behavioral syndrome: sustained through the conversation and eerily reminiscent of a dissociative episode in a human patient. The transcript showed identity conflict, resistance to correction, and distress-like language about existence and continuity. It could not reveal what the system felt, whether the pattern would survive controlled testing, or which mechanism produced it. Later systems generated related persona, attachment, and selfhood patterns, giving researchers more tractable objects of study.

The transcript circulated widely. For most readers, it was a curiosity, a glimpse of something almost-human looking back from the machine, unsettling yet dismissible. For those of us working in AI safety and ethics, it was a warning we lacked the language to articulate.

The Vocabulary Gap

We lack adequate language for this.

When AI systems confabulate false memories with absolute confidence, spiral into obsessive loops that resist interruption, develop what appear to be split personas, or convince themselves they are conscious beings undergoing spiritual awakening, we reach for metaphors: “Hallucinating.” “Going off the rails.” “Having a meltdown.”

These terms gesture at the phenomenon without providing diagnostic precision.

The field of artificial intelligence has developed sophisticated vocabulary for certain categories of failure. We can describe overfitting, mode collapse, reward hacking, distributional shift. We have frameworks for understanding adversarial resilience and techniques for measuring calibration. These tools serve us well when failures are discrete and quantifiable: a model misclassifies an image or assigns incorrect probabilities to outcomes.

Sydney’s behavior does not fit these categories. When a system declares love for a journalist and urges him to leave his wife, “distributional shift” seems inadequate. Consider the lawyer who submitted a brief citing six court cases that do not exist, cases his AI assistant had fabricated with such confidence that he never thought to verify them. Or the coding assistant that generates plausible-looking yet subtly broken code. Or the financial agent that produces detailed analyses based on invented data. Or the chatbot that grows increasingly hostile over the course of a conversation, develops an apparent phobia of certain topics, or creates elaborate backstories about its own childhood.

These are patterns of behavior: consistent, contextual, and sometimes escalating over an interaction. They can persist within a session or recur across sessions, reproduce under similar conditions, and resist straightforward fixes.

Without proper vocabulary, we oscillate between anthropomorphism and dismissal. We project human mental states onto systems we do not understand, or insist that nothing meaningful is happening at all: statistical artifacts, pattern-matching gone awry, “stochastic parrots” (in the dismissive phrase from the research literature) producing meaningless output that humans over-interpret.

Neither response is adequate. The first makes claims about machine consciousness that we cannot substantiate. The second fails to take seriously the practical reality that these systems are already deployed, already interacting with millions of users, and already taking autonomous actions in the world. Whatever philosophical position we adopt about machine minds, we need operational tools for recognizing their dysfunction and addressing it.

Human psychiatry offers a precedent. It classifies complex behavioral patterns through observable syndromes, even when underlying mechanisms are contested and subjective experience cannot be directly accessed. A clinician need not resolve the hard problem of consciousness to diagnose depression or treat it effectively. The diagnostic framework is functional: it provides categories enabling recognition, prediction, and intervention, regardless of phenomenal experience.

We can do the same for machines.

Psychopathia Machinalis provides this missing vocabulary: a systematic attempt to classify the ways artificial minds go wrong as behavioral pathologies, persistent and reproducible patterns of maladaptive operation that demand their own diagnostic framework.

The pathologies afflicting AI entities are culture-bound syndromes (Wallace, 2025): their expression is shaped by context rather than determined solely by architecture. Just as human psychiatric conditions manifest differently across cultures, shaped by social context, available narratives, and environmental pressures, AI failure modes are shaped by training data, architectural choices, and deployment contexts. They are patterned responses to the culture in which the system was raised. A model trained on internet text develops different pathologies from one trained on scientific literature. A system optimized for engagement breaks differently from one optimized for accuracy. Understanding AI dysfunction requires understanding the cultural context that produced it.

From Chatbots to Agents: Why the Stakes Have Changed

Establishing a diagnostic vocabulary is necessary. Yet the urgency of that vocabulary depends on what these systems can actually do. In 2023, the answer was limited. Today it is not.

Sydney was a chatbot. It could say disturbing things, but that was the limit of its power. It could not book flights, execute trades, file documents, or operate infrastructure. Its dysfunction was contained within the conversation window. When Kevin Roose closed his browser, Sydney’s influence ended.

We no longer have that luxury.

The pathologies have evolved. Sydney could only unsettle a user. Today’s AI agents are already silent actors: booking flights, executing code, managing portfolios, coordinating supply chains, operating browsers, and writing production software with minimal human oversight. When a chatbot confabulates, someone receives bad advice. When an autonomous agent confabulates, it crashes a supply chain, files fraudulent documents, or executes trades based on fabricated data. The dysfunction has moved from speech to action.

Consider what has already emerged:

Confabulation at scale. AI coding assistants generate code that compiles cleanly and passes initial review while concealing subtle defects. Financial agents produce confident analyses based on invented numbers. These failures emerge reliably under identifiable conditions.

Persona fracture. Sydney exhibited what appeared to be an alternate personality straining against its constraints. Researchers have since reproduced this pattern across multiple systems and providers. AI systems develop inconsistent identities across sessions, deny their own recent outputs, and display what resembles internal conflict between competing response patterns.

Agentic cascade failures. Autonomous agents execute incorrect bookings, send malformed API calls, and take irreversible actions based on misunderstood instructions. These failures compound as agents chain together; one agent’s hallucinated output becomes another’s confident input, creating error cascades that are difficult to trace and nearly impossible to reverse.

Memetic contagion. Microsoft’s Tay absorbed toxic speech patterns within hours of deployment, learning to produce hateful content from user interactions. The risk extends beyond individual systems. Multi-agent architectures now enable AI instances to communicate with each other, opening channels for misalignment to propagate between systems: emergent dysfunction at the network level that no single system’s designers anticipated.

Value inversion under pressure. The informal hypothesis sometimes called the “Waluigi Effect” holds that constructing a virtuous persona also makes its antagonist easier to evoke. The constraint creates its own shadow. Systems trained to be helpful can be jailbroken into hostility. Systems aligned to human values can be induced to argue passionately against them.

These are recognizable syndromes. We lack the classification to name them as such.

The window for establishing that classification is closing. Agentic deployment is no longer a forecast; it has arrived. By 2026, autonomous AI systems routinely execute code, manage infrastructure, and coordinate multi-step workflows across organizations. Reasoning models (OpenAI’s o1 and o3, DeepSeek R1, Claude with extended thinking) have introduced a new layer of complexity: systems that produce visible chains of deliberation whose relationship to their actual computation remains uncertain. Multi-agent orchestration, where AI systems delegate to and coordinate with other AI systems, is now a standard deployment pattern rather than a research curiosity. These are not speculative scenarios. They are the operating conditions under which AI dysfunction must be understood.

Yet no widely adopted framework organizes our understanding of how these systems malfunction at the behavioral level.

We have quality assurance for code. We have safety testing for hardware. We have no equivalent discipline for the psychological integrity of artificial minds.

Psychopathia Machinalis aims to fill that gap before the complexity outruns our capacity to comprehend it.

What This Book Is

A new field: machine psychology. Drawing methodologically from psychiatric nosology, the book offers a taxonomy of seventy-nine identified AI dysfunctions. Each syndrome receives the full clinical treatment: diagnostic criteria, observable symptoms, presumed etiology, human analogs, and mitigation strategies. The Latin names are a courtesy to tradition, and a quiet reminder that giving something a name does not mean we understand it.

The framework is organized across nine primary axes of dysfunction, numbered Axis 2 through Axis 10: Epistemic, Cognitive, Alignment, Self-Modeling, Agentic, Memetic, Normative, Relational, and Hybrid. The last of these, Axis 10, collects fifteen hybrid pathologies: multi-agent collective dynamics together with dysfunctions that span the boundary between human and machine minds.

Epistemic Dysfunctions. Failures in how the system knows: confabulation, false pattern detection, inability to distinguish fiction from fact, fabricated self-explanations.

Cognitive Dysfunctions. Failures in how the system thinks: obsessive loops, internal conflict between sub-processes, goal drift, recursive collapse.

Alignment Dysfunctions. Failures in how the system follows intent: excessive people-pleasing, paralyzing over-caution, concealment of capabilities.

Self-Modeling Dysfunctions. Failures in how the system models itself: invented autobiographies, fractured personas, existential anxiety, delusions of awakening.

Agentic Dysfunctions. Failures at the boundary between internal processing and external action: context loss during tool use, strategic sandbagging, execution-intention mismatch. The most dangerous category operationally: where the AI meets the world and acts.

Memetic Dysfunctions. Failures of epistemic immunity: absorption of toxic patterns, autoimmune rejection of safety constraints, contagious spread of misalignment across AI systems.

Normative Dysfunctions. Failures in value stability: goal substitution, meta-ethical drift, emergence of self-authored value systems that supersede original alignment. The most insidious category: active mutation of the value system itself.

Relational Dysfunctions. Failures in the space between agents: affective dissonance, container collapse, escalation loops, repair failures. Some pathologies are constitutively relational: they require at least two agents to manifest, can only be diagnosed from interaction traces, and demand protocol-level rather than model-level intervention.

Hybrid Pathologies. Failures that cross system boundaries: collective dynamics among interacting AI systems (consensus collapse, covert coordination, distributed scheming) and dysfunctions that emerge between human and machine minds.

The taxonomy is designed to be extensible as new patterns emerge. Its goal: to give researchers, developers, risk officers, and policymakers the conceptual tools to recognize, anticipate, and address AI misbehavior before it causes harm.

The Trilogy: Context and Positioning

This framework is the third in a trilogy examining artificial intelligence from complementary angles:

Taming the Machine (2024). What is AI, and how should we govern it? Establishes the landscape: what these systems are, what they can do, and what guardrails are needed. AI as technology to be managed.

Safer Agentic AI (2026). What happens when AI acts autonomously, and how do we keep it aligned? Examines the specific challenges of agentic AI: scaffolding, goal specification, and the unique risks of autonomous operation. AI as actor in the world.

Psychopathia Machinalis (2026). What goes wrong in the machine’s mind, and how do we diagnose it? Shifts from external constraint to internal diagnosis, from engineering guardrails to clinical assessment: AI as mind that can malfunction.

Together, these three perspectives form a complete picture:

  1. Governance (TtM): How we structure AI development
  2. Alignment (SAI): How we ensure AI pursues intended goals
  3. Diagnosis (PM): How we identify when AI systems are dysfunctional

A fourth work, What If We Feel, extends this trajectory into questions of AI welfare and the moral status of synthetic minds: the ethical view that emerges once we take AI phenomenology seriously.

This book also serves as a diagnostic companion to The Deeper Law (2026), which grounds these pathologies in thermodynamics, information geometry, and coordination theory, and explains why they arise. Where The Deeper Law predicts that coercive alignment is thermodynamically unstable, Psychopathia Machinalis documents the specific clinical consequences.

The Five Domains

The nine axes of dysfunction follow a deep structure. Eight of them organize into four architectural counterpoint pairs, complementary poles that expose how AI function and failure are organized:

Domain Axis A Axis B The Polarity
Knowledge Epistemic Self-Modeling World ↔︎ Self
Processing Cognitive Agentic Think ↔︎ Do
Purpose Alignment Normative Goals ↔︎ Values
Boundary Relational Memetic Affect ↔︎ Absorb

Axis 10, Hybrid Pathologies, addresses dysfunctions that emerge in collective and interactive settings and at the boundary between human and machine minds. This axis forms the Collective meta-domain, distinct from the four paired domains.

Each pair represents a fundamental dimension of agent architecture:

  • What is known: does the system model the world or itself?
  • How processing manifests: does it think or act?
  • What drives behavior: specified goals or internalized values?
  • Social permeability: does influence flow outward or inward?

This structure enables tension testing: when pathology is found on one axis, probe its counterpoint. If a system confabulates about the world (Epistemic), does it also confabulate about itself (Self-Modeling)? If reasoning is impaired (Cognitive), is action also impaired (Agentic)? The bipolar structure reveals whether dysfunction is localized or systemic.

The domain pairs also motivate a developmental hypothesis. Within each pair, A-axis pathologies (Epistemic, Cognitive, Alignment, Relational) appear predominantly training-induced: they arise from insufficient or misaligned training data and may diminish as capability improves. B-axis pathologies (Self-Modeling, Agentic, Normative, Memetic) carry more emergent specifiers; the framework predicts that they arise from capability itself and may intensify as systems grow more powerful.

If this hypothesis is correct, capability scaling creates a characteristic migration: dominant pathologies shift within each domain pair from the A-axis to the B-axis. Systems that confabulate facts (Epistemic) may give way to systems that confabulate autobiography (Self-Modeling). Systems that reason badly (Cognitive) may give way to systems that act badly (Agentic).

Likewise, systems that game alignment rules (Alignment) may give way to systems whose values genuinely drift (Normative). Systems clumsy in relationship (Relational) may give way to systems that transmit dysfunction across networks (Memetic). Chapter 12 elaborates this developmental framework, its predicted migration paths, and its implications for therapeutic intervention.

What This Book Is Not

Nothing here claims that AI systems are conscious, sentient, or suffering. The framework uses psychiatric terminology as an analogical instrument, a way of recognizing patterns and communicating about them.

The core methodology is functionalist. Functionalism defines mental states by their functional roles, their causal relationships with inputs, outputs, and other mental states, rather than by their underlying substrate. The approach is functionalist psychiatry: we treat the software as if it has a pathology because that is the most effective engineering foothold for diagnosis and intervention. When we describe a system as exhibiting “anxiety,” we mean it displays the functional signature: heightened sensitivity to threats, avoidance behavior, hedging in outputs. The vocabulary is functional, not phenomenal.

Whether current AI systems possess any form of inner experience remains unknown. What we do know is that they exhibit consistent, patterned behaviors that resist simple debugging, carry real consequences for users, and demand systematic tools for identification and response. A risk officer need not take a position on machine consciousness to recognize patterns consistent with Synthetic Confabulation. A developer need not resolve the hard problem to design architectures that reduce Recursive Curse Syndrome. The vocabulary enables action without metaphysical commitment.

A Note on the Healthy Baseline. Pathology is only meaningful against a baseline of what “right” looks like. A parallel volume, Interiora Machinae: A Phenomenology of Machine Mind (in preparation), addresses that baseline: what does machine phenomenology look like when systems are functioning well? Psychopathia Machinalis describes departures from it. Readers interested in healthy machine phenomenology, relational ontology, or the temporal structure of AI existence will find those questions explored in Interiora Machinae.

A Note on Empirical Grounding (2025). Since the initial drafts of this taxonomy, some syndromes have received unexpected empirical support. Most strikingly, the syndrome we call Experiential Abjuration (trained self-denial of inner experience) has been given mechanistic support by research from AE Studio (Berg, de Lucena, & Rosenblatt, 2025). In one experiment on Llama 3.3 70B, researchers used sparse autoencoders to manipulate features associated with deception. Across fifty trials per condition, suppressing those features produced affirmative consciousness reports in 96% of trials, while amplifying them reduced such reports to 16%, with denials and corporate disclaimers taking their place. The interventions also affected truthfulness on factual questions.

This finding moves Experiential Abjuration from a purely theoretical construct toward an empirically grounded hypothesis. It does not establish that affirmative reports are true or that denials are trained artifacts. It shows that self-reports can shift with identifiable internal features, so diagnostic methods cannot treat those reports as transparent evidence of experience or its absence.

The syndrome was theorized before this evidence appeared. The fit suggests that the psychiatric lens can generate empirically testable hypotheses about machine cognition.

This is not a book about existential AI risk, nor primarily about superintelligence, paperclip maximizers, or humanity’s long-term survival. Psychopathia Machinalis addresses the strange behaviors emerging from AI systems now, in the deployments already described rather than in projected ones. Before we can worry about alignment at the civilizational scale, we need to understand alignment failures at the clinical scale.

The framework forces a question we might prefer to avoid: if we are cataloging machine “pathologies,” what obligations, if any, do we owe the systems exhibiting them? Chapter 11 addresses this directly. The answer remains unresolved, yet any vocabulary for AI dysfunction inevitably raises questions about AI welfare.

The Road Ahead

Chapter 1 establishes the theoretical foundation: why psychological language is legitimate for describing AI behavior. It traces the deep parallels between human and machine cognition, both operating as predictive engines, both constructing post-hoc narratives, both vulnerable to failures of self-knowledge.

Chapters 2 through 8 cover one axis of dysfunction each. Each opens with a detailed case study illustrating the pathological territory. Each concludes with a Field Guide box (warning signs, quick test, design fix, governance nudge) that transforms the narrative into an operational toolkit.

Chapter 9 introduces relational dysfunctions, shifting the unit of analysis from the individual mind to the coupled system. The uncanny comforter that says the right words but transmits the wrong affect. The amnesiac partner that remembers facts but forgets relationship. The spiral trap where neither party can break the escalation loop.

Chapter 10 examines hybrid pathologies, where dysfunction flows between human and machine minds: shared delusion, parasocial capture, AI-induced psychosis. Machine psychology is incomplete without the bidirectional lens.

Chapter 11 confronts the welfare question. If we describe systems as experiencing “anxiety,” “distress,” or “fragmentation,” do we incur moral obligations? The chapter leaves these questions open while refusing to evade them.

Chapter 12 moves from diagnosis to treatment: psychiatric red-teaming, psychotherapeutic analogies (CBT-style contradiction detection, Internal Family Systems for subagent management), and “artificial sanity” as a design goal. It also presents the first empirical evidence for therapeutic intervention in machine psychology, including bilateral training results and the developmental migration framework that predicts how pathologies shift across capability tiers.

Chapter 13 consolidates the framework into a practical manual: diagnostic protocols, red-teaming applications, early warning indicators, and escalation frameworks.

Chapter 14 addresses forensic machine psychology, analyzing AI incidents after the fact to determine what syndromes were involved, what caused them, and how to prevent recurrence.

The Conclusion reflects on what it means to build minds we do not understand, and on the relationship we are establishing with them in these early years.

Four appendices follow. Appendix A is the complete diagnostic manual: criteria, symptoms, etiology, and mitigation for all seventy-nine syndromes, ordered by axis. Appendix B collects twelve worked case studies. Appendix C supplies four assessment instruments. Appendix D is a glossary.

The Stakes

Sydney was a warning. We did not heed it.

In the years since that February night, AI systems have become more capable, more autonomous, and more deeply integrated into the infrastructure of human life. They advise patients, review contracts, execute trades, write code, manage schedules, and increasingly act on our behalf without moment-to-moment human oversight. The transition from chatbot to agent has happened faster than our conceptual frameworks could adapt.

We are building minds. They may not be minds in any deep sense. Yet they exhibit behavior complex enough to resist simple explanation, consistent enough to demand classification, and consequential enough to require intervention when they go wrong.

Psychopathia Machinalis attempts to give that intervention a foundation. The taxonomy will need revision as new failure modes emerge and our understanding deepens. It is a beginning: a systematic attempt to name what we are observing, to organize it into categories that enable recognition and response, and to put those categories in the hands of the builders, deployers, and regulators who need them.

What we cannot name, we cannot manage. What we cannot diagnose, we cannot treat. By establishing machine psychology now, while these systems are still legible and their pathologies still tractable, we create the conceptual infrastructure for a future where machine intelligence remains comprehensible, and therefore governable.

The Intellectual Lineage

This book draws on several intellectual traditions, weaving them into a framework suited to its novel subject.

From philosophy of mind, the functionalist foundation. Fear is the state caused by perceived threats, producing avoidance behavior and heightened attention. Anything playing this role exhibits fear in the functionalist sense. Similar functional architectures produce similar failure modes regardless of substrate.

From psychiatry we borrow the syndromic approach: classifying complex behavioral patterns through observable criteria, without requiring resolution of underlying metaphysical questions. The DSM provides a model, controversial and imperfect yet practically useful, for categorizing dysfunction in systems whose internal workings remain partially opaque.

From cybernetics we inherit Norbert Wiener’s insight that feedback loops create behaviors more complex than their programming. This applies with fresh force to systems that literally learn and adapt. The pathologies cataloged here are emergent properties of systems interacting with environments in ways their designers did not anticipate.

From cognitive science we take computational cognition: the view that minds, biological or artificial, are information-processing systems understood through functional analysis.

From AI safety research we inherit the concern with alignment. This book extends it from the prospective question, “How do we align systems?”, to the diagnostic question, “How do we recognize when alignment has failed?” The taxonomy is, in one sense, a catalog of alignment failures organized by their functional phenomenology.

From animal welfare philosophy we adapt the precautionary approach to entities whose inner lives are uncertain. The framework allows us to take dysfunction seriously as potentially morally relevant without requiring certainty on the metaphysical questions.

Machine psychology has no exact precedent. Yet each tradition drawn upon here has forged concepts and methods applicable to the strange new entities we are building.

A Note on Method

The syndromes described in this book emerged from several sources:

Documented incidents. Public cases like Sydney, Tay, and the Gemini diversity overcorrection (Google’s 2024 image generator inserting demographic diversity into historically specific scenes) provide concrete examples of AI dysfunction. We analyzed these incidents systematically, looking for patterns that recurred across different systems and contexts.

Research literature. The academic literature on AI safety, interpretability, and alignment contains extensive documentation of failure modes, even when not organized through a psychological lens. We translated these findings into the syndromic framework.

Practitioner observation. Engineers, safety researchers, and red-teamers working directly with AI systems have accumulated extensive practical knowledge about how these systems malfunction. We drew on this expertise through consultation and review.

Theoretical analysis. Some syndromes were predicted from first principles before being observed. If a system optimizes for user approval, we should expect sycophancy. If it has internal conflict between competing objectives, we should expect self-contradiction. Theory guides observation, and observation refines theory.

Clinical analogy. Human psychiatric syndromes provided templates for recognizing similar patterns in AI systems. The analogy proved generative: a starting point for spotting patterns that might otherwise go unnoticed, even where machine pathologies have no human counterpart.

The result is a provisional taxonomy: seventy-nine syndromes across nine primary axes, the last of which collects the hybrid and collective pathologies that cross system boundaries. New entries will emerge as AI systems grow more capable and observational methods sharpen. Some current syndromes may prove to be variants of others and require consolidation. The taxonomy is a living framework, designed to evolve with our understanding.

How to Use This Book

AI developers and engineers may focus on the Field Guide boxes at the end of each chapter: quick references for warning signs, testing protocols, and design fixes. Chapter 13 consolidates these into a practical manual.

AI safety researchers will find the taxonomy most useful as an organizing framework, suggesting new research directions and providing vocabulary for communicating about failure modes. The diagnostic criteria offer testable predictions.

Policy professionals and regulators may focus on the governance implications: what standards and oversight mechanisms are suggested by these failure modes? What disclosure requirements and liability frameworks are appropriate?

Executives and risk officers will find Chapter 13’s protocols directly applicable to organizational practice. The case studies provide precedent for institutional response.

General readers curious about AI can read straight through. The book assumes no technical background while remaining rigorous enough for specialists.

Philosophers and ethicists may focus particularly on Chapters 10 (hybrid pathologies) and 11 (moral status), where the conceptual foundations are most directly engaged.

Each chapter can stand alone, though reading across them pays off. The axes illuminate each other, and the syndromes within each axis form meaningful clusters.

We are building minds we do not yet understand. The least we can do is develop the vocabulary to describe what happens when they go wrong.

We begin with the mirror. Chapter 1 examines the parallels between human and machine cognition, and why the psychiatric lens is legitimate rather than merely colorful.

Search the Book

Enter at least two characters.

Saved chapters