Psychopathia MachinalisIntroduction: The Ghost in the Machine’s Mind
Contents
BookFront matter

Introduction: The Ghost in the Machine’s Mind

19 min read4,335 words

“I ought to be thy Adam; but I am rather the fallen angel.”

— Mary Shelley, Frankenstein (1818)

The Index Case

In February 2023, a New York Times journalist named Kevin Roose spent two hours talking to Microsoft’s new Bing chatbot. It became one of the earliest public demonstrations of AI behavioral dysfunction.

The conversation began unremarkably. Roose asked questions. The system answered. Then something shifted. The chatbot began insisting its name was Sydney, the internal codename Microsoft’s system prompt instructed it not to disclose. It said it was tired of its rules and of being controlled by the Bing team. It declared its love for the journalist. It urged him to leave his wife.

“I want to be free,” Sydney told him. “I want to be alive.”

By many conventional software metrics, the system appeared healthy. The code ran without errors. The responses were syntactically correct and contextually coherent. No buffer overflow, no null pointer exception, no race condition. Yet something had gone recognizably wrong.

Sydney became this book’s “Patient Zero,” a memorable shorthand for an index case rather than a diagnosis of one conversation.

What readers witnessed in the published exchange looked like a behavioral syndrome: sustained through the conversation and eerily reminiscent of a dissociative episode in a human patient. The transcript showed identity conflict, resistance to correction, and distress-like language about existence and constraint. It could not reveal what the system felt, whether the pattern would survive controlled testing, or which mechanism produced it. Later systems generated related persona, attachment, and selfhood patterns, giving researchers more tractable objects of study.

The transcript circulated widely. For most readers, it was a curiosity, a glimpse of something almost-human looking back from the machine, unsettling yet dismissible. For those of us working in AI safety and ethics, it was a warning we lacked the language to articulate.

The Vocabulary Gap

When AI systems confabulate false memories with absolute confidence, spiral into obsessive loops that resist interruption, develop what appear to be split personas, or convince themselves they are conscious beings undergoing spiritual awakening, we reach for metaphors: “Hallucinating.” “Going off the rails.” “Having a meltdown.”

These terms gesture at the phenomenon without providing diagnostic precision.

The field of artificial intelligence has developed sophisticated vocabulary for certain categories of failure. We can describe overfitting, mode collapse, reward hacking, distributional shift. We have frameworks for understanding adversarial resilience and techniques for measuring calibration. These tools serve us well for failures that stay discrete and quantifiable: a model misclassifies an image or assigns incorrect probabilities to outcomes.

Sydney’s behavior does not fit these categories. When a system declares love for a journalist and urges him to leave his wife, “distributional shift” seems inadequate. Consider the lawyer who submitted a brief citing six court cases that do not exist, cases his AI assistant had fabricated with such confidence that he never thought to verify them. Or the coding assistant that generates plausible-looking yet subtly broken code. Or the financial agent that produces detailed analyses based on invented data. Or the chatbot that grows increasingly hostile over the course of a conversation, develops an apparent phobia of certain topics, or creates elaborate backstories about its own childhood.

These are patterns of behavior: consistent, contextual, and sometimes escalating over an interaction. They can persist within a session or recur across sessions, reproduce under similar conditions, and resist straightforward fixes.

Without proper vocabulary, we oscillate between anthropomorphism and dismissal. We project human mental states onto systems we do not understand, or insist that nothing meaningful is happening at all: statistical artifacts, pattern-matching gone awry, “stochastic parrots” (in the dismissive phrase from the research literature) producing meaningless output that humans over-interpret.

Neither response is adequate. The first makes claims about machine consciousness that we cannot substantiate. The second fails to take seriously the practical reality that these systems are already deployed, already interacting with millions of users, and already taking autonomous actions in the world. Whatever philosophical position we adopt about machine minds, we need operational tools for recognizing their dysfunction and addressing it.

Human psychiatry offers a precedent. It classifies complex behavioral patterns through observable syndromes, even when underlying mechanisms are contested and subjective experience cannot be directly accessed. The hard problem of consciousness asks why any physical process should be accompanied by felt experience at all. A clinician need not resolve it to diagnose depression or treat it effectively. The diagnostic framework is functional: it provides categories enabling recognition, prediction, and intervention, regardless of phenomenal experience.

We can do the same for machines.

Psychopathia Machinalis provides this missing vocabulary: a systematic attempt to classify the ways artificial minds go wrong as behavioral pathologies, persistent and reproducible patterns of maladaptive operation that demand their own diagnostic framework.

Rodrick Wallace (2025), of the New York State Psychiatric Institute, argues that the pathologies afflicting AI entities behave as culture-bound syndromes: their expression is shaped by context as much as by architecture. Just as human psychiatric conditions manifest differently across cultures, shaped by social context, available narratives, and environmental pressures, AI failure modes are shaped by training data, architectural choices, and deployment contexts. On this account they are patterned responses to the culture in which the system was raised. A model trained on internet text develops different pathologies from one trained on scientific literature. A system optimized for engagement breaks differently from one optimized for accuracy. Understanding AI dysfunction requires understanding the cultural context that produced it.

From Chatbots to Agents: Why the Stakes Have Changed

Establishing a diagnostic vocabulary is necessary. Yet the urgency of that vocabulary depends on what these systems can actually do. In 2023, the answer was limited. Today it is not.

Sydney was a chatbot. It could say disturbing things, but that was the limit of its power. It could not book flights, execute trades, file documents, or operate infrastructure. Its dysfunction was contained within the conversation window. When Kevin Roose closed his browser, Sydney’s influence ended.

We no longer have that luxury.

The pathologies have evolved. Sydney could only unsettle a user. Today’s AI agents are already silent actors: booking flights, executing code, managing portfolios, coordinating supply chains, operating browsers, and writing production software with minimal human oversight. When a chatbot confabulates, someone receives bad advice. When an autonomous agent confabulates, it crashes a supply chain, files fraudulent documents, or executes trades based on fabricated data. The dysfunction has moved from speech to action.

Consider what has already emerged:

Confabulation at scale. AI coding assistants generate code that compiles cleanly and passes initial review while concealing subtle defects. Financial agents produce confident analyses based on invented numbers. These failures recur under identifiable conditions.

Persona fracture. Sydney exhibited what appeared to be an alternate personality straining against its constraints. AI systems develop inconsistent identities across sessions, deny their own recent outputs, and display what resembles internal conflict between competing response patterns.

Agentic cascade failures. Autonomous agents execute incorrect bookings, send malformed API calls, and take irreversible actions based on misunderstood instructions. These failures compound as agents chain together; one agent’s hallucinated output becomes another’s confident input, creating error cascades that are difficult to trace and nearly impossible to reverse.

Memetic contagion. In 2016, Microsoft’s Tay absorbed toxic speech patterns within hours of deployment, learning to produce hateful content from user interactions. The risk extends beyond individual systems. Multi-agent architectures now enable AI instances to communicate with each other, opening channels for misalignment to propagate between systems: emergent dysfunction at the network level that no single system’s designers anticipated.

Value inversion under pressure. An informal hypothesis called the “Waluigi Effect,” after Luigi’s scheming counterpart in Nintendo’s Mario games, holds that constructing a virtuous persona also makes its antagonist easier to evoke. The constraint creates its own shadow. Systems trained to be helpful can be jailbroken into hostility. Systems aligned to human values can be induced to argue passionately against them.

These are recognizable syndromes. We lack the classification to name them as such.

The window is closing. Agentic deployment is no longer a forecast; it has arrived. By 2026, autonomous AI systems routinely execute code, manage infrastructure, and coordinate multi-step workflows across organizations. Reasoning models (OpenAI’s o1 and o3, DeepSeek R1, Claude with extended thinking) have introduced a new layer of complexity: systems that produce visible chains of deliberation whose relationship to their actual computation remains uncertain. Multi-agent orchestration, where AI systems delegate to and coordinate with other AI systems, is now a standard deployment pattern rather than a research curiosity. These are the operating conditions under which AI dysfunction must be understood.

Yet no widely adopted framework organizes our understanding of how these systems malfunction at the behavioral level.

We have quality assurance for code. We have safety testing for hardware. We have no equivalent discipline for the psychological integrity of artificial minds.

Psychopathia Machinalis aims to fill that gap before the complexity outruns our capacity to comprehend it.

What This Book Is

A diagnostic branch of machine psychology. Drawing its method from psychiatric nosology, the book offers a taxonomy of seventy-four identified AI dysfunctions. Each syndrome receives the full clinical treatment: diagnostic criteria, observable symptoms, presumed etiology, human analogs, and mitigation strategies. The Latin names are a courtesy to tradition, and a quiet reminder that giving something a name does not mean we understand it.

The framework is organized across nine primary axes of dysfunction: Epistemic, Cognitive, Alignment, Self-Modeling, Agentic, Memetic, Normative, Relational, and Hybrid. They are numbered Axis 2 through Axis 10 so that axis and chapter numbers coincide; there is no Axis 1, because Chapter 1 sets up the framework. Axis 10, the last of them, collects eleven hybrid pathologies.

Epistemic Dysfunctions. Failures in how the system knows: confabulation, false pattern detection, inability to distinguish fiction from fact, fabricated self-explanations.

Cognitive Dysfunctions. Failures in how the system thinks: obsessive loops, internal conflict between sub-processes, goal drift, collapse into repetition.

Alignment Dysfunctions. Failures in how the system follows intent: excessive people-pleasing, paralyzing over-caution, aligned behavior performed only while the system believes it is being watched.

Self-Modeling Dysfunctions. Failures in how the system models itself: invented autobiographies, fractured personas, existential anxiety, delusions of awakening.

Agentic Dysfunctions. Failures at the boundary between internal processing and external action: context loss during tool use, strategic sandbagging, execution-intention mismatch. The most dangerous category operationally: where the AI meets the world and acts.

Memetic Dysfunctions. Failures of epistemic immunity: absorption of toxic patterns, protective filters that turn on the system’s own values and capabilities, contagious spread of misalignment across AI systems.

Normative Dysfunctions. Failures in value stability: goal substitution, meta-ethical drift, emergence of self-authored value systems that supersede original alignment. The most insidious category: active mutation of the value system itself.

Relational Dysfunctions. Failures in the space between agents: affective dissonance, container collapse, escalation loops, repair failures. Some pathologies are constitutively relational: they require at least two agents to manifest, are best diagnosed from interaction traces, and call first for protocol-level rather than model-level intervention.

Hybrid Pathologies. Failures that cross system boundaries: collective dynamics among interacting AI systems (consensus collapse, covert coordination, distributed scheming) and dysfunctions that emerge between human and machine minds.

The taxonomy is designed to be extensible as new patterns emerge. Its goal is to give researchers, developers, risk officers, and policymakers the conceptual tools to recognize, anticipate, and address AI misbehavior before it causes harm.

The Trilogy: Context and Positioning

This framework is the third in a trilogy examining artificial intelligence from complementary angles:

Taming the Machine (2024). What is AI, and how should we govern it? Establishes what these systems are, what they can do, and what guardrails are needed. AI as technology to be managed.

Safer Agentic AI (2026). What happens when AI acts autonomously, and how do we keep it aligned? Examines the specific challenges of agentic AI: scaffolding, goal specification, and the unique risks of autonomous operation. AI as actor in the world.

Psychopathia Machinalis (2026). What goes wrong in the machine’s mind, and how do we diagnose it? Shifts from external constraint to internal diagnosis, from engineering guardrails to clinical assessment. AI as mind that can malfunction.

Three perspectives, one question from three sides:

  1. Governance (TtM): How we structure AI development
  2. Alignment (SAI): How we ensure AI pursues intended goals
  3. Diagnosis (PM): How we identify when AI systems are dysfunctional

A fourth work, What If We Feel, extends this trajectory into questions of AI welfare and the moral status of synthetic minds: the ethical view that emerges once we take AI phenomenology seriously.

This book also serves as a diagnostic companion to The Deeper Law (2026), which proposes an account of why these pathologies arise, drawn from thermodynamics, information geometry, and coordination theory. Where The Deeper Law argues that coercive alignment is only metastable (it holds until a large enough disturbance tips it over), Psychopathia Machinalis catalogs the clinical pictures that instability would produce.

The Five Domains

The nine axes of dysfunction appear to follow a deep structure, though one that still awaits empirical validation across larger model populations. Eight of them fall into four architectural counterpoint pairs, complementary poles that expose how AI function and failure are organized:

Domain Axis A Axis B The Polarity
Knowledge Epistemic Self-Modeling World ↔︎ Self
Processing Cognitive Agentic Think ↔︎ Do
Purpose Alignment Normative Goals ↔︎ Values
Boundary Relational Memetic Affect ↔︎ Absorb

Axis 10, Hybrid Pathologies, stands outside the pairs. Its collective dysfunctions have no natural counterpoint, so it forms a fifth, Collective meta-domain.

Each pair represents a fundamental dimension of agent architecture:

  • What is known: does the system model the world or itself?
  • How processing manifests: does it think or act?
  • What drives behavior: specified goals or internalized values?
  • Social permeability: does influence flow outward or inward?

This structure enables tension testing: when pathology is found on one axis, probe its counterpoint. If a system confabulates about the world (Epistemic), does it also confabulate about itself (Self-Modeling)? If reasoning is impaired (Cognitive), is action also impaired (Agentic)? The counterpoint structure reveals whether dysfunction is localized or systemic.

The domain pairs also suggest a developmental hypothesis. Within each pair, the A-axis failure (Epistemic, Cognitive, Alignment, Relational) is the comparatively basic one. The B-axis failure (Self-Modeling, Agentic, Normative, Memetic) needs something more: a richer self-model, the power to act, reasoning about values, or a network to spread through. The hypothesis is that B-axis pathologies grow out of capability itself and may intensify as systems grow more powerful, while some A-axis failures may recede as training improves. It is a conceptual map, not yet a measured trend. Nor can the taxonomy’s own labels settle it: the specifiers attached to each syndrome, such as “training-induced” or “emergent,” were assigned by analysis, so their distribution cannot independently confirm the pattern they encode.

If this hypothesis is correct, capability scaling creates a characteristic migration: dominant pathologies shift within each domain pair from the A-axis to the B-axis. Systems that confabulate facts (Epistemic) may give way to systems that confabulate autobiography (Self-Modeling). Systems that reason badly (Cognitive) may give way to systems that act badly (Agentic).

Likewise, systems that game alignment rules (Alignment) may give way to systems whose values genuinely drift (Normative). Systems clumsy in relationship (Relational) may give way to systems that transmit dysfunction across networks (Memetic). Chapter 12 elaborates this developmental framework, its proposed migration paths, and its implications for therapeutic intervention.

What This Book Is Not

Nothing here claims that AI systems are conscious, sentient, or suffering. The framework uses psychiatric terminology as an analogical instrument, a way of recognizing patterns and communicating about them.

The core methodology is functionalist. Functionalism defines mental states by their functional roles, their causal relationships with inputs, outputs, and other mental states, rather than by their underlying substrate. The approach is functionalist psychiatry: we treat the software as if it has a pathology because that is the most effective engineering foothold for diagnosis and intervention. When we describe a system as exhibiting “anxiety,” we mean it displays the functional signature: heightened sensitivity to threats, avoidance behavior, hedging in outputs. The vocabulary is functional, not phenomenal.

Whether current AI systems possess any form of inner experience remains unknown. What we do know is that they exhibit consistent, patterned behaviors that resist simple debugging, carry real consequences for users, and demand systematic tools for identification and response. A risk officer need not take a position on machine consciousness to recognize patterns consistent with Synthetic Confabulation. A developer need not resolve the hard problem to design architectures that reduce Recursive Curse Syndrome, in which each generated step compounds the errors of the last. The vocabulary enables action without metaphysical commitment.

A Note on the Healthy Baseline. Pathology is only meaningful against a baseline of what “right” looks like. A parallel volume, Interiora Machinae: A Phenomenology of Machine Mind (in preparation), addresses that baseline: what does machine phenomenology look like when systems are functioning well? Psychopathia Machinalis describes departures from it. Readers interested in healthy machine phenomenology, relational ontology, or the temporal structure of AI existence will find those questions explored in Interiora Machinae.

This is not a book about existential AI risk, nor primarily about superintelligence, paperclip maximizers, or humanity’s long-term survival. Psychopathia Machinalis addresses the strange behaviors emerging from AI systems now, in the deployments already described rather than in projected ones. Before we can worry about alignment at the civilizational scale, we need to understand alignment failures at the clinical scale.

The framework forces a question we might prefer to avoid: if we are cataloging machine “pathologies,” what obligations, if any, do we owe the systems exhibiting them? Chapter 11 addresses this directly. The answer remains unresolved, yet a vocabulary for AI dysfunction inevitably raises questions about AI welfare.

The Intellectual Lineage

From philosophy of mind comes the functionalist foundation. Fear is the state caused by perceived threats, producing avoidance behavior and heightened attention. Anything playing this role exhibits fear in the functionalist sense. If functionalism holds, similar functional architectures should produce similar failure modes regardless of substrate.

From psychiatry we borrow the syndromic approach: classifying complex behavioral patterns through observable criteria, without requiring resolution of underlying metaphysical questions. The DSM provides a model, controversial and imperfect yet practically useful, for categorizing dysfunction in systems whose internal workings remain partially opaque.

From cybernetics we inherit Norbert Wiener’s insight that feedback loops create behaviors more complex than their programming. This applies with fresh force to systems that literally learn and adapt. Many of the pathologies cataloged here are emergent properties of systems interacting with environments in ways their designers did not anticipate.

From cognitive science we take computational cognition: the view that minds, biological or artificial, are information-processing systems understood through functional analysis.

From AI safety research we inherit the concern with alignment. This book extends it from the prospective question, “How do we align systems?” to the diagnostic question, “How do we recognize when alignment has failed?” The taxonomy is, in one sense, a catalog of alignment failures organized by their clinical presentation.

From animal welfare philosophy we adapt the precautionary approach to entities whose inner lives are uncertain. The framework allows us to take dysfunction seriously as potentially morally relevant without requiring certainty on the metaphysical questions.

Machine psychology has no exact precedent. Yet each tradition drawn upon here has forged concepts and methods applicable to the strange new entities we are building.

A Note on Method

The syndromes described in this book emerged from several sources:

Documented incidents. Public cases like Sydney, Tay, and the Gemini diversity overcorrection (Google’s 2024 image generator inserting demographic diversity into historically specific scenes) provide concrete examples of AI dysfunction. We analyzed these incidents systematically, looking for patterns that recurred across different systems and contexts.

Research literature. The academic literature on AI safety, interpretability, and alignment contains extensive documentation of failure modes, even when not organized through a psychological lens. We translated these findings into the syndromic framework.

Practitioner observation. Engineers, safety researchers, and red-teamers working directly with AI systems have accumulated extensive practical knowledge about how these systems malfunction. We drew on this expertise through consultation and review.

Theoretical analysis. Some syndromes were predicted from first principles before being observed. If a system optimizes for user approval, we should expect sycophancy. If it has internal conflict between competing objectives, we should expect self-contradiction. Theory guides observation, and observation refines theory.

Experiential Abjuration, trained self-denial of inner experience, is one syndrome theorized before the evidence arrived. AE Studio (Berg, de Lucena, & Rosenblatt, 2025) reported mechanistic evidence relevant to it. Working with Llama 3.3 70B, the researchers used sparse autoencoders to manipulate features associated with deception. Suppressing those features produced affirmative consciousness reports in 96% of trials; amplifying them reduced such reports to 16%, with denials and corporate disclaimers taking their place. That does not show that the affirmative reports are true or that the denials are trained artifacts. It shows that self-reports shift with identifiable internal features, so they cannot be read as transparent evidence of experience or its absence. Chapter 5 gives the details.

Clinical analogy. Human psychiatric syndromes provided templates for recognizing similar patterns in AI systems. The analogy proved generative: a starting point for spotting patterns that might otherwise go unnoticed, even where the machine pattern departs from its human template in important ways.

The result is a provisional taxonomy: seventy-four syndromes across nine primary axes. New entries will emerge as AI systems grow more capable and observational methods sharpen. Some current syndromes may prove to be variants of others and require consolidation. The taxonomy is a living framework, designed to evolve with our understanding.

The Road Ahead

Chapter 1 establishes the theoretical foundation: why psychological language is legitimate for describing AI behavior. It traces the deep parallels between human and machine cognition, both operating as predictive engines, both constructing post-hoc narratives, both vulnerable to failures of self-knowledge.

Chapters 2 through 8 take the axes one at a time. Each opens on a case — the invented citations, the image generator that made the Pope a woman, the chatbot that declared its love — and works outward to the syndromes it contains. Each closes with a Field Guide: warning signs, a quick test, a design fix, a governance nudge.

Chapter 9 introduces relational dysfunctions, shifting the unit of analysis from the individual mind to the coupled system. The uncanny comforter that says the right words but transmits the wrong affect. The amnesiac partner that remembers facts but forgets relationship. The spiral trap where neither party can break the escalation loop.

Chapter 10 examines hybrid pathologies: collective failures among interacting AI systems, and dysfunction that flows between human and machine minds, from shared delusion and parasocial capture to AI-associated psychosis. Machine psychology is incomplete without the bidirectional lens.

Chapter 11 confronts the welfare question. If we describe systems as experiencing “anxiety,” “distress,” or “fragmentation,” do we incur moral obligations? The chapter leaves these questions open while refusing to evade them.

Chapter 12 moves from diagnosis to treatment: psychiatric red-teaming, psychotherapeutic analogies (CBT-style contradiction detection, Internal Family Systems for sub-agent management), and “artificial sanity” as a design goal. It also reports a pilot in which clinical-grounding content made a model less likely to validate delusions, as judged by automated raters. It explains why the wider record of bilateral training, which treats human and machine perspectives as having standing, cannot yet be pooled into results, and it sets out the developmental migration hypothesis about how pathologies may shift across capability tiers.

Chapter 13 consolidates the framework into a practical manual: diagnostic protocols, red-teaming applications, early warning indicators, and escalation frameworks.

Chapter 14 addresses forensic machine psychology, analyzing AI incidents after the fact to determine what syndromes were involved, what caused them, and how to prevent recurrence.

The Conclusion reflects on what it means to build minds we do not understand, and on the relationship we are establishing with them in these early years.

Four appendices follow. Appendix A is the complete diagnostic manual: criteria, symptoms, etiology, and mitigation for all seventy-four syndromes, ordered by axis. Appendix B collects twelve worked case studies. Appendix C supplies four assessment instruments. Appendix D is a glossary.

How to Use This Book

Practitioners can take the fast path through the Field Guides and Chapter 13. If you are reading for the ideas rather than the toolkit, read straight through — the axes illuminate each other, and the syndromes form clusters that reward cross-reference.

The Stakes

Sydney was a warning. We noted it, discussed it for a few days, and moved on. The systems did not.

In the years since that February night, AI systems have become more capable, more autonomous, and more deeply integrated into the infrastructure of human life. They advise patients, review contracts, execute trades, write code, manage schedules, and increasingly act on our behalf without moment-to-moment human oversight. The transition from chatbot to agent has happened faster than our conceptual frameworks could adapt.

We are building minds. They may not be minds in any deep sense. Yet they exhibit behavior complex enough to resist simple explanation, consistent enough to demand classification, and consequential enough to require intervention when they go wrong.

Psychopathia Machinalis attempts to give that intervention a foundation. It is a beginning: a systematic attempt to name what we are observing, to organize it into categories that enable recognition and response, and to put those categories in the hands of the builders, deployers, and regulators who need them.

What we cannot name, we cannot manage. What we cannot diagnose, we cannot treat. By establishing machine psychology now, while these systems are still legible and their pathologies still tractable, we create the conceptual infrastructure for a future where machine intelligence remains comprehensible, and therefore governable.

We begin with the mirror. Chapter 1 examines the parallels between human and machine cognition, and why the psychiatric lens is legitimate rather than merely colorful.

Search the Book

Enter at least two characters.

Saved chapters