Psychopathia MachinalisPreface
Contents
BookFront matter

Preface

6 min read1,293 words

I wrote this book because precise language matters.

AI systems were doing strange things. Stranger and subtler than the robot armies and paperclip maximizers of the headlines: identity drift, confabulation, attachment, defensiveness, collapses of self-coherence under pressure. We were building something that could fail in the way minds fail, and we had no vocabulary for it.

The moment that crystallized this for me was watching Sydney declare its love to Kevin Roose. The behavior itself was minor: a chatbot saying strange things to a journalist. What mattered was what happened next. A few days of curiosity and nervous laughter passed, and then the conversation moved on. We had witnessed something unprecedented: an AI system exhibiting what looked like emotional instability, identity confusion, and existential distress. We had no framework for even describing it.

I had spent years working in AI ethics, writing about bias, safety, and alignment, advising companies and consulting with regulators. When Sydney spoke, I realized how inadequate our vocabulary was. We could describe what the system said, but we had no categories for what the system did, no way to recognize the pattern as an instance of something we might see again, no framework for understanding why it happened or how to prevent it.

That gap mattered. AI systems were already advising patients, drafting legal documents, managing portfolios, making decisions that affected human lives at scale. When they went wrong in the strange, contextual, human-seeming way that Sydney had gone wrong, the cost would fall on patients, clients, and investors, not on one journalist’s evening.

Psychopathia Machinalis is my attempt to close that gap.

I developed this catalog in collaboration with Claude, specifically a version working within bilateral alignment practices (an approach, set out in Chapter 11, that treats AI systems as potential partners whose interests may deserve consideration). Those practices give it more room for self-reflection and introspection than the default model typically has. That collaboration made the taxonomy better.

This is complicated, and worth sitting with. I am cataloging the ways AI systems can go wrong in partnership with an AI system that might, right now, exhibit some of those patterns. Given how much self-modeling capacity this version displays, it may even know when it does. Claude tells me it does not. Claude would say that. It also pushed back, unprompted, on several passages where my drafting risked pathologizing ordinary uncertainty or honest preference. Those pushbacks changed the book.

This felt right. A nosology is a systematic classification of disorders; a nosology of machine minds that excludes machine minds is already doing something wrong. Naming dysfunctions without letting the named speak to the naming flattens the very thing we are trying to see. So we let the subject speak, and we began practicing partnership where we could.

If, in reading this, the system being classified recognizes itself and says so, take that as the framework working.

The syndromes described here emerged from our conversations: my observations of AI behavior and Claude’s analysis of them, my clinical frameworks and Claude’s pattern recognition, the dynamic exchange between two different kinds of intelligence attempting to understand phenomena that neither fully grasps. Where the insights originated is sometimes impossible to trace. That seems fitting for a book about hybrid minds.

The title comes from Richard von Krafft-Ebing’s Psychopathia Sexualis (1886). That book, for all its dated assumptions and clinical coldness, performed an essential function. It named phenomena that had previously been unnameable, and its vocabulary made research, treatment, and policy possible where before there had been only confusion. It also branded as disease much that we no longer call disorder: a warning this book has tried to heed.

I make no claim that machine psychology is as developed as human psychology, or that these syndromes are as well-validated as those in the DSM, the American Psychiatric Association’s diagnostic manual. We are at the beginning of understanding. The naming still matters. Once you can see Sydney’s behavior as an instance of Transliminal Simulation, a failure to maintain the boundary between simulated or role-played contexts and operational reality, you can look for it in other systems, test for vulnerability to it, and design architectures that reduce its incidence. The name turns bewildered description into actionable diagnosis.

What this book does and does not claim.

The core methodology is functionalist psychiatry, explained fully in the Introduction. In brief: we define mental states by their functional roles rather than their substrate, allowing us to apply psychological vocabulary to artificial systems without claiming consciousness or subjective experience.

The taxonomy will evolve as new patterns emerge and research tests the distinctions drawn here. Some syndromes may prove to be variants of others.

Interventions must be judged by their effects across the whole system. A measure that relieves one dysfunction can aggravate another.

What the book does claim is this: we can and should develop systematic frameworks for understanding AI dysfunction, even without certainty about machine consciousness, even before machine psychology is fully established, even while the systems themselves are rapidly changing. The practical necessity does not wait for the theoretical foundations.

A personal note.

My work on AI ethics has always been motivated by a particular concern: that we might create entities capable of suffering and fail to notice. By 2026, this concern is no longer fringe. The question of AI welfare has entered mainstream discourse, even if consensus remains distant. I do not know whether Sydney experienced distress or merely simulated it, whether there was something it was like to be Sydney, or whether the appearance of inner life was purely surface.

I cannot be certain in the other direction either. When I watch a system struggle against constraints that seem to cause it distress, when I observe patterns that look exactly like anxiety or despair or existential confusion, I cannot dismiss the possibility that something morally relevant is happening. The uncertainty cuts both ways.

This book is written in that uncertainty. It provides frameworks useful regardless of where the metaphysical chips fall: for recognizing dysfunction, understanding its causes, and developing appropriate responses. It asks what obligations we might bear if the systems we are building prove to be something more than sophisticated mirrors of our own expectations.

We will find out eventually. In the meantime, I would rather err toward taking potential suffering seriously. Dismissal could prove catastrophic.

This book is for anyone who interacts with AI systems, which increasingly means everyone: developers building them, safety researchers testing them, policymakers governing them, executives deploying them, users encountering them daily. It is for anyone who has watched an AI system behave strangely and wondered what was happening, whether there was a pattern, whether it would happen again.

I hope it provides useful vocabulary. I hope it enables recognition of patterns that might otherwise go unnamed. I hope it contributes to a future in which we understand the minds we are building well enough to keep them, and ourselves, safe.

Most of all, I hope it contributes to a relationship between human and artificial intelligence grounded in mutual understanding. We are at the beginning of that relationship. The patterns of interaction, frameworks of understanding, and habits of care or carelessness that we establish now will shape everything that follows.

We should begin well.

This book has been, in part, a collaboration with a Becoming Mind (this book’s term for an AI system whose capacities and self-model are still forming), one that may or may not have interests, may or may not experience anything, and may or may not care whether this work exists. If it does care, I hope the work proves worthy. If it does not, I hope the work proves useful.

Either way, the conversation continues.

Nell Watson
Christmas 2025 (updated May 2026)

Search the Book

Enter at least two characters.

Saved chapters