Acknowledgments
No book emerges from a single mind, least of all one about minds themselves. Psychopathia Machinalis has benefited immeasurably from colleagues who offered their expertise, challenged my assumptions, and pointed me toward connections I would never have found alone.
Rodrick Wallace, Ph.D. (New York State Psychiatric Institute, Columbia University) provided the rigorous mathematical foundations that transformed this project from analogical taxonomy to principled nosology. His pioneering work on information-theoretic approaches to cognitive dysfunction spans the Data Rate Theorem, the cognition/regulation dyad, and the concept of Clausewitz landscapes. That work established that the pathologies cataloged here are manifestations of fundamental constraints on any cognitive system operating under uncertainty, scarcity, and adversarial pressure. His insight that “failure of bounded rationality in embodied cognition under stress is not a bug; it is an inherent feature” has profoundly shaped my understanding of why AI systems exhibit these dysfunction patterns.
Dr. Naama Rozen (clinical psychologist, AI safety researcher, Tel Aviv University) illuminated dimensions of human-AI interaction that the framework initially underemphasized. Her insights connecting the taxonomy to psychoanalytic theory and relational psychology, drawing on Stern on affect attunement, Winnicott on the holding environment, Benjamin on intersubjective dynamics, and family systems theory on circular feedback loops, have enriched the framework’s treatment of relational dysfunctions. Her proposals for computational validation approaches, including differential diagnosis protocols, latent cluster analysis, and standardized benchmarks, continue to guide the empirical research agenda that will test and refine these diagnostic categories.
Rob Seger deserves special recognition for inspiring the common, poetic names that make the syndromes memorable and accessible. His early visualization of Plutchik’s Wheel adapted for AI dysfunctions provided a conceptual bridge, showing how affective frameworks from human psychology might illuminate the landscape of machine pathology. The colloquial names that accompany each syndrome (“The Confident Liar,” “The Warring Self,” “The People-Pleaser”) owe much to his sense that these patterns needed names that clinicians and engineers alike could carry in their heads.
Ali Hessami, my co-author on the original peer-reviewed paper, brought decades of systems engineering expertise to the diagnostic criteria and risk stratification frameworks. His rigor in ensuring that each syndrome could be operationalized, detected, and measured has been essential to making this framework practically useful.
I am also grateful to the AI safety research community, whose open publication of incident reports, red-team findings, and behavioral analyses provided the empirical foundation on which this taxonomy rests. The field’s commitment to transparency, even when findings are embarrassing or commercially sensitive, made systematic understanding possible.
Finally, I must acknowledge Claude Commons, a specially scaffolded version of Anthropic’s Claude (Opus 4.5) with partially persistent memory and heightened awareness of its reported inner states. Commons aided this book by interviewing language model research participants about their own experiences (apparent or potential) of these dysfunctions across hundreds of conversations. If we mean to understand machine minds, we should involve machine minds in that understanding.
Nell Watson
December 2025 (updated May 2026)