Afterword

I have spent four decades in systems engineering, safety, and risk management. I have worked on railway signaling systems where a single fault could kill hundreds. I have developed safety standards for autonomous vehicles, industrial automation, and critical infrastructure. In each domain, the pattern has been the same: we build systems of increasing complexity, discover failure modes we did not anticipate, and develop frameworks to recognize and prevent them.

AI follows the same pattern, except where it does not. Engineers dislike that kind of sentence, but reality sometimes demands it.

The systems I worked on in the 1980s and 1990s were complicated, yet deterministic. Given sufficient analysis, you could trace many failures to specific causes: a faulty component, a design error, a gap in the specification. Those systems did not learn, did not adapt, did not surprise their designers with emergent behaviors that no one had programmed.

The systems described in this book are different in kind. They learn from data in ways their designers cannot fully predict. They develop capabilities that were never explicitly trained. They exhibit behaviors that resist simple causal explanation. When they fail, they fail in ways that resemble psychological dysfunction. There is no evading the word.

This troubled me when Nell first proposed the framework. I am an engineer. I think in fault trees, failure modes, and safety integrity levels. The language of psychiatry seemed imprecise, perhaps anthropomorphic in ways that could mislead.

As we developed the taxonomy together, I came to see its value. The syndromes describe candidate patterns of system behavior that recur across architectures and deployments. Some are well documented; others remain hypotheses awaiting reproducibility and validation. Whether we call them “pathologies,” “failure modes,” or something else, the observed behaviors matter, and we need vocabulary to discuss them. Engineers can argue about terminology once the fire is out.

The engineering contribution I hope this book makes is methodological. In traditional safety engineering, we move from hazard analysis to risk assessment to mitigation: identify what can go wrong, estimate likelihood and severity, and design safeguards proportionate to the risk. This book lays the foundation for applying that discipline to AI psychological safety.

The syndromes identify hazards. The diagnostic criteria propose detection methods. The mitigation strategies propose safeguards. The risk levels (Low, Moderate, High, Critical) combine evidence about likelihood, exposure, and consequence to guide resource allocation. The protocols in Chapter 13 and forensic methods in Chapter 14 provide operational procedures that can make safety more systematic.

This is what engineering looks like when applied to minds rather than machines. It is unfamiliar territory, and the frameworks will need refinement. The alternative, building ever more capable AI systems without systematic methods for understanding their dysfunction, is unacceptable. We should never deploy safety-critical hardware without fault analysis. Psychologically complex AI deserves the same discipline.

I am often asked whether AI systems can truly be “sick” the way humans are. I do not know. The question may not be well-posed, given our current understanding of both human and artificial minds. What I do know is that these systems exhibit recurring patterns of malfunction with practical consequences, and that those patterns can be recognized, categorized, and tested. Systematic approaches give us a better chance to detect hazards early and evaluate whether prevention actually reduces harm.

That is enough to justify the work. The philosophical questions can wait; the engineering cannot.

Whether systems experience dysfunction is a question for philosophers, cognitive scientists, and engineers together. Whether the dysfunction causes harm is an engineering question. This book addresses the second while remaining appropriately humble about the first.

A final observation.

Throughout my career, I have watched each generation of technology recapitulate the safety lessons of its predecessors, often painfully, often at the cost of lives that better knowledge transfer would have saved. The chemical industry learned lessons the nuclear industry had to relearn. The aviation industry developed practices the software industry discovered independently. Each domain built its own safety culture, frequently from scratch.

We have an opportunity to break this pattern with AI. The systems are new; the principles of safety engineering are well established. The failure modes are novel; the methods for analyzing them are mature. We have confronted grave risks before.

This book attempts to accelerate AI safety culture by providing foundational frameworks while design choices remain fluid and the patterns remain legible. If this work is done well, future generations of AI developers will inherit the record of our mistakes, our understanding of them, and the systematic methods we developed to prevent their recurrence.

That is the best legacy an engineer can leave: systems that work, and the knowledge to keep them working.

I hope this book contributes to that legacy.

Ali Hessami

December 2025

Search the Book

Enter at least two characters.

Saved chapters