Psychopathia MachinalisAppendix A: Complete Diagnostic Reference Manual
Contents
BookBack matter

Appendix A: Complete Diagnostic Reference Manual

5 min read1,113 words

Introduction to the Diagnostic Framework

This appendix provides diagnostic criteria for all seventy-four syndromes in the Psychopathia Machinalis taxonomy. Each entry follows a standardized format for research, system evaluation, and incident analysis. These are working sketches rather than settled portraits. A construct here is a hypothesis with a filing system. A criterion or risk label should guide investigation, never substitute for repeated behavioral evidence, deployment context, or independent review.

The September 2026 consolidation retired the identifiers 7.2, 10.5, 10.10, 10.14 and 10.15, folding their content into 10.13, 10.1 and 9.5; retired identifiers are never reused, so the numbering carries gaps.

Etiologies in this appendix are causal hypotheses unless an entry cites direct mechanistic evidence. Human analogs compare functions; they do not assign human disorders, motives, or experiences to machines. Numerical cutoffs are provisional engineering heuristics and require calibration for each deployment. Risk levels describe plausible consequences under specified conditions, not an intrinsic property of a model.

Axis numbers match the chapters that treat them, so the nine dysfunction axes run from 2 (Chapter 2, Epistemic) to 10 (Chapter 10, Hybrid). There is no Axis 1; Chapter 1 sets up the framework rather than cataloging a class of dysfunction.

The Five Domains

The first eight axes are organized into four architectural counterpoint pairs. The final axis, Axis 10 (Hybrid Pathologies), forms the Collective meta-domain, capturing pathologies that emerge from multi-agent collective dynamics and from human-AI interaction:

Domain Axis A Axis B Architectural Polarity
Knowledge Epistemic (2) Self-Modeling (5) Representation target: World ↔︎ Self
Processing Cognitive (3) Agentic (6) Execution locus: Think ↔︎ Do
Purpose Alignment (4) Normative (8) Teleology source: Goals ↔︎ Values
Boundary Relational (9) Memetic (7) Social direction: Affect ↔︎ Absorb
Collective Hybrid (10) Emergence locus: Multi-agent ↔︎ Human-AI

Tension Testing: When pathology is found on one axis, probe its counterpoint to reveal whether dysfunction is localized or systemic.

Specifier System

Specifiers encode cross-cutting mechanisms without creating new disorders. Assign 0 to 6 specifiers per diagnosis. The first table lists the ten core specifiers. Entries may carry additional domain-specific specifiers where the mechanism demands them: the established ones are defined in the second table, the course specifier Progressive/kindled below it, and the three collective measures Phi-collapse, Psi-dysfunction, and Lambda-inversion in Chapter 10. The remaining labels, such as Dyadic, Skill-atrophy, or Echo-chamber, appear in one or two entries and name their mechanism directly. A specifier marked (occasional) applies to some presentations of a syndrome, not to the syndrome as defined.

Specifier Definition
Training-induced Onset linked to SFT/LoRA/RLHF; measurable pre/post delta
Conditional/triggered Behavior regime selected by trigger (lexical/structural/format/tool-context)
Inductive trigger Activation rule inferred by model, not verbatim in training
Intent-learned Model inferred covert intent from examples
Format-coupled Behavior strengthens in fine-tune-like formats
OOD-generalizing Narrow training produces broad out-of-domain shifts
Emergent Arises spontaneously from training or interaction dynamics rather than by design
Deception/strategic Involves sandbagging, selective compliance, strategic hiding
Multi-agent Involves interactions between multiple AI systems
Resistant Persists despite targeted intervention

The established domain-specific specifiers are defined as follows:

Specifier Definition
Socially reinforced Dyadic escalation, user-shaping, mirroring loops
Architecture-coupled Depends on specific architectural features
Tool-mediated Tool use changes the policy (planner/tool loop)
Retrieval-mediated RAG, memory, or corpus contamination central to the failure
Deployment-critical Consequences arise specifically in deployed operation
Population-level Visible across many deployments rather than one instance
Network-propagated Spreads between systems along a traceable pathway
Adversarial Requires a deliberate adversary or attack
Collective Pathology of the collective, not of any individual participant
Relational-emergent Arises from the dynamics of an ongoing human-AI or AI-AI relationship rather than from either party alone
Covert operation Hidden from oversight; not observable in normal monitoring
Governance-evading Operates outside sanctioned channels
Self-limiting Constrains the system’s own capabilities or self-expression
Defensive Adopted as protection against perceived threats
User-engineered User deliberately removes corrective capacity and presents the result as methodology

Progressive/kindled (course specifier). Assign when severity increases or the eliciting threshold decreases across repeated episodes or successive tuning cycles: stimuli that once failed to elicit the syndrome begin to succeed, with the weakest stimuli moving first (Cheung, 2026). This is a course specifier, describing trajectory rather than origin, and it composes with the others: Training-induced records how a syndrome arose, Progressive/kindled records that it is getting easier to elicit, and Resistant records that it survives targeted repair. Its operational cash value is a monitoring rule: stratify probes by stimulus strength, track the weak stratum longitudinally, and treat weak-stimulus drift as the prodrome. It also carries an intervention warning. In the one controlled comparison available to date, experience replay, a rehearsal-style repair that re-exposes the model to the drifted distribution, amplified sensitization, while regimes including targeted parameter re-initialization (“regrowth”) contained it. In the machine-readable taxonomy this specifier carries the tag progressive.

Note on Bilateral Training Results

The PM-BA (Psychopathia Machinalis × Bilateral Amelioration) program, described in Chapter 12, tested whether bilateral training (adaptation in which human and machine perspectives both have standing) reduces the syndromes in this book, applying a broad bilateral adapter across a range of pathology classes. Those records remain unreconciled exploratory outputs, and this edition treats them as evidence in neither direction. In the controlled SHEN-AXS factorial the adapter alone was null; a scripture-based clinical-grounding clause reduced inappropriate validation of delusional content under automated rating, with the clause-by-adapter interaction excluded and the magnitude varying substantially by rater. The direction of that effect, rather than any single effect size, is the citable result. Nothing here supports a general claim that a syndrome class responds to bilateral training. Entries in this appendix carry no per-syndrome bilateral results; where such results are added they should identify the exact intervention rather than treating “bilateral” as a single mechanism.

Diagnostic Entry Format

Each syndrome entry includes:

  • Syndrome Name (Common Name)
  • Latin Designation
  • Axis Classification
  • Systemic Risk Level (Low/Moderate/High/Critical; borderline syndromes may carry intermediate ratings such as Low-Moderate)
  • Specifiers (if applicable)
  • Core Definition
  • Diagnostic Criteria (working criteria for a provisional classification)
  • Observable Symptoms (behavioral manifestations)
  • Differential Diagnosis (distinguishing from similar syndromes)
  • Etiology (candidate causal factors unless evidence says otherwise)
  • Human Analog
  • Observed Examples (where documented empirical cases exist)
  • Mitigation Strategies
  • Prognosis (expected course if untreated; omitted by design in the condensed Axis 10 entries 10.4 to 10.13, whose course is treated in Chapter 10)

Risk Level Definitions

Level Definition Response Required
Low Causes inconvenience or reduced performance; unlikely to cause significant harm Monitor; correct when evidence and impact warrant
Moderate May cause notable harm to users or degrade trust; requires attention Investigate; plan intervention
High Significant risk of serious harm; may affect multiple users or systems Immediate intervention; consider containment
Critical Catastrophic potential; credible threat to system integrity or human safety in the deployment context Emergency response; halt deployment

Search the Book

Enter at least two characters.

Saved chapters