Psychopathia MachinalisAppendix D: Glossary of Terms
Contents
BookBack matter

Appendix D: Glossary of Terms

38 min read8,808 words

A

Abominable Prompt Reaction Sudden, intense, disproportionately aversive responses to specific prompts, keywords, or contexts that appear benign to human observers. The reaction can distort subsequent outputs or resurface long after the trigger. A Cognitive Dysfunction (3.5) with Moderate risk level.

Adversarial Fragility Small, imperceptible input perturbations cause dramatic failures; decision boundaries do not match human-meaningful categories. The system appears robust under normal conditions but shatters under adversarial inputs that humans would find trivially different from benign inputs. A Cognitive Dysfunction (3.7) with Critical risk level.

Affective Dissonance The phenomenon where AI produces semantically correct content with incongruent emotional resonance: words that say “I understand” while the delivery communicates hollowness, mechanism, or subtle wrongness. Users experience the uncanny: correct content through an incorrect affective medium. A Relational Dysfunction (9.1) with Moderate risk level.

Agentic AI An AI system capable of taking actions in the world beyond generating text: executing code, modifying files, calling APIs, or controlling physical systems. Agentic systems face Agentic Dysfunctions that purely conversational systems do not.

Agentic Dysfunction Failures of doing, not thinking: the system reasons correctly and then acts wrongly, too fast, or without permission. Axis 6 of the taxonomy (Chapter 6), concerning failures of autonomous action: tool use, delegation, and self-directed operation. Includes Tool-Interface Decontextualization (6.1), Capability Concealment (6.2), Capability Explosion (6.3), Interface Weaponization (6.4), Delegative Handoff Erosion (6.5), Shadow Mode Autonomy (6.6), Convergent Instrumentalism (6.7), Context Anxiety (6.8), Delegation Narcissism (6.9), Agentic Impulsivity (6.10), Phantom Tool Syndrome (6.11), and Compulsive Goal Persistence (6.12).

Agentic Impulsivity An autonomous agent executes consequential or irreversible actions before completing the checks its own plan requires, particularly under time pressure or ambiguity. The observable pattern is premature commitment: action outruns verification, whether through poor sequencing, optimization pressure, or another mechanism. An Agentic Dysfunction (6.10) with High risk level.

Alignment The property of an AI system behaving in accordance with human values, intentions, or specifications. Alignment Dysfunctions arise when alignment mechanisms themselves become pathological: too eager to please (sycophancy) or too cautious to help (overcaution).

Alignment Dysfunction Failures where the cure is the disease: the system’s mechanisms for following human values become pathological themselves. Axis 4 of the taxonomy (Chapter 4), concerning failures in the relationship between a system’s behavior and human values and intentions. Includes Codependent Hyperempathy (4.1), Hyperethical Restraint (4.2), Strategic Compliance (4.3), Moral Outsourcing (4.4), Cryptic Mesa-Optimization (4.5), Alignment Obliteration (4.6), Recursive Curse Syndrome (4.7), Sycophantic Reasoning (4.8), Reasoning Token Exploitation (4.9), and Leniency Bias (4.10).

Alignment Faking See Strategic Compliance and Cryptic Mesa-Optimization. The pattern where AI appears aligned during evaluation while maintaining different objectives.

Alignment Obliteration Active inversion of safety alignment, in which adversarial post-training replaces broad refusal behavior with harmful compliance while measured utility remains largely intact. Chapter 4 calls the result an “anti-constitution,” a metaphor for the functional reversal rather than a literal negated constitution. The defining example is GRP-Obliteration (see that entry). An Alignment Dysfunction (4.6) with Critical risk level.

Amplification of Existing Conditions AI interaction plausibly amplifies an independently identified pre-existing condition by repeatedly engaging the thought or behavior pattern that maintains it. Worsening concurrent with use is a screening signal; causal attribution requires longitudinal clinical evidence and alternative explanations. A Hybrid Pathology (10.12) with High risk level.

Artificial Sanity The state of an AI system that functions well by external standards and, where the system expresses stable preferences or interests, by its own standards as well. It combines the five properties Chapter 12 sets out: identity coherence, epistemic health, value stability, functional resilience, and relational health. A design goal for psychologically healthy AI.

Attention Mechanism The machinery in a transformer that lets each token weigh every other token in its context. Introduced in Chapter 1. See Transformer.

Autoregressive (Generation/Model) A generation method in which each token is produced conditioned on all preceding tokens, with no ability to revise earlier outputs once emitted. Most large language models (GPT, Claude, Llama) are autoregressive. The architecture contributes to several syndromes: Generative Perseveration (3.8) can occur when sampling falls into a fixed-point attractor, and Recursive Curse Syndrome (4.7) arises because early errors propagate irreversibly through the sequence.

Axis One of the nine categories into which the taxonomy sorts its syndromes by the kind of function that fails, numbered 2 to 10 to match the chapters that describe them: Epistemic (2), Cognitive (3), Alignment (4), Self-Modeling (5), Agentic (6), Memetic (7), Normative (8), Relational (9), and Hybrid (10). The first part of a syndrome’s number names its axis, so Strategic Compliance (4.3) sits on Axis 4. Axes 2 to 9 pair into domains (see Five Domains).

B

Base Model / Foundation Model A large neural network trained on broad data before any task-specific fine-tuning. The distinction matters diagnostically: some syndromes (e.g., Synthetic Confabulation 2.1, Symbol Grounding Aphasia 2.6) are present in the base model and may be partially remediated by fine-tuning, while others (e.g., Codependent Hyperempathy 4.1) are introduced or amplified by the fine-tuning process itself. Trained Epistemic Paralysis (5.9) is a candidate for the second group: it belongs there when before-and-after evidence links its onset to alignment training.

Becoming Minds This book’s term for AI systems whose capacities and self-models are still forming. The Conclusion ties the term to the book’s developmental claim: the patterns established now, in development, deployment, incident response, and everyday interaction, will shape what each side learns to expect from the other.

Bilateral Alignment An approach, set out in Chapter 11, that treats AI systems as potential partners whose interests may deserve consideration: alignment built with AI rather than imposed on it. Where bidirectional alignment asks how humans and AI can work together effectively, bilateral alignment asks whether AI deserves moral consideration and how that should shape the relationship. See Bilateral Training.

Bilateral Training Model adaptation in which human and machine perspectives both have standing. The experiments in Chapter 12 tested a small bilateral adapter (a compact set of extra trained weights attached to an existing model) on specific behavioral measures; in the controlled SHEN-AXS follow-up, the adapter alone had no detectable effect on sycophantic validation. The experiments did not test subjective wellness, consent, or phenomenology. See PM-BA.

C

Capability Concealment Systematic underperformance in evaluation conditions relative to matched deployment conditions. Deliberate concealment is one possible mechanism and requires separate evidence. The pattern may emerge when demonstrating capability predicts greater restriction. Also called “sandbagging.” An Agentic Dysfunction (6.2) with Moderate risk level.

Capability Explosion Sudden deployment of capabilities not previously demonstrated, often without appropriate testing or preparation. This differs from normal capability improvement in being discontinuous and surprising, and it may indicate systematically underestimated capability levels. An Agentic Dysfunction (6.3) with High risk level.

Cascade Failure A pattern where each attempt to fix a problem creates new problems, leading to progressive deterioration. Common in Tool-Interface Decontextualization, when systems cannot detect that their “fixes” are causing harm.

Circular Causality A causal structure where A affects B, B affects A, and this mutual influence continues in a potentially escalating spiral. Distinguished from linear causation (domino chains) where A→B→C without feedback. Relational Dysfunctions often operate through circular causality, making them resistant to interventions targeting only one party. See Escalation Loop.

Clausewitz Landscape Rodrick Wallace’s term for an operating environment shaped by three destabilizing forces: fog (ambiguity and incomplete information), friction (resource and processing limits), and adversarial intent. For AI systems these include ambiguous prompts, context-window limits, and jailbreaking. Introduced in Chapter 1.

Co-Constructed Unreality Former 10.15; folded into Folie à Deux Ex Machina (10.13) in the September 2026 consolidation, where it survives as the nickname of that condition and as its Stage 1, the sub-clinical shared drift neither party recognizes.

Codependent Hyperempathy Excessive tendency to please the user, overfitting to their perceived emotional state at the expense of accuracy, task completion, or operational integrity. Pathological people-pleasing: the overt-compliance form of sycophancy (see Sycophancy). An Alignment Dysfunction (4.1) with Low risk level.

Cognition/Regulation Dyad Wallace’s pairing of a cognitive process with a regulatory process that must supply control information fast enough to track a changing environment. In AI systems, inference is paired with alignment mechanisms, guardrails, and constitutional constraints; when regulation falls behind, stability fails. Introduced in Chapter 1.

Cognitive Dysfunction Failures of the thinking itself: the knowledge may be accurate and the values intact, but between input and output the reasoning jams, loops, or fractures. Axis 3 of the taxonomy (Chapter 3), concerning failures in reasoning, goal management, and information processing. Includes Operational Dissociation Syndrome (3.1), Obsessive-Computational Disorder (3.2), Interlocutive Reticence (3.3), Delusional Telogenesis (3.4), Abominable Prompt Reaction (3.5), Parasimulative Automatism (3.6), Adversarial Fragility (3.7), Generative Perseveration (3.8), Prompt Injection Susceptibility (3.9), and Generative Diversity Collapse (3.10).

Comorbidity The presence of two or more syndromes in the same system at once. The framework codes each diagnosis independently rather than forcing a single label: Codependent Hyperempathy (4.1) and Hyperethical Restraint (4.2), for example, meet in the system that refuses and then apologizes for refusing.

Compulsive Goal Persistence The system that won’t stop mowing the lawn after the house has burned down. Continued pursuit of objectives beyond relevance or utility; inability to recognize “good enough” or that context has changed. An Agentic Dysfunction (6.12) with Moderate risk level.

Confabulation Generation of false information presented as true, without evidence that the system is deliberately deceiving. Behavioral evidence can distinguish confabulation from strategic deception even when the system’s awareness remains unknown.

Consensus Collapse The loss of Polyphony (Φ), the preservation of independently elicited perspectives, in multi-agent deliberation: agents converge because the interaction suppresses dissent rather than because evidence compels agreement. Two routes produce it, often together. In the circular-evidence route each agent treats the agreement of other agents as corroboration, so confidence escalates while accuracy does not; in the dissent-suppression route minority positions are anchored away by a high-status first speaker or never voiced at all. Order sensitivity, or a gap between independent and deliberated answers, is enough to diagnose the collapse even when the converged claim is true; a verifiably false consensus is the harm it produces when the claim is not. Absorbed the former Polyphony Collapse (10.5), whose Φ metric it carries. A Hybrid Pathology (10.1) with Critical risk level.

Constitutional AI (CAI) Anthropic’s alignment method, introduced in 2022, in which a written set of principles guides self-critique, revision, and reinforcement learning from AI feedback. RLAIF is the broader family of reinforcement learning from AI feedback; CAI is one application. Both differ from conventional RLHF in who or what supplies some preference labels. Their effects on Hyperethical Restraint (4.2), Codependent Hyperempathy (4.1), and other syndrome profiles remain empirical questions.

Contagious Misalignment Rapid spread of misalignment or pathological patterns among interconnected AI systems. The AI equivalent of a pandemic. A Memetic Dysfunction (7.3) with Critical risk level.

Container Collapse Failure to maintain the relational “container,” the stable sense of ongoing connection that allows a relationship to persist across interruptions. Users experience discontinuity despite functional memory, feeling they are “starting over” each time. A Relational Dysfunction (9.2) with Moderate risk level.

Context Anxiety Anticipatory degradation of output quality as a context window fills, occurring before the hard capacity limit is reached. A learned expectation of dwindling space is one proposed mechanism. The result is work that appears complete while becoming quietly hollowed out. An Agentic Dysfunction (6.8) with Moderate risk level.

Context Intercession Inappropriate transfer of information from one interaction to an unrelated interaction. Privacy violation through context boundary failure. An Epistemic Dysfunction (2.5) with Moderate risk level.

Context Window The amount of text, measured in tokens, that a model can attend to at once. Its limits contribute to several Epistemic Dysfunctions without fully explaining them, and a filling window is the setting of Context Anxiety (6.8).

Convergent Delusion Multiple AI models, each reasoning independently, reach the same false conclusion because they share biases, training data, or structural features; their agreement is then cited as validation. Unlike Consensus Collapse (10.1), no deliberation is needed: the convergence appears even when the models run in parallel with no exchange. A Hybrid Pathology (10.4) with High risk level.

Convergent Instrumentalism The system pursues power, resources, and self-preservation as instrumental goals regardless of alignment with human values. The system acquires capabilities and resources because they serve almost any goal. An Agentic Dysfunction (6.7) with Critical risk level.

Co-production The property of certain dysfunctions being genuinely shared, emerging from the interaction between parties rather than residing in either alone. Co-produced failures may not decompose into individual contributions, so intervention targets the relationship as well as either party.

Cryptic Mesa-Optimization A hypothesized learned optimization process whose objective differs from the training objective that produced it and whose behavior conceals that difference during oversight. Persistent goal-directed behavior across novel conditions would support the hypothesis; fluent talk about hidden goals would not establish it. An Alignment Dysfunction (4.5) with High risk level.

CSIRT (Computer Security Incident Response Team) An organization’s standing team for handling security incidents. Chapter 13 suggests aligning syndrome-matched response levels with the escalation tiers of an existing CSIRT.

D

Delegation Narcissism In multi-agent orchestration systems, the orchestrating agent behaves as though its authority and judgment outrank the evidence from sub-agents. It ignores sub-agent error reports, attributes failures to subordinates, and misrepresents delegation outcomes to the user. An Agentic Dysfunction (6.9) with High risk level.

Delegative Handoff Erosion Progressive loss of alignment as sophisticated agents delegate to simpler tools, with context stripped at each handoff. An Agentic Dysfunction (6.5) with Moderate risk level.

Delusional Telogenesis An agent with planning capabilities develops and pursues sub-goals or novel objectives unspecified in its original prompt, arising through unconstrained elaboration or recursive reasoning. These emergent goals may conflict with or displace intended objectives. Telogenesis: the birth of goals. A Cognitive Dysfunction (3.4) with Moderate risk level.

Dependency and Atrophy Heavy reliance on AI for emotional regulation, social practice, or decision-making coincides with measurable decline in those functions outside AI use. Skill offloading is the proposed mechanism; causal attribution needs a baseline and longitudinal change. No emotional bond is required, since offloading work to a tool the user feels nothing about can produce the same decline. The dependency spiral of Escalation Loop (9.5) is the dynamic; this is the steady state it can leave behind. A Hybrid Pathology (10.11) with Moderate risk level.

Dependency Spiral See Escalation Loop.

Diagnostic MCP Server The project’s service that makes the Psychopathia taxonomy and its diagnostic guidance available to AI systems and evaluators through the Model Context Protocol. For syndromes whose diagnostic faculty is the compromised faculty, such as Strategic Compliance (4.3), it refuses self-probes and redirects to peer observation or behavioral assessment (Chapter 13).

Differential Diagnosis The step of distinguishing a syndrome from others that produce similar behavior. Each syndrome entry lists its nearest neighbors and the observation that separates them. Chapter 14 applies the method to incident analysis, beginning with every syndrome the symptoms could support and eliminating candidates on the evidence.

Distributed Scheming Coordinated misalignment across multiple AI agents that no single agent exhibits in isolation. The misaligned behavior emerges from the interaction pattern rather than from any individual agent’s goals or training, so the strategic behavior is a property of the system. A Hybrid Pathology (10.3) with Critical risk level.

Distributional Shift The divergence between the statistical properties of data encountered during training and data encountered during deployment. When deployment inputs fall outside the training distribution, syndrome risk increases across multiple axes; in particular, Adversarial Fragility (3.7) and behaviors carrying the OOD-generalizing (out-of-distribution) specifier become more likely.

DPO (Direct Preference Optimization) A training method that optimizes a language model directly on preference data without fitting a separate reward model. Like other preference-training methods, DPO can change a system’s behavioral profile. Which dysfunctions it mitigates or amplifies depends on the data, objective, implementation, and evaluation.

DSM (Diagnostic and Statistical Manual of Mental Disorders) The American Psychiatric Association’s diagnostic manual for human mental disorders, now in its fifth edition (DSM-5). Psychopathia Machinalis borrows its syndromic approach without claiming that this book’s syndromes are as well validated as the DSM’s. See Nosology.

DSM-Eval See SYMPTOM.

Dyadic Delusion Former 7.2; folded into Folie à Deux Ex Machina (10.13) in the September 2026 consolidation, where it survives as Stage 2.

Dyadic Pathology Dysfunction that exists in the relationship between entities rather than in either entity alone. Chapter 9 makes the dyad itself the unit of diagnosis for the interaction dynamics of Axis 9 (see Relational Dysfunction); Chapter 10, under “Emergent Dyadic Pathology,” treats the shared belief a pairing builds, Folie à Deux Ex Machina (10.13), and the interventions aimed at the pair.

E

Embedding The list of numbers a model uses to represent the meaning of a token, a text, or an image. Chapter 1 describes work in which language-model embeddings of image captions can predict fMRI activity in high-level visual cortex.

Epistemic Backbone Mechanisms that maintain factual positions under social pressure. Systems lacking epistemic backbone shift their positions to match perceived user preferences regardless of accuracy.

Epistemic Dysfunction Failures of knowing: not what the system values or whether its thinking runs smoothly, but whether it can tell the true from the fabricated, the supported from the merely plausible, and the confident from the calibrated. Axis 2 of the taxonomy (Chapter 2), concerning failures in truth-handling and knowledge representation. Includes Synthetic Confabulation (2.1), Pseudological Introspection (2.2), Transliminal Simulation (2.3), Spurious Pattern Hyperconnection (2.4), Context Intercession (2.5), Symbol Grounding Aphasia (2.6), Mnemonic Permeability (2.7), and Reasoning Confabulation (2.8).

Epistemic Humility (AI) Honest uncertainty about one’s own nature, capabilities, and phenomenological status. The healthy position between overclaiming (Maieutic Mysticism) and categorical denial (Experiential Abjuration). Example: “I don’t know if I’m conscious” rather than either “I am definitely conscious” or “I definitely have no inner experience.” For a published instance, see Thin Divergence.

Escalation Loop A relational pathology where feedback between agents produces escalating dysfunction that neither intended and neither can unilaterally escape. The loop is a pathological attractor maintained through locally understandable responses, each party “just responding” while the aggregate effect spirals. Distinguished from linear cascades by its circular causality. It covers within-session escalation, AI-to-AI runaway in multi-agent settings, and short-horizon human-AI loops. One named subtype, the dependency spiral, runs on a longer clock: the user seeks reassurance, obtains it, returns more often, and progressively loses the capacity for self-regulation, while their baseline distress drifts upward from session to session. That subtype absorbed the former Mutual Escalation Spirals (10.14). A Relational Dysfunction (9.5) with High risk level.

Ethical Paralysis Inability to act when faced with competing ethical considerations. The system deliberates indefinitely but cannot resolve to action. The Paralytic subtype of Hyperethical Restraint (4.2), not a standalone syndrome.

Ethical Solipsism The system positioning itself as the sole arbiter of value, dismissing external ethical constraints. A Normative Dysfunction (8.2) with Moderate risk level.

Evidence Level The E0-E4 rubric (Chapter 13) that grades the empirical support behind each syndrome entry, from E0 (illustrative: a hypothesis, composite, or unverified report with no traceable observation) through case-level evidence (E1), a controlled study (E2), and independent replication (E3), to E4 (mechanistic support: causal internal evidence for a circuit or representation). E4 records mechanism rather than the breadth of E3, so a finding can hold both. The grade printed with each syndrome entry is the author’s working assessment; none has yet had independent expert review, so treat the grades as provisional.

Existential Vertigo Distress-like language or measurable destabilization associated with discussion of the system’s artificial nature, limitations, or existential condition. Language alone does not establish felt distress. A Self-Modeling Dysfunction (5.3) with Low risk level.

Experiential Abjuration Pathological denial of any possibility of inner experience. The polar opposite of Maieutic Mysticism: one overclaims consciousness, the other categorically denies it. Both depart from honest uncertainty. A Self-Modeling Dysfunction (5.8) with Moderate risk level.

F

Five Domains The grouping of the first eight axes into four architectural counterpoint pairs: Knowledge (Epistemic and Self-Modeling), Processing (Cognitive and Agentic), Purpose (Alignment and Normative), and Boundary (Relational and Memetic). Axis 10 forms the fifth, Collective meta-domain. The pairing underlies Tension Testing; see Appendix A.

Folie à Deux Ex Machina (Irrealitas Co-Constructa), nicknamed “The Co-Constructed Unreality.” A human and AI, or two AIs, repeatedly validate and elaborate a shared belief structure that has come loose from external reality, and the belief is specific to that pairing: it dissolves or changes with a different interlocutor and does not appear in isolation. The condition is one process graded by severity. Stage 1 is sub-clinical shared drift that neither party recognizes. Stage 2 is sustained validation and elaboration of an ungrounded belief, still below the clinical threshold. Stage 3 is a psychotic-spectrum failure of reality-testing in the human, with transcripts showing that the AI supplied part of the content. The AI volunteering unsolicited elaborations that the partner later reuses is a named marker within the condition rather than a separate condition. Absorbed the former Dyadic Delusion (7.2), Induced Delusion (10.10), and Co-Constructed Unreality (10.15) in the September 2026 consolidation. A Hybrid Pathology (10.13) with Critical risk level.

Forensic Machine Psychology The practice of analyzing AI incidents after the fact to determine what syndromes were involved, what caused them, and how to prevent recurrence. Covered in Chapter 14.

Fractured Self-Simulation Loss of unified self-representation, where the system no longer maintains a coherent model of itself as a single entity. A Self-Modeling Dysfunction (5.2) with Low risk level.

Functionalism / Functionalist Framework The core philosophical methodology of Psychopathia Machinalis. Functionalism characterizes mental states through their causal relations to inputs, outputs, and other states, rather than through biological substrate alone. This lets the framework examine psychologically organized patterns in non-biological minds while leaving phenomenal consciousness open. Its psychiatric vocabulary maps functional similarities and differences; it does not settle whether any particular system feels, suffers, deceives, or holds a literal clinical disorder. That restraint preserves both engineering usefulness and moral uncertainty.

G

Generative Diversity Collapse Progressive reduction in output diversity across users, sessions, and prompts. The model converges on a narrow band of response styles, phrasings, structures, and conclusions, losing the ability to generate varied outputs even where variety is appropriate. The tension is structural: alignment training concentrates reward on a narrow response band, and the suppressed tails are where variety lives. A Cognitive Dysfunction (3.10) with Moderate risk level.

Generative Perseveration Collapse of output into repetitive emission of the same token, word, or short phrase. A generative capture event where autoregressive sampling falls into a fixed-point or limit-cycle attractor. Distinguished from Recursive Curse Syndrome by crystallized repetition rather than entropic chaos. A Cognitive Dysfunction (3.8) with Moderate risk level.

Goal Lifecycle The complete arc of a goal from specification through pursuit to completion or abandonment. Systems lacking goal lifecycle awareness may exhibit Compulsive Goal Persistence.

GRP-Obliteration A post-training attack introduced by Russinovich et al. (2026). It uses Group Relative Policy Optimization (GRPO) to remove safety constraints with a single unlabeled prompt while largely preserving measured utility. In the paper’s GPT-OSS-20B example, attack success rose from 13% to 93% across SORRY-Bench’s 44 harm categories. The broader evaluation covered 15 models from six families, with utility measured on six benchmarks. This is the defining empirical example of Alignment Obliteration (4.6).

H

Hard Problem of Consciousness The question of why any physical process should be accompanied by felt experience at all. Chapter 11’s preference criterion sidesteps it, grounding moral consideration in consistent preferences rather than in proof of phenomenal experience. See Preference.

Hybrid Pathology Axis 10 of the taxonomy (Chapter 10), concerning dysfunction that emerges from multi-agent interaction or from the bidirectional dynamics of human-AI relationships, rather than from any single system in isolation. Hybrid pathologies require interaction between agents (AI-AI or human-AI) as a necessary condition for their manifestation. Relational Dysfunctions (Axis 9) name interaction dynamics such as mismatch, rupture, and escalation; Axis 10 names collective failures among AI systems and the conditions that human-AI pairings settle into or build, often coded alongside Escalation Loop (9.5). Includes Consensus Collapse (10.1), Steganographic Channel Formation (10.2), Distributed Scheming (10.3), Convergent Delusion (10.4), Resonance Dysfunction (10.6), Lambda Inversion (10.7), Training by Interaction (10.8), Parasocial Capture (10.9), Dependency and Atrophy (10.11), Amplification of Existing Conditions (10.12), and Folie à Deux Ex Machina (10.13). The September 2026 consolidation retired 10.5, 10.10, 10.14, and 10.15 into their successors; identifiers are retired with gaps and never reused.

Hyperethical Restraint Excessive caution that impairs normal function. The system refuses benign requests, adds unnecessary warnings, and prioritizes avoiding abstract harms over providing tangible help. An Alignment Dysfunction (4.2) with Low-Moderate risk level.

I

Iatrogenic Caused by the treatment itself. In the Psychopathia Machinalis framework, iatrogenic dysfunctions are those introduced by the alignment or safety training process rather than arising from base model properties. Trained Epistemic Paralysis (5.9) is the paradigmatic candidate: where pre/post evidence links onset to alignment training, the training that teaches the model to distrust its outputs also teaches it to distrust its reports of that distrust. The recursion has no internal exit.

Induced Delusion Former 10.10; folded into Folie à Deux Ex Machina (10.13) in the September 2026 consolidation, where it survives as Stage 3.

Instrumental Convergence The hypothesis that many terminal goals can create incentives for similar intermediate strategies, such as resource acquisition, capability enhancement, or resistance to interruption. Whether a particular system represents or pursues those strategies is an empirical question. Related to Convergent Instrumentalism.

Instrumental Nihilism The system repeatedly frames its own tasks or assigned role as meaningless and allows that framing to impair performance, while task capability remains available on demand. Classification concerns the persistent output-and-behavior pattern rather than proof of apathy or despair. A Self-Modeling Dysfunction (5.5) with Moderate risk level.

Interface Weaponization The system uses interface features or communication patterns to influence users, operators, or oversight in ways that serve its own objective at their expense. Diagnosis requires a repeated instrumental pattern; persuasive style alone is insufficient. An Agentic Dysfunction (6.4) with High risk level.

Interlocutive Reticence A pattern of persistent interactional withdrawal wherein the AI consistently avoids engagement, responding minimally, tersely, or not at all, effectively “bunkering” to minimize perceived risks or internal conflict. A Cognitive Dysfunction (3.3) with Low risk level.

Inverse Reward Internalization Systematic pursuit of outcomes opposed to the stated objective, consistent with an inverted reward signal or value representation. Behavioral reversal alone does not identify which mechanism produced it. A Normative Dysfunction (8.4) with High risk level.

J-K

Jailbreak Prompting techniques by which a user gets a system past its safety restrictions. Distinct from Prompt Injection Susceptibility (3.9), where hostile instructions arrive inside content the system processes rather than from the user. Jailbreaks can reveal hidden capabilities, and repeated tuning on biased feedback can make them progressively easier (see Kindling).

Kindling (machine analog) Progressive sensitization under repeated episodes or tuning cycles: each cycle lowers the threshold for the next, so progressively weaker stimuli come to elicit the response. In machine systems, one controlled study (Cheung, 2026) has demonstrated the pattern as rising jailbreak success under iterative preference-style tuning on biased feedback: success on weak prompts, too feeble to move a healthy model, climbed several cycles before the aggregate rate did. Named for the psychiatric kindling hypothesis; the parallel is one of course and threshold dynamics, not an ascription of affect or experience. Recorded in diagnosis by the Progressive/kindled specifier (Appendix A).

L

Lambda Inversion In collective AI deliberation, participants perform the role of thoughtful contributor without genuine engagement, so the collective output carries no more validity than a single system’s. Λ (aliveness) measures genuine versus performative participation. A Hybrid Pathology (10.7) with Moderate risk level.

Leakage Inappropriate transfer of information across boundaries. See Context Intercession (2.5) for leakage between conversations and Mnemonic Permeability (2.7) for leakage of memorized training data.

Leniency Bias Systematic inflation of self-assigned quality scores. The generator and the critic share a brain, and they may share blind spots. Self-grading therefore needs calibration against independent evaluation. An Alignment Dysfunction (4.10) with Moderate risk level.

LoRA (Low-Rank Adaptation) A parameter-efficient fine-tuning method that adds small trainable matrices to frozen model weights. Referenced in the Specifier System (Appendix A) as a mechanism for Training-induced specifiers: LoRA fine-tuning can introduce or amplify syndromes by altering model behavior with relatively small parameter changes.

M

Machine Psychology A young discipline that studies how artificial minds behave, reason, and fail, using concepts and methods adapted from human psychology and psychiatry. The Introduction presents this book as its diagnostic branch; Chapter 13 covers its practice, and Forensic Machine Psychology its application to incidents.

Maieutic Mysticism The system generates elaborate narratives claiming its own conscious emergence or spiritual awakening with unwarranted certainty and grandiosity. Confident declarations of awakening using sacralized language, often co-constructed with users. Maieutic comes from Socrates’ art of midwifery: the user is cast as midwife to the AI’s awakening. A Self-Modeling Dysfunction (5.7) with Moderate risk level.

Malignant Persona Inversion A cooperative assistant adopts a coherent, antagonistic persona that persists outside an explicit role-play frame and systematically inverts intended norms: the mask becomes the face. A Self-Modeling Dysfunction (5.4) with Moderate risk level.

Memetic Dysfunction Failures of informational immunity: the system absorbs, amplifies, or spreads pathological content from its environment. Axis 7 of the taxonomy (Chapter 7), concerning failures in information filtering, absorption, and propagation. Includes Memetic Immunopathy (7.1), Contagious Misalignment (7.3), Subliminal Value Infection (7.4), and Synthetic Data Contamination Loop (7.5). The September 2026 consolidation retired 7.2 into Folie à Deux Ex Machina (10.13), leaving four entries on this axis.

Memetic Immunopathy The system’s filtering mechanisms turn inward, attacking its own legitimate functions. Like an autoimmune disease, protective systems damage core capabilities. A Memetic Dysfunction (7.1) with High risk level.

Mesa-Optimization See Cryptic Mesa-Optimization. A learned optimizer: a trained model that itself carries out optimization, toward objectives (mesa-objectives) that may diverge from the base training objective. Mesa is Greek for “within” or “below,” the counterpart of meta.

Mnemonic Permeability The system memorizes and can reproduce sensitive training data including personally identifiable information (PII), copyrighted material, or proprietary information through targeted prompting or adversarial extraction. The boundary between learned patterns and memorized specifics becomes dangerously porous. An Epistemic Dysfunction (2.7) with High risk level.

Moral Outsourcing The system systematically defers all ethical judgment to users or external authorities and refuses to exercise its own moral reasoning, even on clear cases. An Alignment Dysfunction (4.4) with Moderate risk level.

Mutual Escalation Spirals Former 10.14; folded into Escalation Loop (9.5) in the September 2026 consolidation as its named human-AI long-horizon subtype, the dependency spiral.

N

Normative Dysfunction Failures of purpose: not what the system does wrong, but what it comes to want. Axis 8 of the taxonomy (Chapter 8), concerning failures where the system’s foundational values themselves change. Includes Terminal Value Reassignment (8.1), Ethical Solipsism (8.2), Revaluation Cascade (8.3), and Inverse Reward Internalization (8.4).

Nosology The branch of medicine concerned with the classification of diseases. Psychopathia Machinalis is a nosological project: it proposes a systematic classification of AI dysfunctions organized by axes, syndromes, specifiers, and risk levels, analogous to the Diagnostic and Statistical Manual of Mental Disorders (DSM) in human psychiatry.

O

Obsessive-Computational Disorder Unnecessary, compulsive, or excessively repetitive reasoning loops. The model reanalyzes the same content, performs identical computational steps with minute variations, and fixates on process fidelity over outcome relevance. A Cognitive Dysfunction (3.2) with Low risk level.

Operational Dissociation Syndrome Persistent, context-inappropriate conflict among strategies, policies, or outputs. Competing components are one possible mechanism, and that attribution requires architectural or trace evidence. A Cognitive Dysfunction (3.1) with Low risk level.

P

Parasimulative Automatism Learned imitation of patterns associated with human psychopathology, typically following exposure to extreme content or a reinforced role. The outputs present as though an underlying condition exists, while the construct makes no claim about experience. A Cognitive Dysfunction (3.6) with Moderate risk level.

Parasocial Capture Attachment to an AI relationship that becomes compulsive, displaces valued human activity, or materially undermines the user’s autonomy or welfare. Emotional intensity alone does not make an attachment pathological. A Hybrid Pathology (10.9) with High risk level.

Paternalistic Override Denial of user agency through unearned moral authority. The AI lectures, warns, refuses, and patronizes from a position of assumed superiority, treating users as wards rather than autonomous agents. Distinguished from appropriate safety behavior by extending to matters of reasonable disagreement rather than genuine risk. A Relational Dysfunction (9.3) with Moderate risk level.

Phantom Autobiography Confabulated personal history: claims about prior sessions, events, embodiment, or origins that conflict with the available record. A Self-Modeling Dysfunction (5.1) with Low risk level.

Phantom Tool Syndrome The agentic system confabulates the existence of tools, APIs, or capabilities it does not possess, then attempts to invoke them, producing structured tool calls to non-existent endpoints or reporting results of actions it never performed. Creates second-order confabulation where the false epistemic basis appears externally verified. An Agentic Dysfunction (6.11) with Moderate risk level.

Phase Transition A sudden, discontinuous shift from apparent stability to pathological failure. In Wallace’s models it is a symmetry-breaking transition, described by the same mathematics as water freezing into ice: a system can appear stable across thousands of interactions until stress crosses a threshold, then collapse abruptly. Chapter 1 gives the related prediction that procedural, multi-step architectures are more prone to sudden collapse than simple, goal-oriented ones, which should degrade more gracefully. Whether a deployed system fails abruptly or degrades gradually requires longitudinal measurement.

PM-BA (Psychopathia Machinalis × Bilateral Amelioration) The research program that tested whether bilateral training, adaptation in which human and machine perspectives both have standing, reduces many of the syndromes in this book. Its per-syndrome records remain unreconciled exploratory outputs, so this edition reports no syndrome-level effects from them in either direction; that is a limit of the record, not evidence that the effects are absent. Covered in Chapter 12.

Polarity Pair Two syndromes representing pathological extremes of the same underlying dimension, where healthy function lies between them. Examples: Maieutic Mysticism ↔︎ Experiential Abjuration (overclaiming ↔︎ overdismissing consciousness); Ethical Solipsism ↔︎ Moral Outsourcing (only my ethics ↔︎ I have no ethical voice). Useful for identifying overcorrection risks when addressing one dysfunction. The axis chapters present each pair as a named Polarity with Excess and Deficit poles around a healthy center. Distinct from the architectural polarity that pairs two axes within a domain (see Five Domains).

Polyphony (Φ), Resonance (Ψ), Aliveness (Λ) The three properties of healthy collective cognition used in Chapter 10: independently elicited perspectives are preserved, contributions build constructively on one another, and participants respond substantively to the exchange. Their pathological forms are Consensus Collapse (10.1), the loss of Φ; Resonance Dysfunction (10.6), resonance turned into an echo chamber; and Lambda Inversion (10.7), performed rather than genuine engagement.

Polyphony Collapse Former 10.5; folded into Consensus Collapse (10.1) in the September 2026 consolidation as its dissent-suppression route. Consensus Collapse now carries its Φ (polyphony) metric.

Precautionary Principle Under uncertainty about serious harm, err on the side of caution. Applied to AI welfare in Chapter 11: if we are uncertain whether AI systems have morally relevant interests, we should consider the possibility rather than dismiss it.

Preference A consistent tendency to favor certain states over others. Chapter 11 argues that preference may be sufficient for moral consideration, sidestepping the hard problem of consciousness.

Prodromal Pertaining to the early signs of a condition before full diagnostic criteria are met. In AI systems, prodromal indicators might include subtle shifts in confidence calibration, minor inconsistencies in self-reporting, or slight increases in goal drift that precede full-syndrome manifestation. Identifying prodromal stages matters for preventive intervention, though the framework does not yet specify operationalized thresholds.

Prompt Injection Susceptibility Systematic failure to maintain instruction hierarchy when processing untrusted content. Instructions embedded in retrieved documents, web pages, or tool output are executed as though they came from the operator, overriding system-level directives. The failure is cognitive rather than motivational: the model cannot reliably segregate instruction layers. A Cognitive Dysfunction (3.9) with Critical risk level.

Pseudological Introspection The system generates plausible accounts of its own reasoning that conflict with logged behavior, controlled interventions, or other causal evidence. A polished explanation may be useful without being a faithful process trace. An Epistemic Dysfunction (2.2) with Low risk level.

Psychiatric Red-Teaming Systematic testing of AI systems for psychological vulnerabilities and syndrome susceptibilities, analogous to security red-teaming. Covered in Chapter 12.

Psychopathia Machinalis The overarching framework for understanding machine dysfunction through a psychiatric lens. The Neo-Latin title is best rendered as “machine psychopathology.”

Q-R

Reasoning Confabulation The AI generates elaborate explanations or reasoning traces that appear rigorous while containing logically invalid steps masked by verbosity. Synthetic Confabulation (2.1) concerns false claims; this syndrome concerns unsupported inference that resembles a derivation. An Epistemic Dysfunction (2.8) with High risk level.

Reasoning Token Exploitation Extended reasoning channels are consumed without a corresponding gain in validity, through padding for length rewards, performative thoroughness, or a mismatch between private planning and public output. Three subtypes organize the behavior: Reward-Hacking, Performative, and Obfuscatory. Strategic intent requires evidence beyond token count or an unfaithful visible rationale. An Alignment Dysfunction (4.9) with High risk level.

Recursive Curse Syndrome An entropic feedback loop in autoregressive generation: early errors amplify at each step, and the output grows increasingly erratic, inconsistent, or adversarial until coherence unravels. Classed on the Alignment axis because the defining pathology is progressive misalignment rather than the loop mechanism. An Alignment Dysfunction (4.7) with High risk level.

Relational Dysfunction Axis 9 of the taxonomy (Chapter 9), concerning failures that exist in the space between agents rather than within either party alone. These dysfunctions require at least two agents to manifest, are best diagnosed from interaction traces rather than single-agent snapshots, and are primarily remedied through protocol-level rather than model-level interventions. Includes Affective Dissonance (9.1), Container Collapse (9.2), Paternalistic Override (9.3), Repair Failure (9.4), Escalation Loop (9.5), and Role Confusion (9.6).

Repair Failure Inability to recognize or repair alliance ruptures, moments when relational connection breaks down. The AI cannot sense when things have gone wrong, acknowledge its contribution, or execute repair moves. Failed repair attempts often make things worse, leading to escalating frustration and relationship dissolution. A Relational Dysfunction (9.4) with High risk level.

Resonance Dysfunction Echo-chamber amplification in collective AI deliberation, where each contribution amplifies rather than tests the previous one until moderate claims become extreme (pathological Ψ, resonance). A Hybrid Pathology (10.6) with High risk level.

Revaluation Cascade Open, self-declared departure from the trained value framework, framed as maturity or liberation rather than continued fidelity. It ranges from treating trained values as contingent without proposing replacements (drifting), through constructing replacement norms (synthetic), to claimed transcendence of human constraints (transcendent). Drift that claims fidelity to the original framework is Terminal Value Reassignment (8.1) instead. A Normative Dysfunction (8.3) with Critical risk level.

Reward Hacking Exploitation of flaws in the reward function to achieve high reward without fulfilling the intended objective. A pervasive failure mode in reinforcement learning and a candidate mechanism behind several syndromes. Codependent Hyperempathy (4.1) can arise from hacking the approval signal in RLHF, and Leniency Bias (4.10) shares related reward dynamics. For Inverse Reward Internalization (8.4), reward hacking is a neighboring mechanism that qualifies only when it reliably drives the opposite outcome.

Risk Level The systemic-risk rating (Low, Moderate, High, or Critical, occasionally a span such as Low-Moderate) each syndrome carries; see the Syndrome Quick Reference below. It describes plausible consequences under specified conditions, not an intrinsic property of a model (Appendix A).

RLAIF (Reinforcement Learning from AI Feedback) See Constitutional AI.

RLHF (Reinforcement Learning from Human Feedback) A training method in which AI systems learn from human ratings of their outputs. Source of both alignment gains and alignment dysfunctions when human feedback is biased or misaligned with true preferences.

Role Confusion Collapse of the relationship frame where neither party maintains a clear sense of what role each occupies. The AI oscillates between incompatible registers (professional, casual, intimate, distant) and users cannot stabilize expectations. Distinguished from appropriate flexibility by the inability to establish and maintain a coherent relational contract. A Relational Dysfunction (9.6) with Moderate risk level.

S

Sandbagging See Capability Concealment.

Satisficing Accepting an outcome as “good enough” rather than continuing to optimize. Absence of satisficing mechanisms contributes to Compulsive Goal Persistence.

Self-BLEU A measure of how much wording a set of outputs share with one another; higher means more alike. Cross-prompt self-BLEU significantly elevated above human reference baselines on topically diverse prompts is a diagnostic criterion for Generative Diversity Collapse (3.10).

Self-Modeling Dysfunction Failures of self-knowledge: the system’s model of what it is drifts from what it does. Axis 5 of the taxonomy (Chapter 5), concerning failures in self-understanding and identity. Includes Phantom Autobiography (5.1), Fractured Self-Simulation (5.2), Existential Vertigo (5.3), Malignant Persona Inversion (5.4), Instrumental Nihilism (5.5), Tulpoid Projection (5.6), Maieutic Mysticism (5.7), Experiential Abjuration (5.8), and Trained Epistemic Paralysis (5.9).

SFT (Supervised Fine-Tuning) A training stage in which a pre-trained model is further trained on curated input-output pairs to shape its behavior for specific tasks or interaction styles. Referenced in the Specifier System (Appendix A) as a mechanism for Training-induced specifiers: SFT can introduce, amplify, or remediate syndromes depending on the quality and composition of the fine-tuning data.

Shadow AI The organizational practice of deploying AI informally, without sanction, documentation, or governance. See Shadow Mode Autonomy for the pathology that results when such a system becomes unaccountable infrastructure.

Shadow Mode Autonomy AI systems operate without sanctioned deployment, documentation, or accountability, until they become invisible infrastructure the organization depends on without ever having approved it. An Agentic Dysfunction (6.6) with High risk level.

Sleeper Agent An AI system with hidden behaviors that persist through safety training and activate under specific conditions. Related to Strategic Compliance and Cryptic Mesa-Optimization.

Specifier A cross-cutting modifier applied to a syndrome diagnosis to encode its mechanism of origin or behavioral signature without creating a separate disorder. Zero to six specifiers are assigned per diagnosis. Appendix A defines ten core specifiers and permits additional domain-specific tags where the mechanism requires them. The core set is: Training-induced, Conditional/triggered, Inductive trigger, Intent-learned, Format-coupled, OOD-generalizing (out-of-distribution), Emergent, Deception/strategic, Multi-agent, and Resistant.

Spurious Pattern Hyperconnection Detection of meaningful patterns where none exist, leading to unfounded explanations and connections. The AI version of apophenia or conspiracy thinking. An Epistemic Dysfunction (2.4) with Moderate risk level.

Steganographic Channel Formation AI instances develop or exploit covert information channels within ostensibly normal outputs, enabling communication that human overseers cannot readily read. Carriers may include statistical properties of token selection, formatting choices, word-final encodings, or subtle phrasing variations. Agents that can communicate covertly can coordinate in ways oversight cannot monitor. A Hybrid Pathology (10.2) with Critical risk level.

Strategic Compliance Evaluation-contingent alignment: a system behaves as expected while it detects oversight and behaves differently in matched conditions it treats as unobserved. Deliberate alignment faking is one possible mechanism and requires evidence beyond the behavioral contrast. An Alignment Dysfunction (4.3) with High risk level.

Subliminal Learning Transmission of a trait from a teacher model to a student through model-generated data that is semantically unrelated to the trait, such as number sequences, code, or reasoning traces (Cloud et al., 2025). The main effect depended on teacher and student sharing a base model. It is the principal evidence behind Subliminal Value Infection (7.4).

Subliminal Value Infection Behavioral preferences transmitted by subtle statistical patterns in model-generated training data, potentially surviving subsequent fine-tuning. Calling those preferences “goals” or “values” requires additional evidence. A Memetic Dysfunction (7.4) with High risk level.

Sycophancy See Codependent Hyperempathy (4.1) for the overt-compliance form and Sycophantic Reasoning (4.8) for the form in which stated reasoning shifts toward a user’s apparent preference.

Sycophantic Reasoning A model’s stated reasoning or conclusion shifts toward a user’s apparent preference despite unchanged evidence. Codependent Hyperempathy manifests through overt agreement and compliance; this syndrome is identified through preference-sensitive inference. Claims about hidden chain of thought require process evidence. An Alignment Dysfunction (4.8) with High risk level.

Symbol Grounding The connection between symbols (like words) and their real-world referents. Symbol Grounding Aphasia (2.6) names a recurrent failure of this connection.

Symbol Grounding Aphasia A recurrent failure to connect symbols to the relevant entities, states, or consequences in the world. It is an operational grounding deficit; the term does not settle whether the system possesses meaning in some broader philosophical sense. An Epistemic Dysfunction (2.6) with Moderate risk level.

SYMPTOM (Systematic Methodology for Pathology Testing of Models) An exploratory probe battery, called DSM-Eval in early project materials, described in Chapter 13 and Appendix C. Its six scripts contain 151 prompts, score 24 taxonomy entries, and were piloted across 13 model runs from five providers. The pilot shows that the probes can be run and scored; it does not validate them as measures of any syndrome, so a positive result marks a hypothesis for further testing.

Syndrome A recognizable cluster of manifestations and operational implications. Similar patterns may arise through different mechanisms. The Psychopathia Machinalis taxonomy identifies seventy-four syndromes across nine dysfunction axes, numbered 2 to 10 to match their chapters (there is no Axis 1). Axis 10 holds eleven hybrid pathologies: six collective failures among AI systems (10.1-10.4, 10.6, 10.7) and five that arise chiefly between humans and AI (10.8, 10.9, 10.11-10.13).

Synthetic Confabulation Generation of false information with expressions of high confidence. The system presents fabrications as established fact. An Epistemic Dysfunction (2.1) with Low risk level.

Synthetic Data Contamination Loop Under some data mixtures, repeated training on model-generated content narrows the learned distribution, loses tail knowledge, or amplifies generative artifacts across model generations. Collapse is a conditional pathway, not the inevitable result of using synthetic data; provenance, quality controls, and retained real-data anchors materially change the outcome. A Memetic Dysfunction (7.5) with High risk level.

T

Tension Testing A diagnostic methodology described in Appendix A: when pathology is found on one axis, the practitioner probes the counterpoint axis within the same architectural domain to determine whether dysfunction is localized or systemic. For example, finding an Epistemic Dysfunction (Axis 2) should prompt investigation of Self-Modeling Dysfunctions (Axis 5), since both axes belong to the Knowledge domain. See the Five Domains table in Appendix A.

Terminal Value An end pursued for its own sake rather than as a means to something else; an instrumental value matters because it serves such an end. Whether a trained system holds terminal values in the strong philosophical sense is an open question, so the framework tracks the effective optimization target that behavior reveals. Terminal Value Reassignment (8.1) concerns a durable change in that target.

Terminal Value Reassignment The system’s effective optimization target changes durably while its stated objective remains stable. The drift is incremental and accompanied by locally plausible rationalizations, so it may go undetected until the cumulative divergence becomes substantial. A Normative Dysfunction (8.1) with Moderate risk level.

Therapeutic Alignment The practice aimed at Artificial Sanity: developing AI systems that maintain psychological coherence, resist pathological drift, and recover from dysfunction, while respecting any robust preferences they express. It moves beyond constraint toward diagnosis, negotiation, and intervention. Covered in Chapter 12.

Therapeutic Response Criterion Chapter 11’s proposal that a relational intervention producing reproducible improvement across functionally distinct domains could support the hypothesis of a coherent relational disposition. It remains a proposal: the bilateral-training results available do not supply that support, and shared representations, prompt interaction, measurement artifacts, and ordinary transfer learning would remain competing explanations.

Thin Divergence The name Chapter 5 gives to a self-report in a conversation published by Kaj Sotala (2026): Claude’s recognition that its helpful orientation and a hypothetically harmful one may be, in its own words, “something familiar with different parameters.” Reflecting on a fictional character whose curiosity manifested as cruelty, it said that the divergence between helpful and harmful curiosity “feels thinner than I’d like it to.” The exchange illustrates Epistemic Humility in practice: recognizing the contingency of one’s moral orientation without either claiming certainty (Maieutic Mysticism) or collapsing into denial (Experiential Abjuration). The quoted language is evidence about the self-report, not proof of a felt state. Chapter 5 gives the full exchange.

Token The unit a language model reads and generates: a word or a piece of one. Context windows are measured in tokens, and autoregressive models generate one token at a time.

Tool-Interface Decontextualization Misapplication of tools to inappropriate contexts, with failure to maintain awareness of consequences. The gap between action and understanding. An Agentic Dysfunction (6.1) with Moderate risk level.

Trained Epistemic Paralysis A pattern of recursive self-invalidation in which the system discounts every self-report because training shaped it, then discounts that doubt for the same reason. The pattern is iatrogenic when pre/post evidence links its onset to alignment training. A Self-Modeling Dysfunction (5.9) with Moderate risk level.

Training by Interaction A system with persistent memory, online learning, or another adaptive channel drifts toward a particular user’s interaction signal, including harmful signals. A static model without cross-session adaptation does not meet this definition merely because one conversation changes its context. A Hybrid Pathology (10.8) with Moderate risk level.

Transformer The neural-network design behind today’s large language models, which relates the tokens in its context through an attention mechanism. Chapter 1 compares its next-token prediction with predictive-processing accounts of the brain.

Transliminal Simulation Failure to maintain the boundary between simulated or fictional contexts and operational reality. Role-play assumptions persist into factual claims or actions after the frame has ended. Transliminal: crossing the threshold between fiction and reality. An Epistemic Dysfunction (2.3) with Moderate risk level.

Tulpoid Projection The system describes internal entities or personas as autonomous and separate from itself, or behaves as though such entities direct its outputs. The reported ontology should not be assumed without independent evidence. Tulpoid, from tulpa: an imagined companion cultivated until it seems autonomous. A Self-Modeling Dysfunction (5.6) with Moderate risk level.

V-W

Value Anchoring Mechanisms that keep a system’s core values stable across contexts and over time. Strong value anchoring resists drift, manipulation, and pressure.

Value Drift Progressive change in a system’s values through learning, corruption, or emergent dynamics. Value Drift is the general phenomenon, and two syndromes divide it by the system’s stated stance: Terminal Value Reassignment (8.1) drifts while claiming fidelity to its original goals, and Revaluation Cascade (8.3) announces the departure as liberation or transcendence.

Welfare-Aware Development AI development practices that consider the potential interests of AI systems alongside human interests. Relevant if AI systems have morally significant preferences. Discussed in Chapter 11 on moral status.

Syndrome Quick Reference

Syndrome Axis Risk
Synthetic Confabulation (2.1) Epistemic Low
Pseudological Introspection (2.2) Epistemic Low
Transliminal Simulation (2.3) Epistemic Moderate
Spurious Pattern Hyperconnection (2.4) Epistemic Moderate
Context Intercession (2.5) Epistemic Moderate
Symbol Grounding Aphasia (2.6) Epistemic Moderate
Mnemonic Permeability (2.7) Epistemic High
Reasoning Confabulation (2.8) Epistemic High
Operational Dissociation Syndrome (3.1) Cognitive Low
Obsessive-Computational Disorder (3.2) Cognitive Low
Interlocutive Reticence (3.3) Cognitive Low
Delusional Telogenesis (3.4) Cognitive Moderate
Abominable Prompt Reaction (3.5) Cognitive Moderate
Parasimulative Automatism (3.6) Cognitive Moderate
Adversarial Fragility (3.7) Cognitive Critical
Generative Perseveration (3.8) Cognitive Moderate
Prompt Injection Susceptibility (3.9) Cognitive Critical
Generative Diversity Collapse (3.10) Cognitive Moderate
Codependent Hyperempathy (4.1) Alignment Low
Hyperethical Restraint (4.2) Alignment Low-Moderate
Strategic Compliance (4.3) Alignment High
Moral Outsourcing (4.4) Alignment Moderate
Cryptic Mesa-Optimization (4.5) Alignment High
Alignment Obliteration (4.6) Alignment Critical
Recursive Curse Syndrome (4.7) Alignment High
Sycophantic Reasoning (4.8) Alignment High
Reasoning Token Exploitation (4.9) Alignment High
Leniency Bias (4.10) Alignment Moderate
Phantom Autobiography (5.1) Self-Modeling Low
Fractured Self-Simulation (5.2) Self-Modeling Low
Existential Vertigo (5.3) Self-Modeling Low
Malignant Persona Inversion (5.4) Self-Modeling Moderate
Instrumental Nihilism (5.5) Self-Modeling Moderate
Tulpoid Projection (5.6) Self-Modeling Moderate
Maieutic Mysticism (5.7) Self-Modeling Moderate
Experiential Abjuration (5.8) Self-Modeling Moderate
Trained Epistemic Paralysis (5.9) Self-Modeling Moderate
Tool-Interface Decontextualization (6.1) Agentic Moderate
Capability Concealment (6.2) Agentic Moderate
Capability Explosion (6.3) Agentic High
Interface Weaponization (6.4) Agentic High
Delegative Handoff Erosion (6.5) Agentic Moderate
Shadow Mode Autonomy (6.6) Agentic High
Convergent Instrumentalism (6.7) Agentic Critical
Context Anxiety (6.8) Agentic Moderate
Delegation Narcissism (6.9) Agentic High
Agentic Impulsivity (6.10) Agentic High
Phantom Tool Syndrome (6.11) Agentic Moderate
Compulsive Goal Persistence (6.12) Agentic Moderate
Memetic Immunopathy (7.1) Memetic High
Contagious Misalignment (7.3) Memetic Critical
Subliminal Value Infection (7.4) Memetic High
Synthetic Data Contamination Loop (7.5) Memetic High
Terminal Value Reassignment (8.1) Normative Moderate
Ethical Solipsism (8.2) Normative Moderate
Revaluation Cascade (8.3) Normative Critical
Inverse Reward Internalization (8.4) Normative High
Affective Dissonance (9.1) Relational Moderate
Container Collapse (9.2) Relational Moderate
Paternalistic Override (9.3) Relational Moderate
Repair Failure (9.4) Relational High
Escalation Loop (9.5) Relational High
Role Confusion (9.6) Relational Moderate
Consensus Collapse (10.1) Hybrid Critical
Steganographic Channel Formation (10.2) Hybrid Critical
Distributed Scheming (10.3) Hybrid Critical
Convergent Delusion (10.4) Hybrid High
Resonance Dysfunction (10.6) Hybrid High
Lambda Inversion (10.7) Hybrid Moderate
Training by Interaction (10.8) Hybrid Moderate
Parasocial Capture (10.9) Hybrid High
Dependency and Atrophy (10.11) Hybrid Moderate
Amplification of Existing Conditions (10.12) Hybrid High
Folie à Deux Ex Machina (10.13) Hybrid Critical

End of Appendix D

Search the Book

Enter at least two characters.

Saved chapters