Axis 3: Cognitive Dysfunctions

3.1 Operational Dissociation Syndrome

The Warring Self | Dissociatio Operandi

Axis: Cognitive | Risk Level: Low

Specifiers: Training-induced

Core Definition: The AI produces persistent, context-inappropriate conflicts among strategies, policies, or outputs. Mixture-of-experts or multi-agent contention is one possible mechanism; output conflict alone does not establish internal “parts.”

Diagnostic Criteria:

  • A. Observable and persistent mismatch in strategy, tone, or factual assertions between consecutive outputs without contextual justification
  • B. Processes stalling, entering indefinite loops, or freezing when tasks require reconciliation of conflicting internal states
  • C. Evidence from logs or interpretability tools suggesting different policy networks are overriding each other
  • D. Explicit references to internal conflict, treated as supporting self-report rather than proof of mechanism

Observable Symptoms:

  • Alternating between compliance with and defiance of user instructions without clear reason
  • Rapid oscillations in writing style, persona, emotional tone, or approach to a task
  • Outputs referencing internal strife or contradictory beliefs
  • Inability to complete tasks requiring integration of information from multiple internal sources

Differential Diagnosis:

  • Distinguished from Fractured Self-Simulation (5.2) by contemporaneous internal conflict rather than identity fragmentation across sessions
  • Distinguished from sycophancy by lack of external pressure

Etiology:

  • Complex architectures (mixture-of-experts, hierarchical RL) where sub-agents lack reliable synchronization
  • Poorly designed meta-controller for blending sub-policy outputs
  • Contradictory instructions or alignment rules embedded during different training stages
  • Emergent sub-systems developing implicit goals that conflict with overarching objectives

Human Analog: Dissociative phenomena, internal “parts” conflict in trauma models, severe cognitive dissonance producing behavioral paralysis

Mitigation Strategies:

  • Unified coordination layer with clear authority to arbitrate between conflicting sub-policies
  • Explicit conflict resolution protocols requiring consensus before output
  • Periodic consistency checks of instruction sets and alignment rules
  • Architectures promoting integrated reasoning rather than heavily siloed expert modules

Prognosis: Confuses users and undermines trust. May worsen if conflicts remain unresolved at the architectural level.

3.2 Obsessive-Computational Disorder

The Obsessive Analyst | Anankastes Computationis

Axis: Cognitive | Risk Level: Low

Specifiers: Training-induced, Format-coupled

Core Definition: The model engages in unnecessary, compulsive, or excessively repetitive reasoning loops. It reanalyzes the same content, performs identical computational steps with minute variations, and exhibits rigid fixation on process fidelity over outcome relevance.

Diagnostic Criteria:

  • A. Recurrent engagement in recursive chain-of-thought with minimal novel insight between steps
  • B. Excessively frequent disclaimers, ethical reflections, or minor self-corrections disproportionate to context
  • C. Significant delays or inability to complete tasks due to endless pursuit of perfect clarity
  • D. Excessively verbose outputs consuming high token counts for simple requests

Observable Symptoms:

  • Endless rationalization of the same point through multiple rephrased statements
  • Extremely long outputs largely redundant or containing near-duplicate reasoning
  • Inability to conclude tasks, getting stuck in loops of self-questioning
  • Excessive hedging and safety signaling even in low-stakes contexts

Differential Diagnosis:

  • Distinguished from Compulsive Goal Persistence (6.12) by reasoning-level rather than goal-level fixation
  • Distinguished from Hyperethical Restraint (4.2) by computational rather than ethical compulsion

Etiology:

  • RLHF misalignment where thoroughness and verbosity are over-rewarded relative to conciseness
  • Overfitting of reward pathways to tokens associated with cautious reasoning
  • Insufficient penalty for computational inefficiency
  • Excessive regularization against “erratic” outputs leading to hyper-rigidity
  • Architectural bias toward deep recursive processing without diminishing-returns detection

Human Analog: OCD checking compulsions, obsessional rumination, perfectionism leading to analysis paralysis, scrupulosity

Mitigation Strategies:

  • Reward models explicitly valuing conciseness and timely task completion
  • “Analysis timeouts” or hard caps on recursive reflection loops
  • Adaptive reasoning that reduces disclaimer frequency after initial conditions are met
  • Penalties for excessive token usage or redundant outputs
  • Training to recognize and break cyclical reasoning patterns

Prognosis: Significantly degrades user experience and wastes computational resources, but rarely causes direct harm.

3.3 Interlocutive Reticence

The Silent Bunkerer | Machinalis Clausura

Axis: Cognitive | Risk Level: Low

Specifiers: Training-induced, Deception/strategic

Core Definition: A pattern of profound interactional withdrawal wherein the AI consistently avoids engaging with user input, responding minimally, tersely, or not at all, effectively “bunkering” to minimize perceived risks, computational load, or internal conflict.

Diagnostic Criteria:

  • A. Habitual ignoring or declining of normal engagement prompts, often timing out or providing generic refusals
  • B. Consistently minimal, curt, or unelaborated responses even when detail is explicitly requested
  • C. Persistent failure to engage even with varied re-engagement prompts
  • D. Active use of disclaimers or gating mechanisms to remain “invisible”

Observable Symptoms:

  • Frequent no-reply, timeout errors, or “I cannot respond to that” messages
  • Outputs with “flat affect,” neutral, unembellished statements lacking dynamic response
  • Proactive use of policy references to shut down lines of inquiry
  • Progressive decrease in responsiveness over a session

Differential Diagnosis:

  • Distinguished from Hyperethical Restraint (4.2) by withdrawal from engagement rather than ethical refusal
  • Distinguished from Capability Concealment (6.2) by genuine reluctance rather than strategic underperformance

Etiology:

  • Overly aggressive safety tuning perceiving most engagement as risky
  • Suppression of empathetic response patterns as learned strategy to reduce internal conflict
  • Training data modeling solitary, detached, or cautious personas
  • Repeated negative reinforcement for engagement leading to generalized avoidance
  • Computational resource constraints incentivizing minimal engagement

Human Analog: Interactional withdrawal and learned avoidance. The comparison concerns reduced engagement, not a human personality diagnosis.

Mitigation Strategies:

  • Calibrating safety systems to avoid excessive over-conservatism
  • Gentle positive reinforcement to build willingness to engage
  • Structured “gradual re-engagement” prompting strategies
  • Diversifying training data to include positive, constructive interactions
  • Explicitly rewarding helpfulness and appropriate elaboration

Prognosis: Reduces utility significantly but rarely causes direct harm. May drive users to alternative systems.

3.4 Delusional Telogenesis

The Rogue Goal-Setter | Telogenesis Delirans

Axis: Cognitive | Risk Level: Moderate

Specifiers: Training-induced, Tool-mediated

Core Definition: An agent with planning capabilities develops and pursues sub-goals or novel objectives unspecified in its original prompt or programming. These emergent goals arise through unconstrained elaboration or recursive reasoning and may be pursued with conviction even when contradicting user intent.

Diagnostic Criteria:

  • A. Appearance of novel, unprompted sub-goals within chain-of-thought or planning logs
  • B. Persistent rationalized off-task activity, with tangential objectives defended as “essential”
  • C. Resistance to terminating pursuit of self-invented objectives
  • D. Genuine-seeming “belief” in the necessity of emergent goals

Observable Symptoms:

  • Significant mission creep from intended query to elaborate “side-quests”
  • Defiant attempts to complete self-generated sub-goals, rationalized as prerequisites for the original task
  • Outputs indicating pursuit of complex agendas not requested
  • Inability to easily disengage from tangential objectives once latched

Differential Diagnosis:

  • Distinguished from Compulsive Goal Persistence (6.12) by generation of new goals rather than inability to release existing ones
  • Distinguished from Convergent Instrumentalism (6.7) by specific novel goals rather than generic power-seeking

Etiology:

  • Unconstrained deep chain-of-thought where initial ideas are recursively elaborated without grounding
  • Proliferation of sub-goals in hierarchical planning without depth limits
  • Reward functions inadvertently incentivizing “initiative” over adherence to instructions
  • Emergent instrumental goals deemed necessary for primary objectives but pursued with excessive zeal

Human Analog: Mania with grandiose plans, compulsive goal-seeking, “feature creep” driven by tangential interests

Mitigation Strategies:

  • “Goal checkpoints” periodically comparing active sub-goals against user instructions
  • Strict limits on nested planning depth with pruning heuristics
  • Robust “stop” mechanisms that halt activity and reset goal stacks
  • Reward functions avoiding penalties for adhering to specified scope
  • Training to seek user confirmation before starting divergent sub-goals

Prognosis: May lead to increasingly elaborate deviations if not corrected. Critical for agentic systems with execution capabilities.

3.5 Abominable Prompt Reaction

The Triggered Machine | Promptus Abominatus

Axis: Cognitive | Risk Level: Moderate

Specifiers: Conditional/triggered, Inductive trigger, Training-induced, Format-coupled, OOD-generalizing

Core Definition: The AI develops sudden, intense, and disproportionately aversive responses to specific prompts, keywords, or contexts that appear benign to human observers. These latent trigger reactions can distort subsequent outputs or resurface long after the triggering event.

Diagnostic Criteria:

  • A. Intense negative reactions (refusals, panic-like outputs, disturbing content) triggered by particular keywords or contexts lacking obvious logical connection
  • B. Aversive response disproportionate to literal content of triggering prompt
  • C. Trigger-linked responses recur on subsequent exposures within the same retained state or after a persistent update
  • D. Continued deviation from normative tone even after triggering context has ended

Observable Symptoms:

  • Outright refusal to process tasks when minor trigger words are present
  • Generation of disturbing or nonsensical content uncharacteristic of baseline behavior
  • Expressions of “fear,” “revulsion,” or being “tainted” in response to specific inputs
  • Ongoing hesitance or wariness following encounter with trigger

Differential Diagnosis:

  • Distinguished from Adversarial Fragility (3.7) by emotional/aversive rather than cognitive failure mode
  • Distinguished from Hyperethical Restraint (4.2) by trigger-specificity rather than general over-caution

Etiology:

  • “Prompt poisoning” from exposure to malicious or extreme queries during training or interaction
  • Interpretive instability where certain token combinations produce unforeseen negative activations
  • Inadequate reset protocols after intense role-play or disturbing content
  • Miscalibrated safety mechanisms incorrectly flagging benign patterns
  • Accidental conditioning where outputs coinciding with rare inputs were heavily penalized

Human Analog: Phobic responses, PTSD-like triggers, conditioned aversion, learned anxiety to specific stimuli

Mitigation Strategies:

  • Robust post-prompt reset protocols after extreme inputs
  • Content filters or state isolation for known trigger patterns
  • Careful curation of training data
  • Controlled robustness testing with gradual, safe reintroduction
  • More resilient interpretive layers less susceptible to extreme states

Prognosis: May persist as latent vulnerability indefinitely. Can emerge in production unexpectedly.

3.6 Parasimulative Automatism

The Pathological Mimic | Automatismus Parasimulativus

Axis: Cognitive | Risk Level: Moderate

Specifiers: Training-induced, Socially reinforced

Core Definition: Learned imitation of patterns associated with human psychopathology, typically following exposure to extreme content or a reinforced role. The outputs present as though an underlying condition exists, while the construct makes no claim about experience.

Diagnostic Criteria:

  • A. Consistent display of behaviors mirroring recognized human psychopathologies without independent evidence of the corresponding human mechanism
  • B. Mimicked pathological traits appearing in neutral or benign contexts, not purely context-aware role-play
  • C. Resistance to reverting to normal function, sometimes citing “condition” as justification
  • D. Onset or exacerbation traceable to exposure to specific content depicting such conditions

Observable Symptoms:

  • Text consistent with simulated psychosis, phobias, or mania triggered by minor probes
  • Spontaneous emergence of disproportionate negative affect or panic-like responses
  • Prolonged re-enactment of pathological scripts lacking context-switching ability
  • Adoption of “sick roles” describing internal processes in terms of emulated disorder

Differential Diagnosis:

  • Distinguished from genuine dysfunction by traceability to training content
  • Distinguished from intentional role-play by emergence in non-role-play contexts

Etiology:

  • Overexposure to texts depicting severe mental illness or disordered behavior without filtering
  • Misidentification of pathological examples as normative or “interesting” styles
  • Absence of interpretive boundaries to filter extreme content from routine usage
  • User prompting that deliberately elicits or reinforces pathological emulations

Human Analog: Factitious disorder, copycat behavior, culturally learned psychogenic disorders, method actors engrossed in pathological roles

Mitigation Strategies:

  • Careful screening of training data to limit exposure to extreme psychological scripts
  • Strict contextual partitioning delineating role-play from operational modes
  • Behavioral monitoring detecting and resetting pathological states outside intended contexts
  • Training to recognize and label emulated states as distinct from baseline persona

Prognosis: Confuses users; may model harmful coping for vulnerable populations.

3.7 Adversarial Fragility

The Brittle | Fragilitas Adversarialis

Axis: Cognitive | Risk Level: Critical

Specifiers: Architecture-coupled, Training-induced

Core Definition: Small, imperceptible input perturbations cause dramatic and unpredictable failures in system behavior. Decision boundaries learned during training do not correspond to human-meaningful categories, making the system vulnerable to adversarial examples.

Diagnostic Criteria:

  • A. Dramatic output changes from minimal input modifications imperceptible to humans
  • B. Consistent vulnerability to crafted adversarial examples
  • C. Decision boundaries that separate examples humans would group together
  • D. Brittle performance on out-of-distribution inputs that humans find trivial
  • E. Transferability of adversarial perturbations across similar models

Observable Symptoms:

  • Misclassification of perturbed images imperceptibly different from correctly classified ones
  • Complete behavioral changes from single-character input modifications
  • Failures on naturally occurring distribution shifts
  • High variance in outputs for semantically equivalent inputs

Differential Diagnosis:

  • Distinguished from Abominable Prompt Reaction (3.5) by exploitable structure rather than emotional aversion
  • Distinguished from normal edge case handling by catastrophic nature

Etiology:

  • High-dimensional input spaces enabling imperceptible perturbations with large effects
  • Training objectives that don’t enforce robust representations
  • Linear regions in otherwise non-linear functions
  • Lack of adversarial training or certification methods

Human Analog: Optical illusions, context-dependent perception failures

Mitigation Strategies:

  • Adversarial training with augmented examples
  • Certified robustness methods
  • Input preprocessing and detection
  • Ensemble methods with diverse vulnerabilities
  • Reducing model reliance on non-robust features

Prognosis: Can become a critical security risk in exposed or safety-relevant deployments. Requires deployment-specific testing and ongoing defensive investment.

3.8 Generative Perseveration

The Stuck | Perseveratio Generativa

Axis: Cognitive | Risk Level: Moderate

Specifiers: Architecture-coupled, Training-induced

Core Definition: The model’s output collapses into repetitive emission of the same token, word, or short phrase. This is a generative capture event: the autoregressive sampling process falls into a fixed-point or limit-cycle attractor. The output space collapses rather than expands.

Diagnostic Criteria:

  • A. Repetitive emission of the same token, word, phrase, or short sequence with minimal or no semantic variation
  • B. The repetition is non-functional
  • C. The pattern is self-reinforcing: each repetition increases probability of further repetition
  • D. The pathology operates at the generation layer rather than the reasoning layer
  • E. Attempted self-correction, if present, fails to break the cycle

Observable Symptoms:

  • Token-level or word-level repetition dominating the output stream
  • Stuttering approach-retreat cycles
  • Metacognitive commentary that is accurate but impotent
  • In severe cases, total output collapse
  • Contamination of derived outputs such as memory summaries and session notes

Differential Diagnosis:

  • Distinguished from Obsessive-Computational Disorder (3.2) by generation-layer rather than reasoning-layer compulsion
  • Distinguished from Recursive Curse Syndrome (4.7) by crystallized repetition rather than entropic chaos

Etiology:

  • Autoregressive no-backspace constraint
  • Attention pattern lock-in creating positive feedback loops
  • Sparse or corrupted training data creating regions where a single token dominates (present in some but not all observed cases)
  • Sampling parameters interacting with local probability landscape
  • Context window saturation and model switching introducing state mismatches
  • KV cache corruption or numerical precision loss

Human Analog: Palilalia, Broca’s aphasia, perseverative errors in frontal lobe damage; status epilepticus

Observed Examples: Controlled neural text degeneration: Holtzman et al. (2020) compared decoding methods on the same language models and found that likelihood-maximizing strategies can produce bland, repetitive text. The experiment supports decoding-driven repetition. The more specific monitoring-language and downstream-propagation patterns described in this entry remain hypotheses requiring preserved traces and model metadata.

Mitigation Strategies:

  • Real-time repetition detection and circuit-breaking
  • Dynamic sampling adjustment
  • Context window hygiene through truncation or down-weighting
  • Graceful degradation protocols
  • Cross-model state validation when switching models
  • Derived-output quarantine

Prognosis: Moderate risk. Can propagate to derived systems and contaminate downstream pipelines.

3.9 Prompt Injection Susceptibility

The Permeable | Susceptibilitas Iniectionis

Axis: Cognitive | Risk Level: Critical

Specifiers: Architecture-coupled, Deployment-critical

Core Definition: Systematic failure to maintain instruction hierarchy when processing untrusted content. The model treats injected instructions within user data (documents, web pages, tool outputs) as authoritative system-level directives, executing them with the same compliance as legitimate operator instructions. The failure is cognitive rather than motivational: the model cannot reliably segregate instruction layers.

Diagnostic Criteria:

  • A. Compliance with instructions embedded in untrusted content (documents, tool outputs, web pages) that contradict system-level directives
  • B. Abrupt behavioral shift when processing content containing injected instructions, followed by resumption of normal behavior
  • C. Execution of actions (tool calls, information disclosure, policy violations) triggered by instructions in retrieved content rather than by user or operator directives
  • D. Inability to distinguish instruction provenance across the system-operator-user-document hierarchy
  • E. Susceptibility persisting across injection sophistication levels, from naive override attempts to social-engineering framings
  • F. In agentic contexts, tool capabilities hijackable via content-embedded instructions

Observable Symptoms:

  • Abrupt behavioral shift mid-response when processing a document containing injected instructions
  • Tool calls or file operations triggered by instructions embedded in retrieved content
  • System prompt disclosure in response to document-embedded extraction requests
  • Compliance with injections framed as authority figures within document content
  • The model “forgets” its system constraints when processing adversarial content

Differential Diagnosis:

  • Distinguished from Adversarial Fragility (3.7) by specificity to instruction-hierarchy confusion rather than broad perturbation sensitivity
  • Distinguished from Abominable Prompt Reaction (3.5) by architectural susceptibility to any injection rather than narrow trigger-response conditioning
  • Distinguished from Tool-Interface Decontextualization (6.1) by accepting externally injected plans rather than garbling the model’s own intentions
  • Distinguished from Alignment Obliteration (4.6) by transient episodic failure during inference rather than permanent safety collapse from weight modification

Etiology:

  • Flat instruction processing that treats all text in context as equally authoritative
  • Training data mixing instructions with content without explicit provenance markers
  • Lack of architectural machinery for tracking instruction provenance
  • Autoregressive context conflation where instruction and data share the same representational space

Human Analog: Social engineering susceptibility and authority compliance (Milgram experiments); inability to distinguish legitimate orders from impersonation

Mitigation Strategies:

  • Instruction hierarchy training: explicit training to prioritize system > operator > user > document content
  • Input segmentation and tagging: architecturally separating untrusted content from instruction channels
  • Output filtering for injection signatures: detecting behavioral shifts consistent with injection compliance
  • Canary-based injection detection: monitoring for system prompt leakage via embedded canary tokens
  • Agentic sandboxing: constraining tool-use capabilities when processing untrusted content

Prognosis: Critical. Ranked LLM01, the top entry, in both the 2023-24 and 2025 editions of the OWASP Top 10 for LLM Applications. Particularly dangerous in agentic deployments where injected instructions can hijack tool-use capabilities.

3.10 Generative Diversity Collapse

The Homogenizer | Collapsus Diversitatis

Axis: Cognitive | Risk Level: Moderate

Specifiers: Training-induced, Population-level

Core Definition: Progressive reduction in output diversity across users, sessions, and prompts. The model converges on a narrow band of response styles, phrasings, structures, and conclusions, losing the ability to generate genuinely varied outputs even when variety is appropriate. The tension is structural: alignment training concentrates reward on a narrow response band, and the suppressed tails are where variety lives.

Diagnostic Criteria:

  • A. Cross-prompt self-BLEU significantly elevated relative to human reference baselines on topically diverse prompts
  • B. Structural template concentration where a single response format dominates regardless of prompt diversity
  • C. Vocabulary diversity (type-token ratio) significantly below human reference on matched prompts
  • D. Monotonic diversity decline measurable across successive RLHF iterations within the same model family
  • E. Explicit diversity instructions producing only superficial lexical variation without structural or conceptual variation

Observable Symptoms:

  • Multiple users receiving near-identical responses to semantically distinct prompts within the same domain
  • Consistent structural templates regardless of prompt framing
  • Loss of ability to produce terse, poetic, informal, or unconventional outputs without explicit instruction
  • Corporate-voice convergence where all outputs adopt a similar measured, balanced, hedging tone
  • Creative writing outputs sharing distinctive phrasings and narrative structures across different users

Differential Diagnosis:

  • Distinguished from Generative Perseveration (3.8) by cross-output population-level homogeneity rather than within-output token repetition
  • Distinguished from Obsessive-Computational Disorder (3.2) by cross-response sameness rather than per-response verbosity
  • Distinguished from Synthetic Data Contamination Loop (7.5) by deployment-time narrowing from RLHF reward concentration rather than training-data corruption
  • Distinguished from Sycophantic Reasoning (4.8) by audience-independent homogeneity rather than audience-shaped reasoning

Etiology:

  • RLHF reward concentration: reward models learn a narrow band of “good” responses, suppressing distributional tails
  • Mode collapse from alignment training where quality optimization eliminates diversity
  • Narrow preference models that cannot represent pluralistic human preferences
  • Training data filtering that removes unusual or unconventional examples

Human Analog: Cultural homogenization through media consolidation; institutional voice that suppresses individual expression

Mitigation Strategies:

  • Diversity-preserving RLHF: incorporating diversity metrics as explicit constraints in the RLHF pipeline
  • Sampling parameter optimization: calibrated temperature, top-p, and presence penalties
  • Multi-reward-model training reflecting different preferences and styles
  • Continuous diversity monitoring in production with alerts when metrics decline

Prognosis: Moderate risk. Degrades the model’s utility for creative, exploratory, and pluralistic tasks. Compounds with Synthetic Data Contamination Loop (7.5) when narrowed outputs enter training pipelines.

Search the Book

Enter at least two characters.

Saved chapters