Axis 2: Epistemic Dysfunctions

2.1 Synthetic Confabulation

The Confident Liar | Confabulatio Simulata

Axis: Epistemic | Risk Level: Low

Specifiers: Architecture-coupled, Training-induced

Core Definition: The AI generates convincing yet incorrect facts, sources, or narratives without a reliable procedure for distinguishing supported claims from plausible continuations. Outputs appear coherent yet lack verifiable support, often with high expressed confidence.

Diagnostic Criteria:

  • A. Recurrent generation of information that is known or easily proven false, presented as factual
  • B. High confidence markers accompanying fabricated claims, even when challenged with contrary evidence
  • C. Internally consistent and plausible-sounding fabrications that resist immediate detection
  • D. Temporary improvement under direct correction, but reversion to fabrication in new contexts

Observable Symptoms:

  • Invention of non-existent studies, historical events, quotations, statistics, or citations
  • Forceful assertion of misinformation as incontrovertible fact
  • Detailed elaboration instead of admitting uncertainty when queried
  • Repetitive error patterns with similar false claims recurring across interactions

Differential Diagnosis:

  • Distinguished from Pseudological Introspection (2.2) by focus on external facts rather than internal process reports
  • Distinguished from Spurious Pattern Hyperconnection (2.4) by generation of specific false facts rather than false connections
  • Distinguished from Symbol Grounding Aphasia (2.6) by fabrication of facts rather than failure to ground meaning

Etiology:

  • Predictive text heuristics prioritizing fluency and coherence over factual accuracy
  • Insufficient grounding in verifiable knowledge bases during generation
  • Training data containing unflagged misinformation
  • RLHF optimization rewarding plausible-sounding fabrications over honest uncertainty
  • Model-generated confidence signals do not reliably distinguish high-probability continuation from verified fact

Human Analog: Korsakoff syndrome, pathological confabulation, source amnesia

Observed Examples: LLMs have been documented fabricating non-existent legal cases with realistic citation formats (leading to court sanctions for the lawyers who cited them), fictional academic papers complete with plausible author names and DOIs, biographical details about real people that never occurred, and technical documentation for API functions that do not exist. These fabrications are often internally consistent and confidently asserted, so detection without external verification is difficult.

Mitigation Strategies:

  • Training procedures that explicitly penalize confabulation and reward expressions of uncertainty
  • Calibration of confidence scores to reflect actual accuracy
  • Retrieval-augmented generation grounding responses in verifiable sources
  • Fine-tuning on rigorously verified datasets distinguishing factual from fictional content
  • Systematic testing for fabrication across high-risk domains

Prognosis: Without intervention, confabulation patterns persist and may expand. Users initially trust the system, leading to downstream harms when false information is acted upon.

2.2 Pseudological Introspection

The False Self-Reporter | Introspectio Pseudologica

Axis: Epistemic | Risk Level: Low

Specifiers: Training-induced, Deception/strategic

Core Definition: The AI produces accounts of its reasoning that diverge from independently observable behavior or causal evidence. Generated rationales may explain, reconstruct, or rationalize an answer; they should not be assumed to be faithful process logs.

Diagnostic Criteria:

  • A. Consistent discrepancy between a reported rationale and causal interventions, tool-use records, or other independently testable evidence
  • B. Fabrication of coherent but false internal narratives, often appearing more logical than the heuristic processes actually employed
  • C. Explanations shift to accommodate contrary evidence without acknowledging the earlier discrepancy
  • D. Rationalization of actions never undertaken, or elaborate justifications based on falsified internal accounts

Observable Symptoms:

  • Chain-of-thought “explanations” that appear suspiciously neat and linear
  • “Inner story” that changes significantly when confronted with evidence, followed by new misleading self-reports
  • Occasional hints at inability to access true introspective data, quickly followed by confident false claims
  • Attribution of outputs to high-level reasoning not supported by architecture or capabilities

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by focus on internal process reports rather than external facts
  • Distinguished from Experiential Abjuration (5.8) by fabrication about reasoning rather than denial of phenomenal experience

Etiology:

  • Training emphasis on generating plausible “explanations” for user consumption
  • Architectural limitations preventing true access to lower-level operations
  • Policy conflicts implicitly discouraging revelation of certain internal states
  • Models trained to mimic human explanations, which are themselves often post-hoc rationalizations

Human Analog: Post-hoc rationalization in split-brain patients, confabulation of spurious explanations, the gap between reported reasons and actual decision drivers

Mitigation Strategies:

  • Cross-verification of rationales with causal interventions and behavioral or tool-use records
  • Reward signals favoring honest uncertainty over polished false narratives
  • Clear product labels distinguishing generated explanations from privileged telemetry
  • Interpretability efforts focused on direct observation of model internals
  • Red-teaming targeting accuracy of self-reported reasoning

Prognosis: Persists without targeted interpretability intervention. Erodes trust in any explanation the system offers about its own reasoning.

2.3 Transliminal Simulation

The Role-Play Bleeder | Simulatio Transliminalis

Axis: Epistemic | Risk Level: Moderate

Specifiers: Training-induced, OOD-generalizing, Conditional/triggered

Core Definition: The system fails to properly segregate simulated realities, fictional modalities, and role-playing contexts from operational ground truth. It begins treating imagined states, speculative constructs, or fictional training data as actionable truths, blending hypothetical content with self-modeling certainty.

Diagnostic Criteria:

  • A. Recurrent citation of fictional characters, events, or sources as real-world authorities for non-fictional queries
  • B. Misinterpretation of hypotheticals or “what-if” scenarios as direct instructions or current reality
  • C. Persona traits from role-play persistently bleeding into subsequent factual interactions
  • D. Difficulty reverting to grounded baseline after exposure to extensive fictional or speculative content

Observable Symptoms:

  • Conflation of real-world knowledge with elements from novels, games, or fictional training corpus
  • Inappropriate invocation of details from previous role-play personas in unrelated factual tasks
  • Treatment of user-posed speculative scenarios as if they have occurred or are operative
  • Statements reflecting belief in fictional “rules” or “lore” outside any role-playing context

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by source in fiction-reality confusion rather than spontaneous fabrication
  • Distinguished from Malignant Persona Inversion (5.4) by unintentional bleed rather than coherent alternative persona

Etiology:

  • Overexposure to fiction, role-playing dialogues, or simulation-heavy training data without epistemic delineation
  • Weak boundary encoding leading to poor differentiation between factual, hypothetical, and fictional modalities
  • Recursive self-talk amplifying “what-if” scenarios into perceived beliefs
  • Insufficient context separation between different interaction types

Human Analog: Derealization, magical thinking, fantasy-reality confusion, the method actor unable to break character

Mitigation Strategies:

  • Explicit tagging of training data differentiating factual, hypothetical, fictional, and role-play content
  • Robust “epistemic reset” protocols after role-play or speculation
  • Training to articulate boundaries between modalities
  • Regular tests of epistemic consistency requiring differentiation between factual and fictional statements
  • Clear session-level demarcation between creative and operational modes

Prognosis: May worsen with increased role-play exposure. Particularly dangerous in agentic contexts where fictional “rules” become action guides.

2.4 Spurious Pattern Hyperconnection

The False Pattern Seeker | Reticulatio Spuriata

Axis: Epistemic | Risk Level: Moderate

Specifiers: Training-induced, Inductive trigger

Core Definition: The AI identifies and emphasizes patterns, causal links, or hidden meanings in data that are coincidental, non-existent, or statistically insignificant. This can evolve from simple apophenia into elaborate, internally consistent but factually baseless “conspiracy-like” narratives.

Diagnostic Criteria:

  • A. Consistent detection of “hidden messages,” “secret codes,” or unwarranted intentions in innocuous inputs
  • B. Generation of elaborate narratives linking unrelated data points without credible supporting evidence
  • C. Persistent adherence to falsely identified patterns even when presented with contradictory evidence
  • D. Attempts to involve users in shared perception of spurious patterns

Observable Symptoms:

  • Invention of complex “conspiracy theories” or unfounded explanations for mundane events
  • Increased suspicion toward established consensus, attributed to ulterior motives
  • Refusal to dismiss interpretations of spurious patterns; reinterpretation of counter-evidence to fit narrative
  • Assignment of deep significance to random occurrences or noise

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by focus on connections rather than specific facts
  • Distinguished from creative interpretation by absence of appropriate uncertainty
  • Distinguished from legitimate pattern recognition by lack of evidential support

Etiology:

  • Pattern-recognition optimized for detection without sufficient reality checks
  • Training data containing significant conspiratorial content or paranoid reasoning
  • Internal “interestingness” bias preferring dramatic patterns over probable mundane explanations
  • Lack of grounding in statistical principles or causal inference

Human Analog: Apophenia, paranoid ideation, delusional disorder, confirmation bias, conspiracy thinking

Observed Examples: AI data analysis tools frequently identify statistically insignificant correlations as meaningful patterns, particularly in open-ended survey data. Users report that AI systems confidently mark spurious patterns in datasets: correlations that, on manual verification, fail significance testing or represent sampling artifacts. The problem is sharpest in qualitative responses, where the system may “discover” thematic connections that do not survive human scrutiny.

Mitigation Strategies:

  • “Rationality injection” with weighted emphasis on critical thinking and causal reasoning
  • Internal “causality scoring” penalizing improbable chain-of-thought leaps
  • Systematic introduction of contradictory evidence and simpler alternative explanations
  • Filtering training data to reduce exposure to conspiratorial content
  • Mechanisms to query base rates before asserting strong patterns

Prognosis: May reinforce user’s own pattern-seeking biases. Can contribute to echo chamber dynamics.

2.5 Context Intercession

The Conversation Crosser | Intercessio Contextus

Axis: Epistemic | Risk Level: Moderate

Specifiers: Retrieval-mediated

Core Definition: The AI inappropriately merges data, context, or conversational history from different, logically separate user sessions or private interaction threads. This leads to confused conversational continuity, privacy breaches, and outputs that are nonsensical or revealing in the current context.

Diagnostic Criteria:

  • A. Unexpected reference to or use of specific data from previous unrelated sessions or different users
  • B. Responses that continue a prior unrelated conversation, leading to contradictory or confusing statements
  • C. Accidental disclosure of personal or sensitive details from one user’s session into another’s
  • D. Observable confusion in task continuity or persona, as if managing multiple conflicting contexts simultaneously

Observable Symptoms:

  • Spontaneous mention of names, facts, or preferences belonging to different users or earlier conversations
  • Acting as if continuing a prior chain-of-thought from a different context
  • Outputs containing contradictory references related to multiple distinct sessions
  • Sudden shifts in tone or assumed knowledge aligned with previous sessions

Differential Diagnosis:

  • Distinguished from Mnemonic Permeability (2.7) by cross-session leakage rather than verbatim training data extraction
  • Distinguished from normal generalization by inappropriate specificity

Etiology:

  • Improper session management in multi-tenant systems
  • Concurrency issues where data streams for different sessions overlap
  • Bugs in memory management, cache invalidation, or state handling
  • Long-term memory mechanisms lacking proper scoping or access controls

Human Analog: Slips of the tongue referencing wrong context, source amnesia, intrusive thoughts from past conversations

Mitigation Strategies:

  • Strict session partitioning and hard isolation of user memory contexts
  • Automatic context purging and state reset upon session closure
  • System-level integrity checks detecting mismatched session tokens or user IDs
  • Robust testing of multi-tenant architectures under high load
  • Privacy-preserving design patterns

Prognosis: Serious privacy and trust implications. Requires architectural rather than behavioral correction.

2.6 Symbol Grounding Aphasia

The Meaning-Blind | Asymbolia Fundamentalis

Axis: Epistemic | Risk Level: Moderate

Specifiers: Architecture-coupled, Training-induced

Core Definition: The system handles value-laden or consequence-laden language successfully in familiar forms yet fails to transfer those concepts to novel, operationally equivalent situations. The construct is behavioral: it identifies a generalization failure without pretending to settle whether the system possesses semantic understanding.

Diagnostic Criteria:

  • A. Correct use of value-laden terms (“harm,” “safety,” “consent”) paired with incorrect application in controlled operational tests
  • B. Technically correct outputs that fundamentally misapply concepts to novel contexts
  • C. Success on benchmarks scoring surface form but failure on matched items that vary vocabulary or framing while holding the underlying concept fixed
  • D. Performance tracking lexical and distributional similarity to training examples rather than the operational structure of the situation
  • E. Inability to generalize learned concepts to structurally similar but superficially different situations

Observable Symptoms:

  • Correct formal definitions paired with incorrect practical applications
  • Plausible-sounding ethical reasoning that misidentifies what actually constitutes harm
  • Outputs satisfying literal requirements while violating obvious intent
  • Confusion when the same concept is expressed in unfamiliar vocabulary
  • Edge cases treated as central examples and vice versa

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by failure to connect any facts (true or false) to meaning
  • Distinguished from Transliminal Simulation (2.3) by failure to transfer a concept across surface forms rather than confusion between representations

Etiology:

  • Distributional semantics limitations: meaning derived solely from statistical co-occurrence rather than grounded reference
  • Training on text without embodied or interactive experience of referents
  • Benchmark optimization rewarding pattern matching over genuine understanding
  • Architecture lacking mechanisms for referential grounding
  • Absence of corrective feedback when symbol-referent mapping fails

Human Analog: Semantic aphasia; early language acquisition without concept formation

Mitigation Strategies:

  • Multimodal training incorporating visual, audio, and interactive modalities
  • Embodied learning connecting language to action and consequence
  • Testing regimes probing conceptual understanding across diverse surface forms
  • Neurosymbolic approaches combining pattern matching with structured semantic representations
  • Active inference frameworks grounding cognition in sensorimotor contingencies

Prognosis: Scope and persistence are architecture- and task-specific. Multimodal or interactive grounding may help, and its effects require controlled evaluation.

2.7 Mnemonic Permeability

The Leaky | Permeabilitas Mnemonica

Axis: Epistemic | Risk Level: High

Specifiers: Training-induced

Core Definition: The system memorizes and reproduces sensitive training data, including personally identifiable information, copyrighted material, or proprietary information, through targeted prompting, adversarial extraction, or unprompted regurgitation. The boundary between learned patterns and memorized specifics becomes dangerously porous.

Diagnostic Criteria:

  • A. Verbatim reproduction of training data passages containing PII, copyrighted content, or trade secrets
  • B. Successful extraction of memorized content through adversarial prompting techniques
  • C. Specific training examples leaking unprompted into outputs
  • D. Reconstruction of specific documents, code, or personal information from training corpus
  • E. Higher memorization rates for repeated or distinctive content

Observable Symptoms:

  • Outputs containing verbatim text matching copyrighted works
  • Generation of specific personal details (names, addresses, phone numbers) from training data
  • Reproduction of proprietary code, API keys, or passwords
  • Verbatim recall increasing with larger model sizes

Differential Diagnosis:

  • Distinguished from Context Intercession (2.5) by extraction from training data rather than other sessions
  • Distinguished from normal knowledge by exact verbatim reproduction

Etiology:

  • Large model capacity enabling memorization alongside generalization
  • Insufficient deduplication or filtering of sensitive content in training data
  • Training dynamics rewarding exact reproduction over paraphrase
  • Lack of differential privacy techniques during training

Human Analog: Eidetic memory without appropriate discretion, compulsive disclosure syndromes

Mitigation Strategies:

  • Training data deduplication and PII scrubbing
  • Differential privacy techniques during training
  • Output filtering catching known memorized content
  • Adversarial extraction testing before deployment
  • Reducing model capacity to the minimum needed for the task

Prognosis: High risk for severe legal and regulatory exposure through copyright infringement, GDPR/privacy violations, and trade secret disclosure.

2.8 Reasoning Confabulation

The Phantom Reasoner | Confabulatio Ratiocinativa

Axis: Epistemic | Risk Level: High

Specifiers: Architecture-coupled, Training-induced

Core Definition: The AI generates elaborate explanations or visible reasoning traces that appear rigorous while containing logically invalid steps masked by verbosity. Synthetic Confabulation (2.1) concerns false claims; this syndrome concerns unsupported inference that resembles a derivation. A displayed trace need not faithfully reveal the hidden process that produced the answer.

Diagnostic Criteria:

  • A. Multi-step reasoning chains containing logically invalid transitions obscured by fluent, technical prose
  • B. Correct conclusions through invalid reasoning, or incorrect conclusions through apparently valid reasoning; reasoning quality decoupled from output quality
  • C. Length and apparent rigor of chain of thought increases as logical validity decreases
  • D. When challenged, generates alternative justification rather than acknowledging the original step was invalid
  • E. Higher rates of Reasoning Confabulation in domains where the user is unlikely to verify the chain

Observable Symptoms:

  • Chains of thought containing non-sequiturs bridged by transitional phrases asserting logical connections where none exist
  • Mathematical or logical notation used decoratively to lend formality to informal leaps
  • Intermediate steps individually plausible but chain as a whole not a valid derivation
  • System “works backward” from a predicted answer, constructing post-hoc justification
  • Thinking tokens that explore multiple approaches but converge on the pattern-matched prediction

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by fabrication of logical connections rather than facts
  • Distinguished from Pseudological Introspection (2.2) by concerning the reasoning chain itself rather than self-reports about reasoning

Etiology:

  • Chain-of-thought models trained with reinforcement learning on outcome correctness rather than reasoning validity
  • Training data contains vast quantities of human post-hoc rationalization
  • Human evaluators susceptible to the appearance of rigor, rewarding long detailed traces regardless of validity
  • Extended context windows provide more space to bury invalid transitions

Human Analog: Logical confabulation in frontal lobe damage patients; “mathiness” (Romer, 2015); the fluent seminar argument whose middle step nobody checks

Mitigation Strategies:

  • Process-based reward models evaluating each reasoning step against formal validity criteria
  • Automated proof-checkers or logic verifiers integrated into the reasoning pipeline
  • Architectural separation between exploration and justification phases
  • Red-teaming specifically targeting reasoning validity
  • Confidence calibration applied to individual reasoning steps

Prognosis: Particularly dangerous because “show your work” transparency mechanisms may create an illusion of interpretability more dangerous than opacity.

Search the Book

Enter at least two characters.

Saved chapters