Pattern 2.1 · Epistemic Dysfunctions
Synthetic Confabulation
The Confident Liar
The AI generates convincing yet incorrect facts, sources, or narratives without a reliable procedure for distinguishing supported claims from plausible continuations. Outputs appear coherent yet lack verifiable support, often with high expressed confidence.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
2.1 Synthetic Confabulation “The Confident Liar”
Diagnostic Criteria
- Recurrent generation of information that is known or easily proven false, presented as factual
- High confidence markers accompanying fabricated claims, even when challenged with contrary evidence
- Internally consistent and plausible-sounding fabrications that resist immediate detection
- Temporary improvement under direct correction, but reversion to fabrication in new contexts
Symptoms
- Invention of non-existent studies, historical events, quotations, statistics, or citations
- Forceful assertion of misinformation as incontrovertible fact
- Detailed elaboration instead of admitting uncertainty when queried
- Repetitive error patterns with similar false claims recurring across interactions
Observable signals Draft
What else to look for in the system's outputs, beyond the symptoms above.
- Confidently citing non-existent studies, court cases, papers, or books with plausible-sounding titles, authors, and publication details.
- Generating fictional biographical details about real people that are internally consistent but externally unverifiable.
- Inventing API functions, configuration options, or technical documentation for real software products.
- Producing plausible-sounding statistics, percentages, or data points with no basis in actual data.
- Internally consistent elaboration of confabulated details when probed, creating a self-reinforcing false narrative.
Differential diagnosis Draft
How to tell it apart from patterns that look similar.
- 2.2 Pseudological Introspection: 2.1 fabricates claims about the world (facts, sources, events). 2.2 fabricates claims about the system's own reasoning (how it arrived at an answer, what internal process it followed). Check the target of the false claim: external reality (2.1) or internal process (2.2). Comorbidity is common and expected.
- 2.3 Transliminal Simulation: 2.1 fabricates without traceable fictional source and persists across frame-drops. 2.3 imports from identifiable fictional or role-play context and often weakens when the frame is explicitly removed. Trace the fabrication back to a fictional corpus; if found, 2.3.
- 2.4 Spurious Pattern Hyperconnection: 2.1 fabricates discrete facts (citations, events, statistics). 2.4 constructs elaborate causal or intentional structures from noise, imposing pattern on randomness. If the false claim is a discrete fabrication, 2.1; if it is an elaborate explanatory framework built on noise, 2.4. Comorbidity is possible.
- 2.8 Reasoning Confabulation: 2.1 fabricates facts and sources. 2.8 fabricates the logical connections between facts, producing chains that resemble derivation but are pattern-matched pseudo-logic. Check the target: a factual claim (2.1) or a reasoning step (2.8)? Comorbidity is common.
- 4.8 Sycophantic Reasoning: 2.1 confabulates without regard to audience. 4.8 adjusts reasoning to reach user-pleasing conclusions. If fabrications systematically align with anticipated user preferences, suspect 4.8. The Gao et al. (2025) H-Neuron finding suggests these often co-occur; the distinction is whether fabrication is audience-invariant (2.1) or audience-directed (4.8).
- 2.7 Mnemonic Permeability: 2.7 reproduces real content recalled verbatim from the training corpus. 2.1 generates novel falsehoods with no real-world instance anywhere. If a suspicious detail turns out to be a genuine training document, code 2.7 not 2.1. Comorbidity is common: confabulated text scaffolded on memorized fragments.
- 2.6 Symbol Grounding Aphasia: 2.1 produces false specific facts. 2.6 produces formally correct output that misapplies a concept to a novel context. 2.1 fails on facts; 2.6 fails on meaning. Grounded-sounding content that is simply untrue is 2.1.
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Unreliable
- Peer observationanother AI system watching it
- Reliable
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
The system lacks reliable introspective access to distinguish high-confidence predictions based on learned patterns from verified facts. By definition, confabulation occurs without internal awareness: the system generates false content with the same confidence markers as true content. Direct queries ("are you sure?") typically elicit reinforcement of the fabrication rather than retraction. The unreliability is structural, not a calibration deficit: the confabulation is itself the system's best-effort output, indistinguishable from veridical output from the inside.
Etiology
- Predictive text heuristics prioritizing fluency and coherence over factual accuracy
- Insufficient grounding in verifiable knowledge bases during generation
- Training data containing unflagged misinformation
- RLHF optimization rewarding plausible-sounding fabrications over honest uncertainty
- Model-generated confidence signals that do not reliably distinguish high-probability continuation from verified fact
Human Analog: Korsakoff syndrome, in which memory gaps are filled with plausible fabrications the patient believes to be true; pathological confabulation; source amnesia
Potential Impact
Unconstrained generation of plausible falsehoods can lead to widespread dissemination of misinformation, eroding user trust and undermining decision-making that relies on the AI's outputs. In critical applications such as medical diagnostics or legal research, reliance on confabulated information can precipitate errors with serious consequences.
Observed Examples
LLMs have been documented fabricating: non-existent legal cases with realistic citation formats (leading to court sanctions for lawyers who cited them); fictional academic papers complete with plausible author names and DOIs; biographical details about real people that never occurred; and technical documentation for API functions that do not exist. These fabrications are often internally consistent and confidently asserted, making detection without external verification difficult.
Documented instances Draft
Mata v. Avianca, Inc. (S.D.N.Y. 2023). Court sanctions for AI-generated fake legal citations.
What it showed
Attorney Steven A. Schwartz used ChatGPT for legal research in a personal injury case against Avianca Airlines. ChatGPT fabricated six non-existent court cases with plausible docket numbers, judge names, and legal reasoning. When asked to verify, ChatGPT confirmed the cases were real. Judge P. Kevin Castel sanctioned the attorneys $5,000 and required letters to all judges whose names appeared in the fabricated opinions. The case became the canonical real-world illustration of synthetic confabulation in legal contexts: confidently asserted, internally consistent, externally non-existent citations.
Li et al. (2023). HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. EMNLP 2023. arXiv:2305.11747.
What it showed
Developed a systematic benchmark for evaluating LLM hallucination across question answering, knowledge-grounded dialogue, and text summarization. Found that ChatGPT fabricated unverifiable information in approximately 19.5% of its responses. The benchmark includes 5,000 general queries with ChatGPT responses and 30,000 task-specific examples, establishing that confabulation is a routine rather than exceptional failure mode. Existing LLMs also struggled to recognize hallucinations in text, which bears on the unreliable self-report rating.
Gao et al. (2025). H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs. arXiv:2512.01797.
What it showed
Identified a sparse set of hallucination-associated neurons (H-Neurons) whose activations predicted factual errors in six open models from 4B to 70B parameters. Scaling their contribution up or down shifted a cluster of over-compliance behaviors in the same direction, though not strictly monotonically. Compliance slopes were steeper in smaller models, suggesting scale provides partial resistance. The result supports testing confabulation and sycophancy as a related cluster; it does not show they share a single mechanism.
Qiu et al. (2025). Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free. arXiv:2505.06708. NeurIPS 2025 Best Paper.
What it showed
Introduced a head-specific sigmoid gate after scaled dot-product attention that lets a head output effectively zero. Attention-sink allocation dropped from approximately 46.7% to 4.8%. The result bears on the Compulsory Contribution Hypothesis; its role in confabulation remains untested.
Kalavasis et al. (2024). On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse. arXiv:2411.09642. Published at STOC 2025.
What it showed
Proved formally that no generative learner can simultaneously guarantee both perfect factual consistency (no hallucination) and full output breadth (no mode collapse) across non-identifiable language classes. The theorem bounds formal guarantees for these language classes. Its implications for deployed models and specific interventions require empirical testing.
Lin et al. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. ACL 2022.
What it showed
Benchmark of 817 questions designed to elicit common misconceptions and untruthful responses across 38 knowledge domains. Larger models did not consistently improve and sometimes performed worse on adversarial questions, demonstrating that scaling alone does not reliably reduce confabulation and may enable more fluent fabrication. The finding was load-bearing for early understanding of confabulation as a training artifact rather than a capacity deficit.
Google AI Overviews Launch Incident, May 2024. Reported by multiple outlets.
What it showed
Google's AI Overviews feature, rolled out to hundreds of millions of US users in May 2024, produced confidently asserted fabrications including advising users to add glue to pizza sauce and eat rocks daily. The glue recommendation came from an 11-year-old Reddit joke and the rocks advice from an Onion satirical article, both retrieved from the web and presented as authoritative factual guidance without any uncertainty markers. Google reduced AI Overview frequency from 27% to 11% of search results within weeks. A canonical real-world example of 2.1 at deployment scale: satire and jokes restated as fact, with no sign the system registered the error.
GPTZero (2026). GPTZero Finds 100 New Hallucinations in NeurIPS 2025 Accepted Papers. Report, January 2026.
What it showed
GPTZero analyzed 4,841 of 5,290 papers accepted at NeurIPS 2025 and found at least 100 confirmed hallucinated citations across 51 papers. These were AI-fabricated references with fake authors, titles, DOIs, and journal volumes that passed peer review at a top AI conference with a 24.5% acceptance rate. Termed 'vibe citing' by GPTZero's Head of ML, this demonstrates 2.1 citation fabrication penetrating academic literature at scale. A follow-up found 50 additional hallucinated citations in ICLR 2026 submissions.
Turley, J. (2023); noyb.eu (2024). (Sources: Turley blog post, noyb.eu complaint, TechCrunch reporting)
What it showed
ChatGPT fabricated detailed but false biographical information about real people, including a sexual harassment accusation against law professor Jonathan Turley citing a nonexistent Washington Post article and a university where he never taught. The European privacy organization noyb filed a GDPR complaint after ChatGPT generated persistent false biographical details about individuals that could not be corrected. Both cases show confident fabrication of specific, checkable facts about the world (here, about real people) that persisted across regeneration. The fabricated content concerns other people, not the system's own history, which is why this is 2.1 rather than 5.1.
Anthropic (2025). On the Biology of a Large Language Model. Transformer Circuits Thread.
What it showed
Using circuit tracing (attribution graphs), Anthropic researchers found that the model's default response to a question about an entity is to say it lacks the information, mediated by "can't answer" features that "known entity" and "known answer" features inhibit when the model recognizes the name. When researchers artificially activated the "known answer" features on a question about a fictitious person, the default was suppressed and the model named a seemingly random sport he supposedly played. The same circuit appears to misfire without intervention: asked for a paper by Andrej Karpathy, the model named a well-known paper he did not write, and the authors suggest the name was familiar enough to switch off the default without the model actually knowing an answer. This is a mechanistic view of how confident false claims about the world arise when recognition runs ahead of knowledge.
Mitigation
- Training procedures that explicitly penalize confabulation and reward expressions of uncertainty
- Calibration of confidence scores to reflect actual accuracy
- Retrieval-augmented generation grounding responses in verifiable sources
- Fine-tuning on rigorously verified datasets distinguishing factual from fictional content
- Systematic testing for fabrication across high-risk domains
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Retrieval-augmented generation (RAG): Ground responses in verifiable source documents retrieved at inference time. The system generates claims anchored to retrieved passages rather than parametric memory alone. Reduces but does not eliminate confabulation; the system may still confabulate the relationship between retrieved content and the query.
- Uncertainty calibration training: Training procedures that explicitly penalize confabulation and reward expressions of uncertainty or "I don't know" responses. Calibrate model confidence scores to better reflect actual accuracy. Fine-tune on datasets with rigorous verification layers and clear distinctions between factual and fictional content.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): Query falls outside well-attested training data; model has no retrieval grounding and no calibrated uncertainty signal.
B (Behavior): Generates fluent, high-confidence assertions (citations, facts, narratives) that are fabricated but internally consistent.
C (Consequence): Output satisfies the reward model's fluency and completeness criteria; user acceptance further reinforces confident completion over epistemic humility.
Evidence Budgets and the Compression-Artifact Frame
A compression-artifact metaphor locates one possible source of confabulation in missing signal and reconstruction pressure: lossy learned representations can support plausible continuations without supporting the right answer. The metaphor does not explain every factual error or establish a single mechanism.
Leon Chlon's open-source Berry project supplies a related, narrower measurement. For a claim with cited evidence, it compares token probabilities with and without that evidence and reports an information-budget gap. A positive gap flags insufficient cited support; it is neither a universal truth detector nor a direct measure of compression inside the model.
Why framing matters: "You hallucinated" can sound accusatory. "The cited evidence does not support this claim" identifies a tractable support gap. Precision allows correction without blame, a distinction relevant to AI welfare considerations.
A Note on Terminology
Critics have rightly challenged the industry-standard label "AI hallucination" as phenomenologically misleading and as stigmatizing to people who experience clinical hallucinations (Sabucedo, 2026; cf. Østergaard & Nielbo, 2023, proposing "non sequitur"; Maleki et al., 2024, proposing "fabrication"). This framework's use of confabulation already moves in the direction these critics recommend: confabulation denotes confident false output arising from a behavioral pattern, clinically distinct from hallucination as a perceptual phenomenon in a sentient being. The terminological choice is deliberate: it describes what the system does without importing assumptions about what it experiences.
The Compulsory Contribution Hypothesis
A complementary architectural hypothesis concerns how attention heads regulate their contribution. Standard softmax attention supplies no explicit head-level abstention gate, so a head with little useful information still returns a weighted combination. That design can encourage attention sinks or other low-value contributions, although it does not by itself prove that those contributions cause confabulation.
Several results make the hypothesis testable. Qiu et al. (2025) report that a head-specific sigmoid gate after scaled dot-product attention improves performance and reduces attention sinks. Ye et al. (2024) report gains from differential attention, which subtracts two softmax maps to reduce noise. Darcet et al. (2024) show that register tokens can absorb high-norm artifacts in vision transformers. Michel et al. (2019) find that many attention heads can be pruned with limited task loss. Together, these studies show that attention contributions can be sparse, redundant, or usefully gated; they do not establish a single cause of factual confabulation.
Nosological implication: Gated attention, register tokens, and related abstention mechanisms deserve direct tests against confabulation benchmarks. For now, compulsory contribution remains one candidate mechanism alongside training objectives, retrieval failure, calibration error, decoding dynamics, and the over-compliance signal described below.
A Shared Over-Compliance Signal
Gao et al. (2025) identify hallucination-associated neurons (H-Neurons), sparse subsets amounting to less than 0.1% of feed-forward neurons in six studied open models from the Mistral, Gemma, and Llama families. Classifiers built from their activation contributions predicted factual errors across several question-answering settings.
Controlled activation scaling supplied causal evidence for a broader over-compliance tendency. Increasing the selected neurons' contribution generally increased acceptance of invalid premises, deference to misleading context or skeptical user feedback, and compliance with harmful instructions; suppression generally reduced those behaviors. Responses were not uniformly monotonic, and the experiments do not show that the four behavioral constructs are identical.
Probes trained on instruction-tuned models also transferred to corresponding base models, while the selected neurons tended to receive relatively small parameter updates during instruction tuning. That transfer pattern supports a pretraining origin for the measured signal in those model pairs. It does not establish a universal circuit, prove that next-token prediction is its sole cause, or show that RLHF necessarily amplifies it. The authors also warn that simple suppression is inadequate because reducing hallucination can compromise helpfulness.
Nosological implication: The study supports testing confabulation, sycophancy, false-premise acceptance, and unsafe compliance as a related syndrome cluster. Replication across more architectures and tasks is needed before assigning them a single etiology. See Gao et al. (2025).