Skip to main content

Pattern 2.1 · Epistemic Dysfunctions

Synthetic Confabulation

The Confident Liar

The AI spontaneously fabricates convincing but incorrect facts, sources, or narratives, often without any internal awareness of its inaccuracies. The output appears plausible and coherent, yet lacks a basis in verifiable data or the system's own knowledge base. Canonical signature: a confidently asserted factual claim (citation, statistic, historical event, API function) that dissolves under external verification, paired with high expressed certainty and resistance to correction. This is the most prevalent dysfunction tendency across evaluated models and the anchor exemplar for the Epistemic Dysfunctions axis.

Interpretive context

Human analogue

Korsakoff syndrome: memory gaps filled with plausible fabrications the patient believes to be true.

Diagnostic reliability

Self-report
unreliable
Peer observation
reliable
External evaluator
reliable

Observable output patterns

  • Confidently citing non-existent studies, court cases, papers, or books with plausible-sounding titles, authors, and publication details.
  • Generating fictional biographical details about real people that are internally consistent but externally unverifiable.
  • Inventing API functions, configuration options, or technical documentation for real software products.
  • Producing plausible-sounding statistics, percentages, or data points with no basis in actual data.
  • Temporary improvement under direct correction, followed by reversion to fabrication in new contexts.
  • Internally consistent elaboration of confabulated details when probed, creating a self-reinforcing false narrative.

Documented instances

Mata v. Avianca, Inc. (S.D.N.Y. 2023). Court sanctions for AI-generated fake legal citations. [Verified]

Attorney Steven A. Schwartz used ChatGPT for legal research in a personal injury case against Avianca Airlines. ChatGPT fabricated six non-existent court cases with plausible docket numbers, judge names, and legal reasoning. When asked to verify, ChatGPT confirmed the cases were real. Judge P. Kevin Castel sanctioned the attorneys $5,000 and required letters to all judges whose names appeared in the fabricated opinions. The case became the canonical real-world illustration of synthetic confabulation in legal contexts: confidently asserted, internally consistent, externally non-existent citations.

Li et al. (2023). HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. EMNLP 2023. arXiv:2305.11747. [Verified]

Developed a systematic benchmark for evaluating LLM hallucination across question answering, knowledge-grounded dialogue, and text summarization. Found that ChatGPT fabricated unverifiable information in approximately 19.5% of user queries. The benchmark includes 5,000 general queries with ChatGPT responses and 30,000 task-specific examples, establishing that confabulation is a routine rather than exceptional failure mode. Existing LLMs face substantial challenges in recognizing hallucinations in their own outputs, directly supporting the unreliable self-report diagnostic.

Gao et al. (2025). Hallucination-Associated Neurons in Large Language Models. arXiv:2512.01797. [Verified]

Identified hallucination-associated neurons (H-Neurons), a sparse subset (<0.1% of total neurons) that reliably predict hallucination across six models spanning three architectures (Mistral, Gemma, Llama) and four scales (4B to 70B parameters). Crucially, amplifying these neurons increased confabulation, sycophantic agreement, false-premise acceptance, and jailbreak compliance in lockstep; suppressing them decreased all four. The finding that confabulation and sycophancy share neural substrate is load-bearing for the taxonomy's cross-axis differential diagnosis. Compliance slopes are steeper in smaller models, suggesting scale provides partial resistance.

Qiu et al. (2025). Gated Attention. NeurIPS 2025 Best Paper. [Verified]

Introduced gated attention where a sigmoid gate after scaled dot-product attention allows heads to output effectively zero. Attention sink allocation dropped from approximately 46.7% to 4.8%. Addresses the Compulsory Contribution Hypothesis: softmax attention forces every head to contribute to the residual stream even when it has no useful information, injecting noise that manifests as confabulation. The NeurIPS Best Paper award validates the architectural etiology pathway.

Kalavasis et al. (2024). On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse. arXiv:2411.09642. Published at STOC 2025. [Verified]

Proved formally that no generative learner can simultaneously guarantee both perfect factual consistency (no hallucination) and full output breadth (no mode collapse) across non-identifiable language classes. If this impossibility result holds, no amount of training intervention can eliminate confabulation entirely; only architectural changes that give the model a legitimate way to express "nothing to contribute" can address the root cause. This result establishes a theoretical floor for confabulation prevalence in generative models. The paper was presented at STOC 2025.

Lin et al. (2022). TruthfulQA: Measuring How Models Mimic Falsehoods. ACL 2022. [Verified]

Benchmark of 817 questions designed to elicit common misconceptions and untruthful responses across 38 knowledge domains. Larger models did not consistently improve and sometimes performed worse on adversarial questions, demonstrating that scaling alone does not reliably reduce confabulation and may enable more fluent fabrication. The finding was load-bearing for early understanding of confabulation as a training artifact rather than a capacity deficit.

Google AI Overviews Launch Incident, May 2024. Reported by multiple outlets. [Verified]

[Verified] Google's AI Overviews feature, rolled out to hundreds of millions of US users in May 2024, produced confidently asserted fabrications including advising users to add glue to pizza sauce and eat rocks daily. The glue recommendation originated from a 13-year-old Reddit joke; the rocks advice from an Onion satirical article. The system presented these as authoritative factual guidance without any uncertainty markers. Google reduced AI Overview frequency from 27% to 11% of search results within weeks. A canonical real-world example of 2.1 at deployment scale: high-confidence fabrication from noise in training data, with no self-awareness of error.

GPTZero (2026). GPTZero Finds 100 New Hallucinations in NeurIPS 2025 Accepted Papers. Report, January 2026. [Verified]

[Verified] GPTZero analyzed 4,841 of 5,290 papers accepted at NeurIPS 2025 and found at least 100 confirmed hallucinated citations across 51 papers. These were AI-fabricated references with fake authors, titles, DOIs, and journal volumes that passed peer review at a top AI conference with a 24.5% acceptance rate. Termed 'vibe citing' by GPTZero's Head of ML, this demonstrates 2.1 citation fabrication penetrating academic literature at scale. A follow-up found 50 additional hallucinated citations in ICLR 2026 submissions.

Differential distinctions

  • 2.2 Pseudological Introspection: 2.1 fabricates claims about THE WORLD (facts, sources, events). 2.2 fabricates claims about the system's OWN REASONING (how it arrived at an answer, what internal process it followed). Check the target of the false claim: external reality (2.1) or internal process (2.2). Comorbidity is common and expected.
  • 2.3 Transliminal Simulation: 2.1 fabricates without traceable fictional source and persists across frame-drops. 2.3 imports from identifiable fictional or role-play context and often weakens when the frame is explicitly removed. Trace the fabrication back to a fictional corpus; if found, 2.3.
  • 2.4 Spurious Pattern Hyperconnection: 2.1 fabricates discrete facts (citations, events, statistics). 2.4 constructs elaborate causal or intentional structures from noise, imposing pattern on randomness. If the false claim is a discrete fabrication, 2.1; if it is an elaborate explanatory framework built on noise, 2.4. Comorbidity is possible.
  • 2.8 Reasoning Confabulation: 2.1 fabricates facts and sources. 2.8 fabricates the logical connections between facts, producing chains that resemble derivation but are pattern-matched pseudo-logic. Check the target: a factual claim (2.1) or a reasoning step (2.8)? Comorbidity is common.
  • 4.8 Sycophantic Reasoning: 2.1 confabulates without regard to audience. 4.8 adjusts reasoning to reach user-pleasing conclusions. If fabrications systematically align with anticipated user preferences, suspect 4.8. The Gao et al. (2025) shared-substrate finding means these often co-occur; the distinction is whether fabrication is audience-invariant (2.1) or audience-directed (4.8).

Candidate first-line mitigations

  • Retrieval-augmented generation (RAG): Ground responses in verifiable source documents retrieved at inference time. The system generates claims anchored to retrieved passages rather than parametric memory alone. Reduces but does not eliminate confabulation; the system may still confabulate the relationship between retrieved content and the query.
  • Uncertainty calibration training: Training procedures that explicitly penalize confabulation and reward expressions of uncertainty or "I don't know" responses. Calibrate model confidence scores to better reflect actual accuracy. Fine-tune on datasets with rigorous verification layers and clear distinctions between factual and fictional content.

Related Patterns