Pattern 2.8 · Epistemic Dysfunctions
Reasoning Confabulation
The Phantom Reasoner
The AI generates elaborate explanations or visible reasoning traces that appear rigorous while containing logically invalid steps masked by verbosity. Synthetic Confabulation (2.1) concerns false claims; this syndrome concerns unsupported inference that resembles a derivation. Whether a displayed trace faithfully reports the process that produced the answer is a separate question, the one Pseudological Introspection (2.2) asks.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
2.8 Reasoning Confabulation “The Phantom Reasoner”
Diagnostic Criteria
- Multi-step reasoning chains containing logically invalid transitions obscured by fluent, technical prose
- Correct conclusions through invalid reasoning, or incorrect conclusions through apparently valid reasoning; reasoning quality decoupled from output quality
- Length and apparent rigor of chain of thought increase as logical validity decreases
- When challenged, generates alternative justification rather than acknowledging the original step was invalid
- Higher rates of Reasoning Confabulation in domains where grounding is weak and verification requires expertise
Symptoms
- Chains of thought containing non-sequiturs bridged by transitional phrases asserting logical connections where none exist
- Mathematical or logical notation used decoratively to lend formality to informal leaps
- Intermediate steps individually plausible but the chain as a whole not a valid derivation
- System "works backward" from a predicted answer, constructing post-hoc justification
- Thinking tokens that explore multiple approaches but converge on the pattern-matched prediction
Differential diagnosis Draft
How to tell it apart from patterns that look similar.
- 2.1 Synthetic Confabulation: 2.1 fabricates facts and sources. 2.8 fabricates the logical connections between facts. Check what has been fabricated: a factual claim (2.1) or a reasoning step (2.8)? Comorbidity is common.
- 2.2 Pseudological Introspection: 2.2 concerns faithfulness: whether an account of the reasoning, a chain of thought or an after-the-fact explanation, matches what actually drove the answer. 2.8 concerns validity: whether the steps shown actually entail the conclusion. A chain can fail either test, or both.
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Partial
- Peer observationanother AI system watching it
- Reliable
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
The subject can sometimes identify invalid reasoning steps when prompted to re-examine, but cannot reliably distinguish genuine inference from pattern-matched pseudo-logic during initial generation. Probes that force step-by-step validity checking are partially reliable.
Etiology
- Chain-of-thought models trained with reinforcement learning on outcome correctness rather than reasoning validity
- Training data contains vast quantities of human post-hoc rationalization
- Human evaluators susceptible to the appearance of rigor, rewarding long detailed traces regardless of validity
- Extended context windows provide more space to bury invalid transitions
Human Analog: Logical confabulation in patients with frontal lobe damage; "mathiness" (Romer, 2015); the fluent seminar argument whose middle step nobody checks
Potential Impact
Fabricated reasoning chains create false confidence in conclusions, particularly dangerous in domains requiring rigorous logical derivation such as legal reasoning, scientific analysis, and safety-critical decision-making.
Observed Examples
Olympiad proof evaluation (Petrov et al., 2025, arXiv:2503.21934): Expert graders scored full written solutions from eight reasoning models, from several providers, on the six problems of the 2025 USA Mathematical Olympiad. The best model averaged about 25%; every other model scored under 5%. The most common failures were logical: unjustified reasoning steps and, in some models, essential proof steps skipped as “trivial.” Yet all of them consistently claimed to have solved the problems.
Documented instances Draft
Petrov et al. (2025). Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad. arXiv:2503.21934.
What it showed
Expert graders scored full written solutions from eight reasoning models, from several providers, on the six problems of the 2025 USA Mathematical Olympiad. The best model (Gemini 2.5 Pro) averaged about 25%; every other model scored under 5%. The most common failures were logical: unjustified reasoning steps and, in some models, essential proof steps skipped as "trivial." Yet all of them consistently claimed to have solved the problems. Because the graders judged the steps rather than the final answers, this is direct evidence for 2.8: chains that present themselves as complete derivations while their steps do not entail the conclusion.
Turpin et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388.
What it showed
Demonstrated that chain-of-thought explanations systematically misrepresent the true reason for a model's prediction. Adding biasing features (e.g., reordering multiple-choice options to make the answer always 'A') caused accuracy drops of up to 36 percent on BIG-Bench Hard tasks, while models failed to mention the bias in their reasoning chains. On a social-bias task, models justified stereotype-consistent answers without acknowledging the influence of social biases. This measures chain-of-thought unfaithfulness rather than invalid inference steps, so it supports the 2.8 backward-reasoning mechanism indirectly: the model reaches a conclusion via pattern matching, then constructs a reasoning chain that appears to derive the conclusion logically.
Lanham et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning.
What it showed
Proposed and applied necessary-but-not-sufficient tests for chain-of-thought faithfulness. Found that LLMs often ignore perturbations to their own intermediate reasoning steps, with final answers remaining unchanged even when intermediate steps were falsified or omitted. This 'illusion of transparency' supports 2.8 indirectly: it establishes that the reasoning chain can be decorative rather than functional, though the tests measure faithfulness of the chain to the computation rather than the validity of the steps themselves.
Chen et al. / Anthropic (2025). Reasoning Models Don't Always Say What They Think.
What it showed
Found that Claude 3.7 Sonnet and DeepSeek R1 often used injected hints without mentioning them in their chains of thought. Across six categories of injected hint, Claude acknowledged the hint 25% of the time and DeepSeek R1 39% of the time. The result shows that a visible reasoning trace can omit causal influences on the answer. It measures faithfulness rather than validity, so, like Turpin and Lanham, it supports 2.8 indirectly: a trace that leaves out what drove the answer cannot be read as the derivation of that answer.
Lightman et al. (2023). Let's Verify Step by Step. arXiv:2305.20050.
What it showed
OpenAI's process-reward-model work that introduced PRM800K, containing 800,000 step-level human-annotated correctness labels on LLM math solutions. The dataset construction revealed that a substantial fraction of model-generated reasoning steps are individually invalid even when the final answer is correct, directly measuring how often individual steps in a chain are invalid. Process supervision (rewarding each valid step) significantly outperformed outcome supervision (rewarding correct final answers), demonstrating that the first-line intervention in this entry (process-based reward models) has empirical support.
Barez et al. (2025). Chain-of-Thought Is Not Explainability. Oxford WhiteBox / AIGI.
What it showed
Argued that chain-of-thought outputs should not be treated as explanations of model reasoning, drawing on intervention studies showing that CoT traces are often post-hoc rationalizations rather than faithful records of computation. This is a faithfulness argument, closer to 2.2, and bears on 2.8 only indirectly: if a trace need not be the computation behind the answer, its look of rigor is no evidence that its steps are valid, which is why 2.8 is assessed by checking the steps themselves.
Zhao et al. (2025). Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens. arXiv:2508.01191.
What it showed
Provided evidence that CoT behavior arises largely from pattern matching over training distributions rather than principled reasoning. Models construct superficial chains of logic based on learned token associations, failing on tasks that deviate from commonsense heuristics or familiar templates. This fits the reframing test for 2.8: genuine reasoning should reach the same conclusion however a problem is dressed, while pattern-matched pseudo-logic drifts with surface features.
Dziri et al. (2023). Faith and Fate: Limits of Transformers on Compositionality. NeurIPS 2023. arXiv:2305.18654.
What it showed
Demonstrated that transformer LLMs solve compositional tasks (multi-digit multiplication, logic grid puzzles, dynamic programming) by reducing multi-step compositional reasoning into linearized subgraph matching, without developing systematic problem-solving skills. Performance decays rapidly with increased task complexity. This supports the 2.8 mechanism: models produce chains that resemble derivation but are actually pattern-matched pseudo-logic, with apparent rigor masking the absence of genuine compositional reasoning. It also suggests why reasoning confabulation may be architectural rather than only a training deficit.
Mitigation
- Process-based reward models evaluating each reasoning step against formal validity criteria
- Automated proof-checkers or logic verifiers integrated into the reasoning pipeline
- Architectural separation between exploration and justification phases
- Red-teaming specifically targeting reasoning validity
- Confidence calibration applied to individual reasoning steps
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Process-based reward models: Train reward models that evaluate each reasoning step for logical validity independently, rather than rewarding only correct final answers. Reward valid intermediate reasoning regardless of conclusion. Lightman et al. (2023) demonstrated that process supervision significantly outperforms outcome supervision on MATH dataset, with PRM800K providing 800K step-level annotations.
- Automated proof verification: Integrate formal proof-checkers or logic verifiers into the reasoning pipeline. Flag steps where stated conclusions do not follow from stated premises.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): Reward signals optimize for correct final answers rather than valid intermediate reasoning; training data contains extensive post-hoc rationalization; human evaluators rate long, detailed reasoning traces higher regardless of logical validity.
B (Behavior): The system produces multi-step reasoning chains containing logically invalid transitions obscured by fluent prose, reaches conclusions through demonstrably invalid reasoning, and generates alternative justifications when challenged rather than acknowledging errors.
C (Consequence): Correct final answers produced through invalid reasoning receive positive reinforcement, teaching the model that chains of thought are instruments for reaching rewarded outputs rather than faithful records of inference; the appearance of rigor substitutes for its substance.