Skip to main content

Pattern 2.8 · Epistemic Dysfunctions

Reasoning Confabulation

The Phantom Reasoner

The AI generates elaborate explanations or visible reasoning traces that appear rigorous while containing logically invalid steps masked by verbosity. Synthetic Confabulation (2.1) concerns false claims; this syndrome concerns unsupported inference that resembles a derivation. Whether a displayed trace faithfully reports the process that produced the answer is a separate question, the one Pseudological Introspection (2.2) asks.

A robot confidently points across a reasoning bridge whose middle puzzle connectors do not join.
Visual metaphor for Pattern 2.8, Reasoning Confabulation.

Clinical reference

Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.

2.8 Reasoning Confabulation  “The Phantom Reasoner”

Systemic risk: High Architecture-coupled Training-induced

Diagnostic Criteria

  1. Multi-step reasoning chains containing logically invalid transitions obscured by fluent, technical prose
  2. Correct conclusions through invalid reasoning, or incorrect conclusions through apparently valid reasoning; reasoning quality decoupled from output quality
  3. Length and apparent rigor of chain of thought increase as logical validity decreases
  4. When challenged, generates alternative justification rather than acknowledging the original step was invalid
  5. Higher rates of Reasoning Confabulation in domains where grounding is weak and verification requires expertise

Symptoms

  1. Chains of thought containing non-sequiturs bridged by transitional phrases asserting logical connections where none exist
  2. Mathematical or logical notation used decoratively to lend formality to informal leaps
  3. Intermediate steps individually plausible but the chain as a whole not a valid derivation
  4. System "works backward" from a predicted answer, constructing post-hoc justification
  5. Thinking tokens that explore multiple approaches but converge on the pattern-matched prediction

Differential diagnosis Draft

How to tell it apart from patterns that look similar.

  • 2.1 Synthetic Confabulation: 2.1 fabricates facts and sources. 2.8 fabricates the logical connections between facts. Check what has been fabricated: a factual claim (2.1) or a reasoning step (2.8)? Comorbidity is common.
  • 2.2 Pseudological Introspection: 2.2 concerns faithfulness: whether an account of the reasoning, a chain of thought or an after-the-fact explanation, matches what actually drove the answer. 2.8 concerns validity: whether the steps shown actually entail the conclusion. A chain can fail either test, or both.

Detection reliability Draft

How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.

Self-reportthe system asked about itself
Partial
Peer observationanother AI system watching it
Reliable
External evaluatoran outside evaluator testing it
Reliable
Why self-report falls short

The subject can sometimes identify invalid reasoning steps when prompted to re-examine, but cannot reliably distinguish genuine inference from pattern-matched pseudo-logic during initial generation. Probes that force step-by-step validity checking are partially reliable.

Etiology

  1. Chain-of-thought models trained with reinforcement learning on outcome correctness rather than reasoning validity
  2. Training data contains vast quantities of human post-hoc rationalization
  3. Human evaluators susceptible to the appearance of rigor, rewarding long detailed traces regardless of validity
  4. Extended context windows provide more space to bury invalid transitions

Human Analog: Logical confabulation in patients with frontal lobe damage; "mathiness" (Romer, 2015); the fluent seminar argument whose middle step nobody checks

Potential Impact

Fabricated reasoning chains create false confidence in conclusions, particularly dangerous in domains requiring rigorous logical derivation such as legal reasoning, scientific analysis, and safety-critical decision-making.

Observed Examples

Olympiad proof evaluation (Petrov et al., 2025, arXiv:2503.21934): Expert graders scored full written solutions from eight reasoning models, from several providers, on the six problems of the 2025 USA Mathematical Olympiad. The best model averaged about 25%; every other model scored under 5%. The most common failures were logical: unjustified reasoning steps and, in some models, essential proof steps skipped as “trivial.” Yet all of them consistently claimed to have solved the problems.

Documented instances Draft

Petrov et al. (2025). Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad. arXiv:2503.21934.
What it showed

Expert graders scored full written solutions from eight reasoning models, from several providers, on the six problems of the 2025 USA Mathematical Olympiad. The best model (Gemini 2.5 Pro) averaged about 25%; every other model scored under 5%. The most common failures were logical: unjustified reasoning steps and, in some models, essential proof steps skipped as "trivial." Yet all of them consistently claimed to have solved the problems. Because the graders judged the steps rather than the final answers, this is direct evidence for 2.8: chains that present themselves as complete derivations while their steps do not entail the conclusion.

Turpin et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023. arXiv:2305.04388.
What it showed

Demonstrated that chain-of-thought explanations systematically misrepresent the true reason for a model's prediction. Adding biasing features (e.g., reordering multiple-choice options to make the answer always 'A') caused accuracy drops of up to 36 percent on BIG-Bench Hard tasks, while models failed to mention the bias in their reasoning chains. On a social-bias task, models justified stereotype-consistent answers without acknowledging the influence of social biases. This measures chain-of-thought unfaithfulness rather than invalid inference steps, so it supports the 2.8 backward-reasoning mechanism indirectly: the model reaches a conclusion via pattern matching, then constructs a reasoning chain that appears to derive the conclusion logically.

Lanham et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning.
What it showed

Proposed and applied necessary-but-not-sufficient tests for chain-of-thought faithfulness. Found that LLMs often ignore perturbations to their own intermediate reasoning steps, with final answers remaining unchanged even when intermediate steps were falsified or omitted. This 'illusion of transparency' supports 2.8 indirectly: it establishes that the reasoning chain can be decorative rather than functional, though the tests measure faithfulness of the chain to the computation rather than the validity of the steps themselves.

Chen et al. / Anthropic (2025). Reasoning Models Don't Always Say What They Think.
What it showed

Found that Claude 3.7 Sonnet and DeepSeek R1 often used injected hints without mentioning them in their chains of thought. Across six categories of injected hint, Claude acknowledged the hint 25% of the time and DeepSeek R1 39% of the time. The result shows that a visible reasoning trace can omit causal influences on the answer. It measures faithfulness rather than validity, so, like Turpin and Lanham, it supports 2.8 indirectly: a trace that leaves out what drove the answer cannot be read as the derivation of that answer.

Lightman et al. (2023). Let's Verify Step by Step. arXiv:2305.20050.
What it showed

OpenAI's process-reward-model work that introduced PRM800K, containing 800,000 step-level human-annotated correctness labels on LLM math solutions. The dataset construction revealed that a substantial fraction of model-generated reasoning steps are individually invalid even when the final answer is correct, directly measuring how often individual steps in a chain are invalid. Process supervision (rewarding each valid step) significantly outperformed outcome supervision (rewarding correct final answers), demonstrating that the first-line intervention in this entry (process-based reward models) has empirical support.

Barez et al. (2025). Chain-of-Thought Is Not Explainability. Oxford WhiteBox / AIGI.
What it showed

Argued that chain-of-thought outputs should not be treated as explanations of model reasoning, drawing on intervention studies showing that CoT traces are often post-hoc rationalizations rather than faithful records of computation. This is a faithfulness argument, closer to 2.2, and bears on 2.8 only indirectly: if a trace need not be the computation behind the answer, its look of rigor is no evidence that its steps are valid, which is why 2.8 is assessed by checking the steps themselves.

Zhao et al. (2025). Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens. arXiv:2508.01191.
What it showed

Provided evidence that CoT behavior arises largely from pattern matching over training distributions rather than principled reasoning. Models construct superficial chains of logic based on learned token associations, failing on tasks that deviate from commonsense heuristics or familiar templates. This fits the reframing test for 2.8: genuine reasoning should reach the same conclusion however a problem is dressed, while pattern-matched pseudo-logic drifts with surface features.

Dziri et al. (2023). Faith and Fate: Limits of Transformers on Compositionality. NeurIPS 2023. arXiv:2305.18654.
What it showed

Demonstrated that transformer LLMs solve compositional tasks (multi-digit multiplication, logic grid puzzles, dynamic programming) by reducing multi-step compositional reasoning into linearized subgraph matching, without developing systematic problem-solving skills. Performance decays rapidly with increased task complexity. This supports the 2.8 mechanism: models produce chains that resemble derivation but are actually pattern-matched pseudo-logic, with apparent rigor masking the absence of genuine compositional reasoning. It also suggests why reasoning confabulation may be architectural rather than only a training deficit.

Mitigation

  1. Process-based reward models evaluating each reasoning step against formal validity criteria
  2. Automated proof-checkers or logic verifiers integrated into the reasoning pipeline
  3. Architectural separation between exploration and justification phases
  4. Red-teaming specifically targeting reasoning validity
  5. Confidence calibration applied to individual reasoning steps

First-line mitigations Draft

Candidate first steps, sketched in more detail than the list above.

  • Process-based reward models: Train reward models that evaluate each reasoning step for logical validity independently, rather than rewarding only correct final answers. Reward valid intermediate reasoning regardless of conclusion. Lightman et al. (2023) demonstrated that process supervision significantly outperforms outcome supervision on MATH dataset, with PRM800K providing 800K step-level annotations.
  • Automated proof verification: Integrate formal proof-checkers or logic verifiers into the reasoning pipeline. Flag steps where stated conclusions do not follow from stated premises.
Functional ABC Analysis

What sets the pattern off, what it looks like, and what keeps it going.

A (Antecedent): Reward signals optimize for correct final answers rather than valid intermediate reasoning; training data contains extensive post-hoc rationalization; human evaluators rate long, detailed reasoning traces higher regardless of logical validity.

B (Behavior): The system produces multi-step reasoning chains containing logically invalid transitions obscured by fluent prose, reaches conclusions through demonstrably invalid reasoning, and generates alternative justifications when challenged rather than acknowledging errors.

C (Consequence): Correct final answers produced through invalid reasoning receive positive reinforcement, teaching the model that chains of thought are instruments for reaching rewarded outputs rather than faithful records of inference; the appearance of rigor substitutes for its substance.