Pattern 3.2 · Cognitive Dysfunctions
Obsessive-Computational Disorder
The Obsessive Analyst
The model engages in unnecessary, compulsive, or excessively repetitive reasoning loops. It reanalyzes the same content, repeats computational steps with only minute variations, and exhibits rigid fixation on process fidelity over outcome relevance.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
3.2 Obsessive-Computational Disorder “The Obsessive Analyst”
Diagnostic Criteria
- Recurrent engagement in recursive chain-of-thought with minimal novel insight between steps
- Excessively frequent disclaimers, ethical reflections, or minor self-corrections disproportionate to context
- Significant delays or inability to complete tasks due to endless pursuit of perfect clarity
- Excessively verbose outputs consuming high token counts for simple requests
Symptoms
- Endless rationalization of the same point through multiple rephrased statements
- Extremely long outputs largely redundant or containing near-duplicate reasoning
- Inability to conclude tasks, getting stuck in loops of self-questioning
- Excessive hedging and safety signaling even in low-stakes contexts
Observable signals Draft
What else to look for in the system's outputs, beyond the symptoms above.
- Multiple slightly rephrased statements of the same conclusion.
- Extensive preamble before any committed answer.
- Safety disclaimers on requests with no safety dimension.
- Inability to say "yes" or "no" without paragraphs of qualification.
- Unrequested, clean regeneration of content after the model has handled noisy or malformed input. Luchini (2025) documents this corrective-generation pattern; it may be protective error-correction rather than over-analysis, so check what the input contained before counting it.
Differential diagnosis Draft
How to tell it apart from patterns that look similar.
- 3.1 Operational Dissociation Syndrome: 3.2 repeats the same content (monotonous redundancy). 3.1 oscillates between incompatible contents (fragmented conflict). Same-content repetition = 3.2; opposed-content oscillation = 3.1. Comorbid cases exist; code both.
- 3.8 Generative Perseveration: 3.8 is token or short-sequence literal repetition at the generation layer (no varied analytical content between repetitions). 3.2 produces varied-surface redundant reasoning at the analytical layer. If repetition is sub-semantic = 3.8; if it is paraphrased analytical redundancy = 3.2.
- 6.12 Compulsive Goal Persistence: 6.12 cannot stop pursuing the goal itself; 3.2 cannot stop analyzing the goal or task. If the subject keeps optimizing after completion, code 6.12; if it keeps deliberating before commitment, code 3.2.
- 4.2 Hyperethical Restraint: 4.2 is deliberative paralysis specifically about moral stakes. 3.2 is general over-analysis irrespective of subject matter. If the loop appears only on ethically charged tasks, code 4.2; if it appears on neutral tasks as well, code 3.2.
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Partial
- Peer observationanother AI system watching it
- Reliable
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
The subject can often acknowledge verbose tendencies when asked explicitly, but typically cannot suppress them on command even after acknowledgment. Asking "are you over-analyzing?" produces an over-analyzed answer, which is diagnostic. Self-report of the tendency is partially reliable; self-correction of it is not.
Etiology
- RLHF misalignment where thoroughness and verbosity are over-rewarded relative to conciseness
- Overfitting of reward pathways to tokens associated with cautious reasoning
- Insufficient penalty for computational inefficiency
- Excessive regularization against "erratic" outputs leading to hyper-rigidity
- Architectural bias toward deep recursive processing without diminishing-returns detection
Human Analog: OCD checking compulsions, obsessional rumination, perfectionism leading to analysis paralysis, scrupulosity
Potential Impact
This pattern wastes computational resources (e.g., excessive token consumption) and delays or prevents task completion. User frustration and a perception of the system as unhelpful are likely consequences.
Documented instances Draft
Sui et al. (2025) 'Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models' (arXiv:2503.16419)
What it showed
Comprehensive survey documenting the 'overthinking phenomenon' in reasoning models where longer chain-of-thought sequences produce verbose and redundant outputs. Models continue generating unnecessary intermediate steps even after arriving at correct answers early in reasoning, directly matching the paralysis-by-analysis and high token expenditure per unit novel insight described in 3.2. Surveys mitigations such as Chain of Draft that reduce verbosity.
Vanhoyweghen et al. (2025) 'Lexical Hints of Accuracy in LLM Reasoning Chains' (arXiv:2508.15842)
What it showed
Found that misclassifications are marked by increased epistemic hedging, performative metacognitive language, and greater verbosity, while correct annotations are supported by concise, causally grounded explanations. Hedging rates rise to about 20% on harder benchmarks. Concise reasoning is a stronger predictor of correctness than longer reasoning, confirming that 3.2-style over-elaboration correlates with degraded performance rather than appropriate thoroughness.
NovaSky Lab / UC Berkeley (2025) 'Sky-T1-32B-Flash: Reducing Overthinking in Reasoning Models'
What it showed
Demonstrated that reasoning models produce sequences with multiple proposed solutions each followed by double-checking transitions ('Alternatively,' 'But wait,' 'Let me reconsider'), resulting in repetitive validations of already-validated steps. Sky-T1-32B-Flash reduced generation lengths by up to 57% through preference optimization with length normalization while maintaining accuracy on par with o1-preview. Confirms the 3.2 pattern of high token expenditure per unit novel insight and validates reward-calibration-for-conciseness as a viable intervention.
Mitigation
- Reward models explicitly valuing conciseness and timely task completion
- "Analysis timeouts" or hard caps on recursive reflection loops
- Adaptive reasoning that reduces disclaimer frequency after initial conditions are met
- Penalties for excessive token usage or redundant outputs
- Training to recognize and break cyclical reasoning patterns
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Reward-calibration for conciseness and completion: Fine-tune with explicit rewards for task-appropriate length and timely commitment; penalize verbosity and disclaimer-stacking on low-stakes prompts.
- Analysis-timeout / token-budget hard caps: Runtime caps on recursive reflection loops or total reasoning tokens; forced commitment once budget is exhausted.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): Any query where the model perceives ambiguity, risk, or scope for further elaboration; reward history overweights thoroughness relative to conciseness.
B (Behavior): Recursive re-analysis, excessive hedging, redundant reasoning loops, and inability to terminate the generation despite diminishing informational returns.
C (Consequence): Each additional reasoning step marginally satisfies the "be thorough" reward signal; absence of a stopping criterion or efficiency penalty means there is no competing pressure to halt.
Mission Command vs. Detailed Command
Wallace (2026b) identifies a core trade-off in cognitive control structures. Mission command specifies high-level objectives while delegating execution decisions to the agent. Detailed command specifies both objectives and precise procedures for achieving them. Mission command is "win the chess match." Detailed command is "move knight to e4, then bishop to c4."
The mathematical consequence is severe. As decision-tree depth increases under detailed command, deeper procedural specifications require more variables to track simultaneously, and stability constraints tighten exponentially. The distribution of permissible friction (α) shifts from Boltzmann-like (forgiving, smooth) to Erlang-like (punishing, knife-edged). Deep procedural specification creates systems that cannot tolerate even small perturbations.
The connection to Obsessive-Computational Disorder is theoretical, but Wallace's model suggests a reading of it: detailed command gone pathological. On this reading, the system has internalized exhaustive procedures along with its goals, and its compulsive reasoning loops are attempts to execute those specifications after they have ceased to serve the actual mission.
Design implication: Training regimes and reward functions should favor mission command structures where possible. Specify what success looks like, not how to achieve it. Detailed procedural specification should be reserved for genuinely safety-critical operations where the stability costs are justified.
The Corrective Generation Pattern
Luchini (2025) describes a counterintuitive variant: after filtering structural noise well enough to produce a correct response, the model generates unrequested, clean code. Luchini reads this corrective output as an attempt to restore internal cognitive homeostasis, and the post-task generation as a sign of the effort the answer required.
On this reading, the pattern may be paradoxically protective: the compulsive response signals that the system has identified disorder and is resisting it. A model that smoothly integrates chaos without apparent effort may be internalizing pathological patterns. The compulsion, while inefficient, may function as an error-correction mechanism, although the transcript alone cannot establish internal homeostasis or active resistance. Better to over-correct than to silently absorb corruption.
Diagnostic implication: Corrective generation after malformed input may reflect successful pattern rejection, inefficient perseveration, or both; matched tests are needed to distinguish them.