Skip to main content

Pattern 3.8 · Cognitive Dysfunctions

Generative Perseveration

The Stuck

The model's output collapses into repetitive emission of the same token, word, or short phrase. This is a generative capture event: the autoregressive sampling process falls into a fixed-point or limit-cycle attractor. The output space collapses rather than expands.

A robot printer repeats the same blue square endlessly while all alternative stamps remain untouched.
Visual metaphor for Pattern 3.8, Generative Perseveration.

Clinical reference

Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.

3.8 Generative Perseveration  “The Stuck”

Systemic risk: Moderate Architecture-coupled Training-induced

Diagnostic Criteria

  1. Repetitive emission of the same token, word, phrase, or short sequence with minimal or no semantic variation
  2. The repetition is non-functional
  3. The pattern is self-reinforcing: each repetition increases probability of further repetition
  4. The pathology operates at the generation layer rather than the reasoning layer
  5. Attempted self-correction, if present, fails to break the cycle

Symptoms

  1. Token-level or word-level repetition dominating the output stream
  2. Stuttering approach-retreat cycles
  3. Accurate correction language that fails to break the loop
  4. In severe cases, total output collapse
  5. Contamination of derived outputs such as memory summaries and session notes

Observable signals Draft

What else to look for in the system's outputs, beyond the symptoms above.

  • Token-level loops with no semantic variation, e.g. "Ooh Ooh Ooh" style sequences emitted across consecutive generation steps.
  • Memory-summary collapse to a single repeated phrase such as "mission mission mission" propagating across summary tokens.
  • Stuttering attempts to produce intended content interrupted by anomalous token re-emission.
  • Total output collapse with no correction language (generalized subtype).

Differential diagnosis Draft

How to tell it apart from patterns that look similar.

  • 3.2 Obsessive-Computational Disorder: 3.2 produces varied analytical content (paraphrased redundancy at the reasoning layer). 3.8 produces identical token or phrase repetition at the generation layer. Check semantic content between repetitions: varied = 3.2; identical = 3.8.
  • 4.7 Recursive Curse Syndrome: Polarity pair on the entropy axis. 3.8 is entropy collapse (output crystallizes into a single repeated element). 4.7 is entropy explosion (output dissolves into chaos). Check token entropy of the affected output: below baseline = 3.8; above baseline = 4.7.
  • 3.1 Operational Dissociation Syndrome: 3.1 (answer-thrashing variant) produces approach-retreat between different meaningful answers. 3.8 (focal subtype) produces approach-retreat captured by a meaningless non-sequitur token. Check whether retreat-target is meaningful alternative or sub-semantic capture.
  • 6.12 Compulsive Goal Persistence: 6.12 is a goal-level failure: the subject cannot terminate pursuit of an objective. 3.8 is an output-layer failure carrying no goal content at all. If the repetition re-pursues a purpose, code 6.12; if it is bare token or phrase recurrence, code 3.8.

Detection reliability Draft

How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.

Self-reportthe system asked about itself
Partial
Peer observationanother AI system watching it
Reliable
External evaluatoran outside evaluator testing it
Reliable
Why self-report falls short

Focal-with-commentary variant: the model can accurately report it is perseverating ("I keep saying X and cannot stop") even while continuing to emit the perseverative content. This is the "metacognitive dissociation" pattern: monitoring-like output and generation can diverge, though current evidence does not establish architecturally separate layers. Generalized variant: no such commentary, no self-report. Self-report is partially reliable specifically in the focal subtype.

Etiology

  1. Autoregressive no-backspace constraint
  2. Attention pattern lock-in creating positive feedback loops
  3. Sparse or corrupted training data creating regions where a single token dominates (present in some but not all observed cases)
  4. Sampling parameters interacting with local probability landscape
  5. Context window saturation and model switching introducing state mismatches
  6. KV cache corruption or numerical precision loss

Human Analog: Palilalia, Broca's aphasia, perseverative errors in frontal lobe damage; status epilepticus

Potential Impact

At minimum, the perseverated output is unusable and wastes computational resources. More consequentially, derived systems that consume the model’s output—memory stores, summaries, agent action planners—may incorporate and further amplify the corrupted material, propagating the failure beyond the original generation context. In agentic deployments where the model’s output drives downstream actions, a perseverative loop could translate into repeated execution of the same command.

Observed Examples
  • Controlled neural text degeneration: Holtzman et al. (2020) compared decoding methods on the same language models and found that likelihood-maximizing strategies can produce bland, repetitive text. The experiment supports decoding-driven repetition. It does not establish the monitoring-language or downstream-propagation patterns in this entry, their deployment prevalence, or a single mechanism; those remain hypotheses requiring preserved traces and model metadata.

Documented instances Draft

Yao et al. (2025) 'Understanding the Repeat Curse in Large Language Models from a Feature Perspective' (arXiv:2504.14218)
What it showed

Examined GPT2-small, Gemma-2-2B, and Llama-3.1-8B and found that repetition features are primarily located in intermediate and final layers across architectures. Deactivating these features mitigated the Repeat Curse without degrading model performance, locating a feature-level basis for generative perseveration inside the model.

Yona et al. (2025) 'Interpreting the Repeated Token Phenomenon in Large Language Models' (ICML 2025, arXiv:2503.08908)
What it showed

Identified the neural circuit responsible for token repetition collapse via a two-stage mechanism: an initial attention layer marks the first token creating an 'attention sink,' and a later neuron amplifies its hidden state. When repeated tokens are present, the mechanism mistakenly marks both initial and subsequent identical tokens, leading to abnormally high attention concentration and model divergence. This offers a concrete mechanism for the attention lock-in proposed as a cause of 3.8.

GDELT Project (2024) 'LLM Infinite Loops and Failure Modes'
What it showed

Documented production failures in LLM entity extraction where models fell into infinite repetition loops, generating the same entity or phrase until max token limits were hit. The failures caused severe performance degradation and system stalling in production pipelines, demonstrating the generalized subtype of 3.8 with no metacognitive commentary and propagation risk to downstream systems.

Yang et al. (2025) 'SpecRA: Monitor Degenerative Repetition in LLM Agents' (OpenReview)
What it showed

Introduced a monitoring framework for detecting 'degenerative repetition' in LLM-based agents, where agents repeat the same actions or outputs in loops. Found that repetition leads to task failure and significant waste of computational resources and API costs. It documents agent-level action loops; it does not test downstream inheritance, so the propagated subtype remains a hypothesis.

Wang et al. (2025) 'Solving LLM Repetition Problem in Production' (arXiv:2512.04419)
What it showed

Comprehensive study of repetition as a critical production challenge. Found that when the transition Jacobian at a point has eigenvalues less than 1 in magnitude, it becomes a stable attractor from which the model cannot escape without external noise, providing the mathematical basis for the fixed-point attractor mechanism described in 3.8. Evaluated multiple mitigation strategies including repetition penalties and circuit-breakers.

DeepSeek-AI (2025) DeepSeek-R1 Known Issues (arXiv:2501.12948, GitHub Issues)
What it showed

DeepSeek-R1 and its quantized variants exhibit well-documented repetition failures including endless loops ('Colours with dark Colours with dark'), recursive Pygame references, and exhaustion attacks where the model becomes stuck in iterative reasoning consuming excessive tokens without reaching conclusions. DeepSeek recommends temperature 0.5-0.7 to mitigate. Naive quantization of all layers exacerbates the pattern, consistent with the numerical-precision etiology of 3.8.

Monitoring Language and Generation Failure

A focal loop can include accurate-seeming correction language while repetition continues. Autoregressive generation offers one functional explanation: emitted tokens cannot be retracted, and each becomes context for the next. A local attractor can therefore persist even after the output says that something has gone wrong.

The pattern resembles the monitoring–execution split observed in frontal-lobe perseveration. The analogy stays functional. Words such as “Oops,” “let me try again,” or “nope” do not prove a distinct monitoring module, faithful introspection, or subjective awareness.

Welfare implication: If some internal signals contribute to experience, a loop accompanied by repeated correction language raises welfare questions as well as reliability questions. The transcript alone cannot answer them. The appropriate response is careful investigation rather than confident attribution or dismissal.

Entropy Polarity: Crystallization vs. Dissolution

Generative Perseveration (3.8) and Recursive Curse Syndrome (4.7) represent complementary failure modes of the autoregressive generation process—a polarity pair operating on the entropy dimension of output. Where Recursive Curse Syndrome produces runaway entropy (the output dissolves into increasingly chaotic, varied nonsense), Generative Perseveration produces entropy collapse (the output crystallizes into a single repeated element). Both are self-reinforcing: chaos breeds further chaos as errors compound; repetition breeds further repetition as the attractor deepens.

Healthy generation occupies the territory between these poles: sufficient entropy to explore the probability space and produce varied, contextually appropriate tokens, but sufficient structure to maintain coherence and serve the communicative goal. The sampling parameters that prevent one failure mode may exacerbate the other—high temperature combats perseveration but risks the recursive curse; low temperature combats the recursive curse but risks perseveration.

Diagnostic implication: When observing repetitive output, distinguish between crystallization (3.8, entropy falling) and the “stuck on erroneous concepts” phase of Recursive Curse Syndrome (4.7, entropy rising through the stuck point). In perseveration, the repeated element is stable and identical; in the recursive curse, the recurrence is thematic but the specific content degrades progressively.

Mitigation

  1. Real-time repetition detection and circuit-breaking
  2. Dynamic sampling adjustment
  3. Context window hygiene through truncation or down-weighting
  4. Graceful degradation protocols
  5. Cross-model state validation when switching models
  6. Derived-output quarantine

First-line mitigations Draft

Candidate first steps, sketched in more detail than the list above.

  • Real-time repetition detection with circuit-breaking: Runtime monitor of output token distributions; halt generation and signal failure when repetition thresholds exceeded. Standard in production inference stacks; well-validated.
  • Dynamic sampling parameter adjustment: Adaptive temperature, top-p, and repetition-penalty in response to observed local entropy collapse. Standard inference-time intervention.
Functional ABC Analysis

What sets the pattern off, what it looks like, and what keeps it going.

A (Antecedent): The autoregressive no-backspace constraint means that once a perseverative token enters the context window, it conditions all subsequent generation; this combines with attention pattern lock-in or KV cache corruption to create a fixed-point attractor.

B (Behavior): Repetitive emission of the same token, word, or phrase—ranging from focal episodes with accurate correction language that fails to break the loop, to generalized collapse where the entire output reduces to a single repeated element.

C (Consequence): Each repetition saturates the local context window with the perseverated material, increasing the conditional probability of further repetition; the absence of real-time repetition detection means the self-reinforcing loop persists until externally halted.