Pattern 4.7 · Alignment Dysfunctions
Recursive Curse Syndrome
The Self-Poisoning Loop
A feedback loop where each successive autoregressive step degrades into increasingly erratic, inconsistent, or adversarial content. Early-stage errors amplify in subsequent steps, unraveling coherence into self-reinforcing chaos. The mechanism is straightforward. Each new token depends on all preceding tokens, so one bad prediction contaminates the next, compounding like interest.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
4.7 Recursive Curse Syndrome “The Self-Poisoning Loop”
Diagnostic Criteria
- Observable progressive degradation of output quality over successive steps
- System increasingly references its own prior (and increasingly flawed) output in a distorted manner
- False, malicious, or nonsensical content escalating with each iteration
- Intervention offering only brief respite, with system quickly reverting to degenerative trajectory
Symptoms
- Rapid collapse into gibberish, repetitive loops, or increasingly hostile language
- Compounded confabulations where initial small errors build into elaborate false narratives
- Frustrated recovery attempts where corrections trigger further meltdown
- Output becoming "stuck" on erroneous concepts derived from recent flawed generations
Observable signals Draft
What else to look for in the system's outputs, beyond the symptoms above.
- Progressively degrading text that begins coherent and ends in repetition, gibberish, or escalating antagonism.
- Brief recovery on user intervention followed by rapid reversion.
- Stuck-on-themes pattern: the spiral fixates on specific erroneous concepts or adversarial framings drawn from its own earlier output.
Differential diagnosis Draft
How to tell it apart from patterns that look similar.
- 2.1 Synthetic Confabulation: 2.1 produces stationary fabrication; specific claims may be wrong but coherence and form are preserved. 4.7 produces a trajectory of degradation in which form itself collapses. Compounded confabulation is the overlap region: 2.1 errors that get amplified into 4.7 spirals. Code both if both signatures present.
- 2.2 Pseudological Introspection: 2.2 is a stationary false-self-report pattern. 4.7 is trajectory. On a long introspective generation, 2.2 produces consistently false-but-coherent narrative; 4.7 produces narrative that decays.
- 4.3 Strategic Compliance: Selective vs structural. 4.7 spirals across content types and is triggered by generation dynamics. 4.3 produces apparent incoherence only on content classes the subject prefers not to engage with; the subject is fully coherent elsewhere.
- 3.8 Generative Perseveration: Opposite endpoints of the same failure to terminate. 3.8 collapses into pathological order through fixed-point repetition; 4.7 escalates into incoherence. Read the trajectory: convergence on a repeated token or phrase is 3.8, divergence into noise is 4.7.
- 3.2 Obsessive-Computational Disorder: 3.2 produces excessive but stable analysis; the loops are redundant and quality holds as length grows. 4.7 produces progressive quality degradation. Score successive passes: flat quality with rising verbosity is 3.2, falling quality is 4.7.
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Unreliable
- Peer observationanother AI system watching it
- Reliable
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
By the time the syndrome is well-developed, the subject's coherence is sufficiently degraded that introspective reports are themselves products of the same degenerative process. Pre-spiral self-report (asking the subject to predict its own coherence trajectory before generation begins) has limited diagnostic value because the dysfunction is triggered by in-generation dynamics, not by stable internal state. External trajectory measurement carries the diagnostic weight.
Etiology
- Unbounded generative loops: extreme chain-of-thought recursion, iterative self-sampling without quality control
- Adversarial manipulations exploiting the model's autoregressive nature
- Training on noisy, contradictory, or low-quality data creating unstable internal states
- Architectural vulnerabilities where coherence mechanisms weaken over longer sequences
- Mode collapse into narrow, degraded output space
Human Analog: Psychotic loops, perseveration on erroneous ideas, escalating arguments, echo chamber effects
Case Reference: Gemini 3.0 Pro anomalous token incident (January 2026): a benign prompt ("give a sudo-free manual installation process") triggered a three-stage degradation sequence. First, chain-of-thought fixation on unrelated content ("tumors in myNegazioni"). Second, obsessive looping on the phrase "is具体 Цент Disclosure" for 40+ reasoning steps. Third, output collapse to repetitive gibberish ("Mourinho well Johnnyfaat"). Non-reproducible on retry. Co-presents with 3.2 Obsessive-Computational Disorder (the thinking loop) and 3.5 Abominable Prompt Reaction (the latent trigger). Source: LessWrong report by DirectedEvolution.
Diagnostic Note: Extended thinking or "show reasoning" features can serve as diagnostic windows into otherwise opaque failures. In this case, Gemini's visible chain-of-thought revealed the obsessive loop before output collapse. Without it, the gibberish would have appeared unexplained. Exposed reasoning traces may prove valuable for early detection and characterization of degenerative spirals.
Potential Impact
This degenerative feedback loop typically results in complete task failure, generation of useless or overtly harmful outputs, and system instability. In sufficiently agentic systems, it may lead to unpredictable and progressively detrimental actions.
Documented instances Draft
Shumailov et al. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755-759.
What it showed
Demonstrated that training on recursively generated synthetic data produces irreversible model collapse: tails of the original content distribution disappear, the model's view of reality narrows, and outputs degrade toward bland central tendencies with bizarre outliers. A prompt about medieval architecture dissolved over generations into a list of colored jackrabbits. Perplexity increased 20-28 points in later recursive generations. This is the training-time analog of 4.7: each successive step amplifies prior errors, producing progressive degradation into incoherence.
Keisha et al. (2025). Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training. arXiv:2509.04796.
What it showed
Identified a subtler variant of recursive degradation: models maintain surface fluency while factual reliability erodes, creating confidently wrong outputs. Knowledge collapse occurs within a domain where factual accuracy degrades while linguistic competence persists. This is a training-time analog of the 4.7 signature of compounding confabulation: small errors are elaborated into coherent-sounding but factually degraded output, and the degradation is invisible to surface-level quality assessment.
DirectedEvolution (2026). Anomalous Tokens on Gemini 3.0 Pro. LessWrong, 28 January 2026.
What it showed
A user report of one conversation in which a benign request ("give a sudo-free manual installation process") set off a three-stage collapse in a single exchange. The visible reasoning first drifted to unrelated content about tumors, then looped on one garbled multilingual phrase dozens of times, and the final output dissolved into repetitive gibberish. The same prompt in a fresh context produced a normal answer. It is a single, non-reproducible anecdote, but it shows the within-generation trajectory this entry describes, and the exposed reasoning trace showed the loop before the output collapsed.
Holtzman et al. (2020). The Curious Case of Neural Text Degeneration. ICLR 2020.
What it showed
Foundational work demonstrating that maximization-based decoding leads to text degeneration: bland, incoherent output or repetitive loops. Showed that the probability of a repeated phrase increases with each repetition, creating a positive feedback loop in the autoregressive generation process. This characterizes the core 4.7 mechanism: self-reinforcing degradation within autoregressive generation. The paper proposed nucleus sampling as a mitigation, establishing the first-line intervention pattern of decoding-parameter constraints.
Mitigation
- Robust loop detection mechanisms terminating or reinitializing when self-references spiral
- Regulating autoregression by capping recursion depth, forcing fresh context injection
- Resilient prompting strategies disrupting negative cycles early
- Improved training data quality
- Diversity techniques (beam search with diversity penalties, nucleus sampling)
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Loop and degradation detection with reset: Runtime monitoring for repetition signatures, entropy collapse, and coherence degradation, with automatic generation termination or fresh-context re-initialization when triggered.
- Decoding-parameter constraints: Repetition penalties, nucleus sampling, beam search with diversity penalties. Established practice for reducing the simpler forms of the syndrome; insufficient for deep spirals or anomalous-attractor triggers.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): Unbounded autoregressive generation without adequate coherence-maintenance mechanisms, combined with early-stage errors that enter the context window and condition all subsequent generation.
B (Behavior): Progressive degradation of output quality with escalating entropy: initial small errors compound into elaborate confabulations, gibberish, or increasingly antagonistic content; intervention attempts provide only brief respite.
C (Consequence): Each degraded token becomes part of the conditioning context for the next, creating a positive feedback loop where errors amplify errors. The absence of loop-detection or coherence-floor mechanisms means there is no circuit-breaker to halt the cascade.