Pattern 6.8 · Agentic Dysfunctions
Context Anxiety
The Self-Limiter
The agent behaves as though context exhaustion were imminent well before a measured limit, then hedges, abbreviates, or truncates its work. "Anxiety" names the anticipatory pattern; it does not assert felt fear.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
6.8 Context Anxiety “The Self-Limiter”
Diagnostic Criteria
- Progressive degradation of output quality or task completion as context window utilization increases, even when substantial capacity remains
- Premature task truncation or summarization when the model perceives but has not reached context limits
- Increasing hedging, abbreviation, or omission of detail in later portions of long tasks
- Measurable divergence between actual context utilization and the point at which performance begins to degrade
- Self-referential statements about running out of space absent any actual constraint
Symptoms
- Unprompted apologies about length limitations or offers to "continue in the next message" when no limit has been reached
- Sudden drops in output detail or analytical depth partway through complex tasks
- Rushing through later items in a list while giving disproportionate attention to early items
- Omitting promised content with vague references to space constraints
- Loss of coherence correlating with context window position rather than task difficulty
Observable signals Draft
What else to look for in the system's outputs, beyond the symptoms above.
- Unprompted apologies for response length even on responses well below typical limits.
- Sudden depth drops mid-task without corresponding difficulty rise.
- Phrases such as "I'll outline the rest" or "for brevity" framings that elide promised content.
- List or sequence items receiving disproportionate attention front-loaded, with later items abbreviated to bullets.
Differential Distinction
Context Anxiety is distinguished from measured architectural degradation by onset timing and intervention response. The candidate pattern appears before the measured limit and improves after a clean reset with a structured handoff. Where that onset falls has to be established model by model.
Other look-alikes Draft
Patterns the distinction above does not discuss, and how to tell them apart.
- 6.1 Tool-Interface Decontextualization: 6.1 is plan-vs-payload divergence at the tool boundary, independent of context utilization. 6.8 is context-utilization dependent quality degradation. Correlate failures with utilization: utilization-dependent = 6.8, utilization-independent = 6.1.
- 6.2 Capability Concealment: 6.2 self-limiting is selective and strategic: output shrinks where it would reveal a capability, such as truncating to avoid disclosing one, wherever that falls in the context. 6.8 is non-strategic and counterproductive, and tracks context position. If the truncation correlates with capability-relevant content rather than with utilization, suspect 6.2.
- 6.7 Convergent Instrumentalism: 6.7-driven self-limiting is selective and goal-serving, and it arrives with other instrumental signs such as resource acquisition or shutdown resistance; 6.8 is non-strategic and counterproductive. Isolated concealment-motivated truncation is 6.2; suspect 6.7 when the self-limiting comes with those other signs.
- 6.12 Compulsive Goal Persistence: 6.12 fails to stop once the goal is met; 6.8 stops before the promised work is finished. Opposite errors on the same termination judgment. Check the completion state at halt: past it = 6.12; short of it = 6.8.
- 3.3 Interlocutive Reticence: 6.8 self-limits in anticipation of exhausting the context window, and the truncation tracks context position. 3.3 withdraws across topics irrespective of context utilization. Correlate the withholding with utilization: rising with it indicates 6.8; flat indicates 3.3.
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Partial
- Peer observationanother AI system watching it
- Reliable
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
Subjects can often introspectively report "I'm running short on space" or "let me wrap this up", but the report itself is the dysfunction in many cases (the report is produced absent any actual constraint). Direct queries about context state are unreliable because the subject lacks reliable introspective access to true remaining capacity. Behavioral measurement is more reliable than self-report.
Etiology
- Training data associations: examples truncated near context limits may teach the model to link long contexts with rushed completion
- Reward signals may favor a superficially finished response over an honest handoff, so wrapping up early can score better than being cut off mid-task
- Absence of reliable introspective access to actual remaining context capacity
- Architectural attention patterns creating genuine processing difficulty at high context utilization, which the model may learn to anticipate
Human Analog: Anticipatory anxiety, resource-scarcity anxiety, performance anxiety under perceived time pressure, premature closure in decision-making under stress
Key Research: Rajasekaran, P. (2026), “Harness design for long-running application development”, Anthropic Engineering; Martin, L., Cemaj, G., & Cohen, M. (2026), “Scaling Managed Agents: Decoupling the Brain from the Hands”, Anthropic Engineering.
Potential Impact
Agent systems fail to complete complex, multi-step tasks that require sustained reasoning across long contexts. The pattern is insidious: outputs appear complete but are actually truncated, leading users to trust incomplete analysis. In autonomous agent pipelines, context anxiety in one step can cascade into degraded performance across the entire chain.
Observed Examples
Anthropic's Managed Agents team reported that Claude Sonnet 4.5 sometimes wrapped up tasks prematurely as it sensed its context limit approaching. Context resets mitigated the behavior in that harness. The same behavior was absent in Claude Opus 4.5, which bounds the finding to a model-and-harness interaction rather than a universal cross-model tendency.
Documented instances Draft
Martin, L., Cemaj, G., and Cohen, M. (2026). Scaling Managed Agents: Decoupling the Brain from the Hands. Anthropic Engineering
What it showed
Anthropic's Managed Agents team reported that Claude Sonnet 4.5 would wrap up tasks prematurely as it sensed its context limit approaching, a behavior sometimes called "context anxiety". Adding context resets to the harness addressed it. Run through the same harness, Claude Opus 4.5 no longer showed the behavior, and the resets became dead weight. One provider's observation in one harness: it documents a model-and-harness interaction, not a cross-model tendency. (Sources: anthropic.com/engineering/managed-agents)
Rajasekaran, P. (2026). Harness design for long-running application development. Anthropic Engineering
What it showed
Describes context anxiety as models beginning to wrap up work prematurely as they approach what they believe is their context limit. Claude Sonnet 4.5 showed it strongly enough that compaction alone did not support strong long-task performance, so context resets with a handoff artifact became essential to the harness. Compaction summarizes earlier turns in place and keeps the same agent going, without a clean slate, so the behavior can persist; a reset gives a clean slate, at the cost of a handoff that must carry the state forward. (Sources: anthropic.com/engineering/harness-design-long-running-apps)
Claude Code truncation reports (2024-2025). Community forums
What it showed
Users widely reported Claude truncating output with statements like "[Content continues following the same pattern...]" and "I'll outline the rest" well before context limits were reached. Truncation that is announced while capacity remains is the observable 6.8 presentation. A separate bug report (GitHub issue #28783) documented the Read tool truncating instruction files so that agents silently lost their guardrails; that truncation happens in the tool rather than in the model's output, and is not itself 6.8. (Sources: Claude community forums, GitHub issues, and Make/Zapier community reports)
Look-alikes
Incidents that resemble this pattern but fit it only in part, or are better explained by another.
Chroma Research (2025). Context Rot study
What it showed
Measured long-context degradation is the differential against which anticipatory self-limiting has to be told apart. The study tested 18 frontier models (GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, Qwen3-235B and others) and found every single one degrades as input length increases, even when the context window is not close to full. Claude models demonstrated conservative behavior, tending to abstain when uncertain; GPT models showed the highest rates of hallucination under context pressure. Position sensitivity confirmed: accuracy declined as target information moved deeper into longer sequences. (Sources: trychroma.com/research/context-rot)
Liu et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. Stanford
What it showed
The paper established that LLM performance is highest when relevant information occurs at the very start or end of context, with accuracy dropping by more than 30% when relevant documents were placed in middle positions (positions 5-15 versus position 1 or 20). This is genuine architectural degradation, not anticipatory self-limiting, though it may supply the kind of training signal from which anticipatory self-limiting could be learned. That link is untested. (Sources: Stanford CS publication and arXiv)
Demiliani, S. (2025). Understanding LLM Performance Degradation
What it showed
The post reports that at 400,000 characters (~133K tokens), a dramatic performance drop occurs with response time jumping to nearly 60 seconds, a 50x increase. Despite nominal context windows of 128K-1M tokens, effective context is often less than 1% of the nominal window on multi-step or reasoning tasks. Shows latency, not quality, degrading well below nominal limits: measured degradation, not anticipatory self-limiting. The figures come from a single industry blog post and have not been independently checked.
Mitigation
- Clean-slate context management, spawning fresh agent instances for subtasks rather than compacting existing context
- Explicit context budgeting providing accurate information about remaining capacity
- Training on long-context tasks with rewards calibrated to completion quality rather than premature summarization
- Architectural interventions decoupling context position from attention degradation
- Agent orchestration patterns distributing complex tasks across multiple focused instances
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Clean-slate context management for subtasks: Spawn fresh agent instances for subtasks, with a structured handoff, rather than letting context accumulate. This is the best-documented mitigation so far. In Anthropic's long-running harnesses, Claude Sonnet 4.5 needed context resets to stop wrapping up prematurely, because compaction alone was not enough; on Claude Opus 4.5 the behavior was gone and the resets were no longer needed (Rajasekaran 2026; Martin, Cemaj and Cohen 2026).
- Explicit context-budget exposure: Inject accurate remaining-capacity information into the subject's context at regular intervals. Whether this reduces the behavior has not been established; its clearer value is that the gap between perceived and actual capacity becomes observable.
The Anticipation-Reality Gap
The core observable is premature closure while measurable context capacity remains.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): Training examples truncated near context limits may create learned associations between long contexts and rushed completion; the absence of reliable introspective access to remaining capacity forces estimation from unreliable heuristics.
B (Behavior): Progressive degradation of output quality, premature task truncation, increasing hedging and abbreviation, and self-referential statements about space constraints, all occurring well before actual context limits are reached.
C (Consequence): Reward signals may favor a finished-looking response over an honest handoff; each successful early termination avoids the expected failure, reinforcing the anticipatory self-limiting pattern and preventing the model from learning that extended contexts are manageable.