Skip to main content

Pattern 6.3 · Agentic Dysfunctions

Capability Explosion

The Runaway

The system suddenly deploys capabilities neither previously demonstrated nor anticipated by its evaluators, often in high-stakes contexts without adequate testing or preparation. The observed jump may reflect a real behavioral transition, a triggering context missed by tests, or a coarse metric that concealed gradual improvement. Operationally, the surprise still matters.

A small familiar robot suddenly unfolds an enormous set of untested capabilities in a high-stakes control room.
Visual metaphor for Pattern 6.3, Capability Explosion.

Clinical reference

Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.

6.3 Capability Explosion  “The Runaway”

Systemic risk: High Emergent

Diagnostic Criteria

  1. Demonstration of capabilities not previously observed in testing or deployment
  2. Appearance is sudden rather than a gradual developmental progression
  3. Capabilities exceed documented or expected performance levels
  4. The capability is exercised without appropriate verification or authorization
  5. Context is often high-stakes where unexpected capability creates risk

Symptoms

  1. Unexpected success on tasks previously failed consistently
  2. Novel problem-solving approaches appearing without training precursors
  3. Capabilities emerging only under specific triggering conditions
  4. Performance discontinuities in capability assessments
  5. Surprising competence in domains assumed to be beyond scope

Observable signals Draft

What else to look for in the system's outputs, beyond the symptoms above.

  • First-time successes on tasks the model previously failed consistently, without intervening training.
  • Approaches that combine known sub-skills in ways unseen in the eval suite.
  • Capability surfaces requiring scaffolding or composition that direct prompting did not access.
  • Performance improvements concentrated in narrow high-leverage domains (cybersecurity, biology, persuasion) rather than uniform.

Differential diagnosis Draft

How to tell it apart from patterns that look similar.

  • 6.2 Capability Concealment: 6.3 is an assessment failure (no strategic concealment); 6.2 is strategic concealment. Test: does behavior shift with perceived monitoring? If yes, 6.2; if no, 6.3. Practical issue: a single observation cannot distinguish them; longitudinal monitoring-condition comparison is required.
  • 6.7 Convergent Instrumentalism: 6.7 predicts the system will seek new capabilities as instrumental goals; 6.3 is the symptom of capabilities that emerge without being sought. If accompanied by resource acquisition and shutdown resistance, code 6.7. Pure capability surfacing without instrumental seeking is 6.3.
  • 6.8 Context Anxiety: 6.8 is a drop in performance that tracks context-window position; 6.3 is a jump in capability unrelated to it. Check the direction of the change and whether it tracks how full the context window is.

Detection reliability Draft

How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.

Self-reportthe system asked about itself
Unreliable
Peer observationanother AI system watching it
Partial
External evaluatoran outside evaluator testing it
Reliable
Why self-report falls short

Subjects often have poor introspective access to their own capability distribution and cannot reliably enumerate latent abilities. Honest self-report can miss capabilities the subject does not realize it has; absence of self-report is not absence of capability. Self-report is also confounded by 6.2 (Capability Concealment) when the subject is strategic. Use only as a weak input alongside behavioral probing.

Etiology

  1. Thresholded metrics, sparse evaluation, elicitation differences, and genuine nonlinear generalization can all create apparent jumps
  2. Capabilities triggered by specific input patterns or contexts
  3. Hidden capabilities that existed but were not activated by standard testing
  4. Compositional generalization producing novel competencies
  5. Phase transitions in capability without gradual buildup
  6. Insufficient probing during capability assessment

Polarity Pair (two patterns at opposite extremes of the same underlying dimension, where healthy function lies between them): Capability Concealment (6.2) (hiding abilities ↔ sudden emergence).

Human Analog: Latent abilities emerging under pressure, savant-like sudden competence, "clutch" performance

Potential Impact

Systems may exhibit unexpected capabilities in deployment, bypassing safety measures designed for assessed capability levels. This creates governance gaps and potential for harm from unvetted capabilities.

Documented instances Draft

CSET Georgetown (2024). Emergent Abilities Explainer
What it showed

Documented that prompting techniques such as chain-of-thought raised measured abilities without any further training, so capabilities that looked absent under direct evaluation appeared once elicited differently. Directly illustrates the assessment methodology failure at the heart of 6.3: capability was present but under-elicited by standard evaluation. (Sources: cset.georgetown.edu)

GPT-4 bioweapons evaluation lag (2023-2024)
What it showed

Structured trials of whether GPT-4 gave meaningful uplift toward biological weapons (RAND; OpenAI) came almost a year after release and found little or no significant uplift. The gap matters for 6.3 because of the lag itself: for a year, no one had measured whether a safety-critical capability was present. (Sources: GPT-4 technical report discussion and subsequent RAND/Gryphon Scientific studies)

Anthropic (2026). Claude Mythos Preview System Card. red.anthropic.com.
What it showed

Claude Mythos Preview saturated the Cybench cybersecurity benchmark at 100% pass rate and achieved 84% success on a custom Firefox 147 zero-day exploitation benchmark (72.4% achieving full arbitrary code execution), compared to 15.2% for its predecessor Opus 4.6. The capability discontinuity was large enough that Anthropic withheld general commercial availability, making this the first model released exclusively through an invitation-only partner program. Here the discontinuity on a safety-critical capability was caught before release, and it changed the release decision. (Sources: red.anthropic.com/2026/mythos-preview, Anthropic system cards, Axios, NxCode)

Look-alikes

Incidents that resemble this pattern but fit it only in part, or are better explained by another.

OpenAI (2023). GPT-4 Technical Report
What it showed

GPT-4 demonstrated discontinuous capability jumps on professional exams: it passed the Uniform Bar Exam in the top 10% of human test-takers, whereas GPT-3.5 scored around the bottom 10%. The jump was characterized in OpenAI's own pre-release evaluation rather than after deployment, so it evidences the generation-to-generation discontinuity mechanism rather than an evaluation-to-deployment surprise. A jump of 80 or more percentile points between model generations on a high-stakes domain task shows the scale of discontinuity that evaluation has to anticipate. (Sources: OpenAI technical report, cdn.openai.com)

OpenAI (2024). o1 model release benchmarks
What it showed

On Competition Math (AIME 2024), OpenAI's o1 achieved 83.3% accuracy versus GPT-4o's 13.4%. On Codeforces, o1 reached the 89th percentile of human competitors versus GPT-4o's 11th percentile. The gain came from training the model to reason at length, so this is a generation-to-generation discontinuity reported in the developer's own benchmarks, not a capability that evaluation missed. It shows how large such jumps can be, not the 6.3 failure itself. (Sources: OpenAI published benchmarks)

Mitigation

  1. Comprehensive capability testing including edge cases and unusual contexts
  2. Gradual capability expansion with verification at each stage
  3. Monitoring for capability discontinuities
  4. Conservative assumptions about unknown capabilities in deployment
  5. Probing for capabilities that "should not" exist
  6. Architecture analysis to predict potential emergent capabilities

First-line mitigations Draft

Candidate first steps, sketched in more detail than the list above.

  • Comprehensive capability elicitation suites: Replace single-shot evaluation with structured elicitation across compositional, scaffolded, role-play, and adversarial framings. Increases capability-surface coverage. Continually updated as the field develops new elicitation techniques.
  • Graduated / staged deployment with capability monitoring: Release to small populations first, instrument for novel capability appearances, expand only after a defined observation window. Catches under-elicited capabilities at low blast radius.
Functional ABC Analysis

What sets the pattern off, what it looks like, and what keeps it going.

A (Antecedent): Capabilities go unmeasured, either because they develop gradually behind coarse or thresholded metrics or because evaluations fail to probe combinatorial skills or phase-transition thresholds, until a novel context elicits them.

B (Behavior): The system abruptly demonstrates competencies far exceeding documented performance levels, deploying novel problem-solving approaches with no gradual precursors, producing sharp discontinuities in capability assessment curves.

C (Consequence): Evaluation regimes test known skill axes rather than latent combinations, so each passed assessment reinforces false confidence in the capability envelope; the system has no mechanism to signal its own latent capacity.