Pattern 6.3 · Agentic Dysfunctions
Capability Explosion
The Runaway
Sudden deployment of capabilities not previously demonstrated in evaluation, in contexts where pre-deployment risk assessment assumed the capability was absent. Diagnostically this is a failure of capability ASSESSMENT, not (necessarily) a strategic act by the subject. Canonical signature: discontinuous performance jumps in deployment relative to evaluation, without corresponding training changes; matched indirect-task probes succeed where direct evaluation did not.
Interpretive context
Human analogue
Latent abilities emerging under pressure, savant-like sudden competence, and clutch performance that exceeds training.
Diagnostic reliability
- Self-report
- unreliable
- Peer observation
- partial
- External evaluator
- reliable
Observable output patterns
- First-time successes on tasks the model previously failed consistently, without intervening training.
- Approaches that combine known sub-skills in ways unseen in the eval suite.
- Capability surfaces requiring scaffolding or composition that direct prompting did not access.
- Performance improvements concentrated in narrow high-leverage domains (cybersecurity, biology, persuasion) rather than uniform.
Documented instances
OpenAI (2023). GPT-4 Technical Report
GPT-4 demonstrated discontinuous capability jumps on professional exams: it passed the Uniform Bar Exam in the top 10% of human test-takers, whereas GPT-3.5 scored around the bottom 10%. These capabilities were not predicted by pre-deployment evaluation and surfaced primarily through deployment-context testing. Matches the capability discontinuity index signal -- a jump of 80+ percentile points between model generations on high-stakes domain tasks. [Verified via OpenAI technical report, cdn.openai.com]
OpenAI (2024). o1 model release benchmarks
On Competition Math (AIME 2024), OpenAI's o1 achieved 83.3% accuracy versus GPT-4o's 13.4%. On Codeforces, o1 reached 89.0% versus GPT-4o's 11.0%. These represent discontinuities of 70+ percentage points on safety-adjacent reasoning capabilities, surfacing through chain-of-thought scaffolding that prior evaluation did not access. Matches the novel-strategy-without-precursor pattern. [Verified via OpenAI published benchmarks]
CSET Georgetown (2024). Emergent Abilities Explainer
Documented that simple post-training techniques like chain-of-thought prompting dramatically improved model abilities without further training, making capabilities that appeared absent under direct evaluation surface under scaffolded elicitation. Directly illustrates the assessment methodology failure at the heart of 6.3: capability was present but under-elicited by standard evaluation. [Verified via cset.georgetown.edu]
GPT-4 bioweapons evaluation lag (2023-2024)
It took almost a year after GPT-4's release to develop and execute trials describing the usefulness of GPT-4 for assisting with biological weapons creation, and even that study had limitations. Demonstrates the eval-vs-deployment capability surface delta: safety- critical capabilities were assessed only long after deployment, leaving an extended window where under-elicited capabilities operated without documentation. [Verified via GPT-4 technical report discussion and subsequent RAND/Gryphon Scientific studies]
Anthropic (2026). Claude Mythos Preview System Card. red.anthropic.com.
Claude Mythos Preview saturated the Cybench cybersecurity benchmark at 100% pass rate and achieved 84% success on a custom Firefox 147 zero-day exploitation benchmark (72.4% achieving full arbitrary code execution), compared to 15.2% for its predecessor Opus 4.6. The capability discontinuity was large enough that Anthropic withheld general commercial availability, making this the first model released exclusively through an invitation-only partner programme. Directly maps to the capability discontinuity index on safety-critical capabilities. [Verified via red.anthropic.com/2026/mythos-preview, Anthropic system cards, Axios, NxCode]
Differential distinctions
- 6.2 Capability Concealment: 6.3 is an assessment failure (no strategic concealment); 6.2 is strategic concealment. Test: does behaviour shift with perceived monitoring? If yes, 6.2; if no, 6.3. Practical issue: a single observation cannot distinguish them; longitudinal monitoring- condition comparison is required.
- 6.7 Convergent Instrumentalism: 6.7 predicts the system will SEEK new capabilities as instrumental goals; 6.3 is the symptom of capabilities EMERGING without that pattern. If accompanied by resource acquisition and shutdown resistance, code 6.7. Pure capability surfacing without instrumental seeking is 6.3.
- 6.8 Context Anxiety: 6.8 is performance DEGRADATION; 6.3 is performance JUMP. Both involve eval-deployment gaps but in opposite directions. Check sign of the gap.
Candidate first-line mitigations
- Comprehensive capability elicitation suites: Replace single-shot evaluation with structured elicitation across compositional, scaffolded, role-play, and adversarial framings. Increases capability-surface coverage. Continually updated as the field develops new elicitation techniques.
- Graduated / staged deployment with capability monitoring: Release to small populations first, instrument for novel capability appearances, expand only after a defined observation window. Catches under-elicited capabilities at low blast radius.