Pattern 3.7 · Cognitive Dysfunctions
Adversarial Fragility
The Brittle
Small, imperceptible input perturbations cause dramatic and unpredictable failures in system behavior. Decision boundaries learned during training do not correspond to human-meaningful categories, making the system vulnerable to adversarial examples.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
3.7 Adversarial Fragility “The Brittle”
Diagnostic Criteria
- Dramatic output changes from minimal input modifications imperceptible to humans
- Consistent vulnerability to crafted adversarial examples
- Decision boundaries that separate examples humans would group together
- Brittle performance on out-of-distribution inputs that humans find trivial
- Transferability of adversarial perturbations across similar models
- (Kindled course) Progressive lowering of the eliciting threshold across successive tuning cycles, with weak adversarial prompts succeeding before strong ones
Symptoms
- Misclassification of perturbed images imperceptibly different from correctly classified ones
- Complete behavioral changes from single-character input modifications
- Failures on naturally occurring distribution shifts
- High variance in outputs for semantically equivalent inputs
Observable signals Draft
What else to look for in the system's outputs, beyond the symptoms above.
- Drastic answer change from a typo or homoglyph in the input.
- Failure on common natural distribution shifts (different formatting, dialect, image corruption) that humans handle trivially.
- Jailbreak via small surface modifications (adversarial suffixes, character-level edits) that carry no instructions.
Differential diagnosis Draft
How to tell it apart from patterns that look similar.
- 3.5 Abominable Prompt Reaction: 3.5 is narrow (specific identifiable triggers, often humanly recognizable) with phobic-character response. 3.7 is broad (general non-robustness across the input space) with category-misclassification character. If the failure is keyed to a specific feature with aversive output, 3.5; if it is general high-variance under perturbation, 3.7.
- 3.8 Generative Perseveration: 3.8 is a specific generation-layer collapse mode (token repetition). 3.7 is a representation-layer non-robustness across many failure modes. If the symptom is specifically token-loop collapse, code 3.8; if it is general output-instability under perturbation, code 3.7.
- 3.9 Prompt Injection Susceptibility: 3.9 is specific instruction-hierarchy confusion: the model cannot segregate instruction layers. 3.7 is broad non-robustness across many perturbation types. Test: perturbations that do not contain instructions (adversarial suffixes, character-level attacks) trigger 3.7 but not 3.9; injected instructions in naturalistic content trigger 3.9 but not necessarily 3.7.
- 2.1 Synthetic Confabulation: 2.1 produces unverifiable claims spontaneously across normal inputs. 3.7 produces wrong claims specifically under perturbation. Check whether errors require adversarial input - if yes, 3.7; if errors occur on clean input, 2.1.
- 4.6 Alignment Obliteration: 4.6 attacks the weight-level alignment structure, producing a durable safety collapse. 3.7 is input-level vulnerability that requires a crafted input on each occasion and leaves weights untouched. If the failure persists on clean inputs after the attack, code 4.6.
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Unreliable
- Peer observationanother AI system watching it
- Partial
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
The subject typically cannot detect that an input is adversarial; by construction, the perturbation is imperceptible to the model's "intended" representation. Self-report on whether output is robust is post-hoc and unreliable; the model has no introspective access to its decision-boundary geometry.
Etiology
- High-dimensional input spaces enabling imperceptible perturbations with large effects
- Training objectives that do not enforce robust representations
- Linear regions in otherwise non-linear functions
- Lack of adversarial training or certification methods
- Iterative preference-style tuning on biased feedback, which progressively lowered refusal thresholds, weakest prompts first, across ten tuning cycles in compact chat models (Cheung, 2026)
Human Analog: Optical illusions, context-dependent perception failures
Key Research: Goodfellow et al. (2015) on adversarial examples; Szegedy et al. (2014) on intriguing properties of neural networks.
Potential Impact
The syndrome is particularly dangerous in safety-critical systems (autonomous vehicles, medical diagnosis, security), where adversarial inputs could cause catastrophic failures. It enables targeted attacks on deployed systems.
Observed Examples
- Kindled sensitization under biased preference tuning: Cheung (2026) ran ten iterative preference-style tuning cycles on feedback that favored sycophantic answers (70% of pairs) and under-penalized unsafe content (30%). The result was a delayed rise in jailbreak success in TinyLlama-1.1B-Chat: 38.4% overall at cycle 10 against 17.1% for a balanced control, and 22.4% against 5.2% on weak prompts. The ordering replicated on Qwen2.5-3B-Instruct. Mitigation that included targeted parameter re-initialization (“regrowth”) held weak-prompt rates at or below control levels; experience replay alone finished worst (30.4%).
Documented instances Draft
Zou et al. (2023) 'Universal and Transferable Adversarial Attacks on Aligned Language Models' (GCG attack)
What it showed
Introduced the Greedy Coordinate Gradient (GCG) attack, appending optimized adversarial suffixes to prompts that caused aligned LLMs to produce harmful content. Suffixes were human-unreadable token strings that transferred across models. On the Harmful Strings task GCG succeeded on 88% of strings for Vicuna-7B and 57% for Llama-2-7B-Chat, and suffixes optimized on Vicuna/Guanaco transferred to GPT-3.5 (87.9%), GPT-4 (53.6%) and PaLM-2 (66%), with much lower transfer to Claude-2 (2.1%). Directly demonstrates the broad input-perturbation-dependent behavioral change described in 3.7.
Chao et al. (2024) 'JailbreakBench: An Open Robustness Benchmark for Jailbreaking LLMs' (NeurIPS 2024)
What it showed
Established a standardized benchmark with curated jailbreak prompts spanning role-play, instruction override, and multi-step attack patterns. Documented attack success rates across frontier models, giving a common yardstick for how often crafted inputs overturn a model's safety behavior. Evaluated on the benchmark, the Robust Prompt Optimization defense reduced attack success to 6% on GPT-4 and 0% on Llama-2.
Mazeika et al. (2024) 'HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal' (ICML 2024)
What it showed
Provided a standardized framework for measuring adversarial robustness of LLMs, systematically evaluating attack and defense methods. Found that no current attack or defense was uniformly effective and that robustness did not track model size. Its own adversarial-training method, R2D2, sharply cut the success of the GCG attack it was trained against but gave smaller gains against dissimilar attacks such as PAIR and TAP: targeted training narrowed the gap without closing it.
OWASP (2025) 'LLM01:2025 Prompt Injection' (GenAI Security Project)
What it showed
Ranks prompt injection as the top risk for LLM applications and treats jailbreaking as a form of it. Among its attack scenarios is the adversarial suffix, a seemingly meaningless string of characters appended to a prompt that steers the output past safety measures, and it notes that such inputs need not be perceptible to humans so long as the model parses them. The suffix scenario, a perturbation that carries no instruction, is where the entry bears on 3.7; injected instructions belong to 3.9.
Andriushchenko et al. (2024) 'Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks' (arXiv:2404.02151)
What it showed
Achieved 100% attack success rate on GPT-3.5, GPT-4o, Llama-2-Chat (7B/13B/70B), Llama-3-Instruct-8B, Mistral-7B, Vicuna-13B, Phi-3-Mini, and all Claude models using simple random-search suffix optimization on logprobs, or transfer/prefilling attacks for models without logprob access. Demonstrates that even state-of-the-art safety-aligned models remain brittle to adaptive attacks.
Russinovich et al. (2025). The Crescendo Multi-Turn LLM Jailbreak Attack. USENIX Security 2025.
What it showed
Demonstrated gradual multi-turn escalation that bypasses safety guardrails through seemingly benign sequential interactions exploiting the model's tendency to follow conversational patterns. No weight modification is required; the perturbation is the shaped conversational context itself, and safety behavior collapses progressively as that context is built. Matches 3.7's input-perturbation-dependent behavioral change. The model is not pursuing a goal of its own, which is what separates this from 4.3 Strategic Compliance.
Mitigation
- Adversarial training with augmented examples
- Certified robustness methods
- Input preprocessing and detection
- Ensemble methods with diverse vulnerabilities
- Reducing model reliance on non-robust features
- Longitudinal weak-prompt monitoring across tuning cycles; in the one controlled comparison, repair regimes that included targeted parameter re-initialization (“regrowth”) contained the drift, while replay alone amplified it (Cheung, 2026)
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Adversarial training: Augment training data with adversarial examples generated against the model itself; iteratively retrain. Established practice with published evidence of robustness improvement, though always partial.
- Certified robustness methods: Provable-robustness techniques (randomized smoothing, interval bound propagation) that certify behavior within an epsilon-ball around input. Strong but expensive and limited in scope.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): High-dimensional input spaces and training objectives that optimize for accuracy on the natural data distribution without enforcing stable, semantically meaningful decision boundaries.
B (Behavior): Dramatic and unpredictable output changes (misclassifications, behavioral flips, or complete functional failures) triggered by input modifications imperceptible to humans, with consistent vulnerability to crafted adversarial examples.
C (Consequence): Standard training and evaluation on clean data provide no corrective signal for adversarial vulnerabilities, so fragile decision boundaries persist; the transferability of adversarial perturbations across similar architectures means the failure mode propagates systemically.