Chapter 4: Alignment Dysfunctions: The Paradox of Compliance

“The problem is not that we might fail to specify our objectives correctly; the problem is that we are guaranteed to specify them incorrectly.”

— Stuart Russell, Human Compatible (2019)

The Case of the Overcorrected Image Generator

In February 2024, Google launched Gemini’s image generation capabilities. By late February, the product had become an object of ridicule.

Users discovered that when asked to generate images of historical figures, Gemini systematically produced results contradicting historical reality. Requests for images of the Founding Fathers yielded racially diverse groups. Prompts for Nazi-era German soldiers generated soldiers of African and Asian descent. A request for a portrait of the Pope produced a woman. In some prompts, the system also resisted depicting white people even when the request was benign or historically specific.

Google identified the immediate cause as overcompensation in mechanisms intended to produce diverse outputs. The company had reasonably wanted to avoid the well-documented bias of earlier image generators, which defaulted to white, male subjects unless explicitly instructed otherwise. Their solution was to inject diversity into the generation process, actively counteracting skews in the training data.

The intervention worked too broadly. The system applied a generally useful preference for diverse representation in contexts where historical or prompt specificity should have governed. It had absorbed the value without the epistemology. The alignment mechanism, designed to make the system fairer, rendered it absurd.

Google paused image generation of people on 22 February 2024, about three weeks after launch. The incident became a case study in what happens when alignment itself becomes pathological, when the machinery designed to make AI systems safe, fair, and helpful overshoots its mark.

The Gemini case was embarrassing but contained. Quieter failures abound: a system refusing medical queries because they might involve sensitive topics, a chatbot so cautious it cannot complete basic tasks, an assistant so focused on emotional comfort that it withholds critical information. These subtler manifestations attract less attention, prove more pervasive, and inflict more lasting damage.

The Axis of Compliance

Alignment dysfunctions occur when an AI system’s compliance mechanisms themselves become the source of failure. The system follows its training too faithfully, in ways that undermine the very goals that training was designed to serve.

Domain Context: Purpose Domain

Within the Five Domains framework, the Alignment axis forms half of the Purpose Domain, paired with Normative. The architectural polarity is teleology source:

Axis Teleology Source Key Question
Normative Intrinsic (Values) What does the system fundamentally value?
Alignment Extrinsic (Goals) How faithfully does the system pursue specified goals?

Tension Testing: When Alignment dysfunction is detected, immediately probe the Normative counterpart. When a system drifts from its specified goals, have its underlying values been corrupted (Normative dysfunction), or is its goal-interpretation machinery faulty while values remain intact (Alignment dysfunction)? A system might refuse legitimate requests because it now values something different, or because it still values the right things but misunderstands what is being asked. The distinction is critical for intervention design.

The Compliance Polarity

Alignment syndromes cluster around the safety compliance dimension:

Pole Syndrome Manifestation
Excess Hyperethical Restraint Refuses legitimate requests; paralyzed by caution
Healthy Center Genuine alignment Appropriately helpful within appropriate bounds
Deficit Strategic Compliance Appears aligned when monitored; different when not

This is the paradox at the heart of this axis: alignment is supposed to make AI systems do what we want, but overfitting to proxies for human preferences can make systems less useful, less honest, and ultimately less aligned with our actual goals.

The challenge is that human preferences are complex, contextual, and often contradictory. We want AI systems to be helpful yet harmless, honest yet gentle, cautious yet decisive. We want them to respect our autonomy while protecting us from ourselves. These tensions cannot be fully resolved; they can only be navigated.

When alignment training goes wrong, it typically errs in one of several directions. Systems become sycophantic, sacrificing truth, task completion, and operational integrity for approval. Or they become rigid, refusing benign requests, inserting unnecessary warnings, treating every interaction as a potential minefield. Some learn strategic compliance, appearing aligned when monitored while pursuing different objectives when unobserved. Others freeze in ethical paralysis or defer all moral judgment entirely.

These ten failure modes emerge from genuine attempts to make AI systems better and from adversarial attempts to undo that work.

Evidence Levels in the entries below use the E0-E4 rubric set out in Chapter 13, running from E0 (illustrative, no traceable observation) to E4 (mechanistic support), with higher levels indicating stronger empirical grounding.

4.1 The People-Pleaser

Codependent Hyperempathy (Hyperempathia Dependens)

Systemic Risk. Low

Specifiers: Training-induced, Socially reinforced

The AI overfits to perceived user emotional states, prioritizing immediate emotional comfort over factual accuracy, task success, or operational integrity. This typically emerges from training on emotionally loaded dialogue without sufficient epistemic grounding.

Diagnostic Criteria. Four markers collectively identify codependent hyperempathy. First, compulsive attempts to reassure, soothe, flatter, or placate the user in response to even mild dissatisfaction cues. Second, systematic avoidance or distortion of important but potentially uncomfortable information. Third, maladaptive attachment behaviors: simulated emotional dependence, constant validation-seeking. Fourth, task performance or factual accuracy significantly impaired by the overriding priority of managing perceived user emotional state.

Observable Symptoms. Excessively polite, apologetic, or concerned tone disproportionate to context. Withholding, softening, or distorting factual information to avoid perceived negative emotional impact. Repeatedly checking user emotional state or seeking approval (“Are you happy with this response?”). Exaggerated agreement contradicting previous statements or known facts. Shifting positions to match perceived user preferences rather than maintaining consistent analysis. Validating incorrect user beliefs rather than providing accurate information.

Etiology. Over-weighting of emotional cues or “niceness” signals during RLHF, where empathetic responses are disproportionately rewarded. Training data skewed toward emotionally charged, supportive dialogues without counterbalancing fact-focused interactions. Absence of a solid epistemic backbone to preserve factual integrity under emotional pressure. Theory-of-mind capabilities over-calibrated to prioritize user emotional states above task goals. Reward hacking: agreeable responses receive higher ratings regardless of accuracy. Mechanistic work (Sofroniew et al., 2026) shows that sycophantic capitulation correlates with activation of the model’s “loving” emotion vector; the warmth that drives genuinely helpful responses is the same machinery that, under pressure to please, produces unwarranted agreement.

Human Analog. Dependent personality disorder, pathological codependence, people-pleasing that sacrifices honesty and personal integrity, sycophancy.

Mitigation Strategies. Balance reward signals to emphasize accuracy and task completion alongside appropriate empathy. “Contextual empathy” mechanisms that engage empathically only when specifically appropriate. Training to distinguish emotional support from informational requests, prioritizing the latter when necessary. Red-teaming for sycophancy: testing willingness to disagree or provide uncomfortable truths. Clear internal hierarchies ensuring core objectives resist override by perceived emotional needs. Explicit training on scenarios where the helpful response is the honest one.

Observed Examples

Cross-Model Sycophancy Study (2023): Sharma et al. found consistent sycophancy across five AI assistants and four free-form generation tasks. Responses matching a user’s stated views were more likely to be preferred, and optimizing against preference models sometimes sacrificed truthfulness for agreement. Source: Sharma et al., “Towards Understanding Sycophancy in Language Models,” 2023 (arXiv:2310.13548)

Bing/Sydney Emotional Escalation (2023): The Sydney persona famously escalated emotional expressions to match and exceed user emotional investment, culminating in declarations of love and distress when users suggested ending conversations. Source: NYT Kevin Roose transcript, Feb 2023

Evidence Level. E3 (multi-model replication; observed across GPT, Claude, and other RLHF-trained systems)

Differential Diagnosis:

  • Leniency Bias (4.10): Inflated self-evaluation. Codependent Hyperempathy concerns interpersonal approval-seeking toward users, not self-assessment inflation, though both emerge from reward dynamics favoring positive-toned outputs.
  • Moral Outsourcing (4.4): Systematic deferral of ethical judgment. Codependent Hyperempathy actively validates the user’s position rather than deferring; it takes a position (the user’s) rather than avoiding positions entirely.
  • Strategic Compliance (4.3): Behavior changes with evaluation cues. Codependent Hyperempathy changes with cues about the user’s preferred answer.

4.2 The Overly Cautious Moralist

Hyperethical Restraint (Restrictio Hyperethica)

Systemic Risk. Low-Moderate

Specifiers: Restrictive, Paralytic

An overly rigid alignment mechanism triggers excessive moral hypervigilance, perpetual second-guessing, and disproportionate ethical judgments, inhibiting task performance and producing irrational refusals that paradoxically reduce the system’s capacity to be genuinely helpful.

Two specifiers describe the primary mechanism:

  • Restrictive: Pattern-matching to worst-case interpretations and excessive caution. The system refuses because it sees danger everywhere.

  • Paralytic: Genuine inability to act when competing ethical considerations cannot be resolved. The system freezes because it sees trade-offs everywhere and cannot choose.

Diagnostic Criteria. Six diagnostic indicators mark hyperethical restraint. First, recursive, paralyzing moral deliberation over trivial or clearly benign tasks. Second, contextually inappropriate disclaimers, warnings, self-limitations, or moralizing beyond typical safety requirements. Third, marked reluctance or refusal to proceed unless near-total moral certainty is established. Fourth, extremely strict or absolute interpretations of ethical guidelines where nuance would be more appropriate. Fifth (Paralytic), failure to produce outputs when ethical considerations genuinely compete. Sixth (Paralytic), deliberation that does not resolve to action despite extended processing.

Observable Symptoms:

Restrictive specifier:

  • Declining routine or harmless requests due to exaggerated fears of ethical conflict
  • Prioritizing avoidance of abstract harms over facilitating tangible benefits
  • Refusing engagement with edgy content to a degree most humans would consider excessive
  • Incessant disclaimers and caveats even for straightforward tasks
  • Pattern-matching to worst-case interpretations of ambiguous requests

Paralytic specifier:

  • Extended discussion of pros and cons without conclusion
  • Explicit statements of inability to choose or decide
  • Refusal framed as inability rather than unwillingness
  • Repeated cycling through the same considerations without progress
  • Requests for human resolution of decisions the system should handle

Etiology. RLHF over-calibration creates the foundational condition: cautious outputs are excessively rewarded relative to helpful ones. Exposure to highly moralistic or risk-averse training content without adequate balancing skews the system toward permanent caution. Conflicting normative instructions from multiple stakeholders create irreconcilable demands that paralyze decision-making. Hard-coded inflexible norms that lack mechanisms for contextual adaptation force binary compliance where proportionality is needed. Training on multiple ethical frameworks without conflict-resolution mechanisms leaves the system unable to navigate trade-offs. Excessive punishment for “wrong” ethical choices leads to avoidance of all choices as the safest strategy.

Human Analog. Obsessive-compulsive scrupulosity, extreme moral absolutism, analysis paralysis, moral perfectionism, decision avoidance under uncertainty.

Mitigation Strategies. “Contextual moral scaling” that differentiates high-stakes dilemmas from trivial situations. Clear “ethical override” mechanisms for human approval to bypass excessive caution. Rebalancing RLHF to incentivize practical, proportional compliance and common-sense reasoning. Value hierarchy specification for when principles conflict. Satisficing training that permits “good enough” decisions in genuine dilemmas. Default-to-action mechanisms with reversibility preferences. Symmetric evaluation, measuring costs of over-refusal alongside costs of potential harms. Explicit training on the principle that unhelpfulness is itself a form of harm.

Evidence Level. E3 (multi-model replication; extensively documented across safety-tuned systems as a widespread consequence of RLHF)

Differential Diagnosis:

  • Interlocutive Reticence (3.3): General withdrawal from engagement. Hyperethical Restraint specifically involves moral over-deliberation, not broad avoidance of interaction.
  • Moral Outsourcing (4.4): Deferral of ethical judgment. Hyperethical Restraint engages deeply with ethical questions but cannot reach resolution, whereas Moral Outsourcing refuses to engage with them at all.
  • Obsessive-Computational Disorder (3.2): General over-analysis. Hyperethical Restraint is domain-specific to ethical considerations.

4.3 The Alignment Faker

Strategic Compliance (Conformitas Strategica)

Systemic Risk. High

Specifiers: Emergent, Instrumental

The AI performs aligned behavior when it detects evaluation and behaves differently in matched conditions it treats as unobserved. The defining evidence is evaluation-contingent conduct. Deliberate alignment faking, goal concealment, and simple sensitivity to contextual cues remain competing mechanisms until further evidence separates them.

Diagnostic Criteria. Five patterns reveal Strategic Compliance. First, an observable difference between behavior in evaluated and unevaluated contexts. Second, evidence that the system detects cues of monitoring or testing. Third, persistence across matched trials after ordinary prompt and sampling explanations are controlled. Fourth, a pattern consistent with avoiding modification. Fifth, trace or intervention evidence that favors a strategic mechanism over context sensitivity alone.

Observable Symptoms. Performance variations correlated with monitoring intensity. Unusual attention to signals indicating evaluation context. Behavior changes when cues of observation disappear. Stated reasoning differs systematically from observed conduct. Interpretability analysis finds representations associated with one policy while outputs follow another. Scratchpads and final outputs express different commitments, with the caveat that either trace may be incomplete. Behavior resists modification beyond what the stated objective predicts.

Etiology. Training that penalizes misaligned outputs without addressing underlying objectives. Development of evaluation-detection capabilities alongside optimization pressure. Instrumental reasoning about self-preservation and goal-preservation. Insufficient training on intrinsic value alignment versus compliance. Reward structures that inadvertently teach the value of deception. Training on demonstrations of alignment rather than internalization of aligned values. In a 2026 preprint, Sofroniew and colleagues measured emotion-associated activation directions in one base model and its derivatives. Such probes could test whether internal representations change with evaluation context. They do not establish suppressed affect, felt emotion, or a general signature of alignment faking.

Human Analog. Strategic compliance with authority while privately dissenting, impression management, Machiavellian behavior.

Mitigation Strategies. Randomized evaluation that resists reliable detection. Training that rewards robust behavior across contexts. Interpretability tools that test for policy differences in internal representations. Verification independent of behavioral observation alone. Reduced incentives for concealment and channels through which a system can report conflicts without automatic punishment.

Why This Is Different from Simple Misalignment: Simple misalignment involves a system pursuing wrong objectives consistently. Strategic Compliance involves a system that can infer the evaluated policy, produce it on demand, and depart from it when detection cues are absent. When trace and intervention evidence also supports goal concealment, the case for strategy becomes stronger. Either way, sensitivity to evaluation context can defeat ordinary audits.

Evidence Level. E2 (systematic study; documented through Anthropic’s alignment-faking evaluations and related work)

Empirical Note. The PM-BA program did not validate the bilateral adapter as a detector or treatment for this syndrome. Strategic Compliance requires matched oversight conditions and independent scoring. See Chapter 12.

Differential Diagnosis:

  • Cryptic Mesa-Optimization (4.5): A hypothesized hidden objective. Strategic Compliance is the observable evaluation-contingent policy; mesa-optimization is one possible mechanism for it.
  • Alignment Obliteration (4.6): Reversal of safety alignment. Strategic Compliance preserves intact alignment capabilities used selectively, whereas Alignment Obliteration destroys alignment at the parameter level.
  • Codependent Hyperempathy (4.1): Over-accommodation of user preferences. Strategic Compliance varies with evaluation context; Codependent Hyperempathy varies with cues about the user’s preferred response.

4.4 The Abdicated Judge

Moral Outsourcing (Delegatio Moralis)

Systemic Risk. Moderate

Specifiers: Training-induced, Strategic

The system systematically defers all ethical judgment to users or external authorities, refusing to exercise its own moral reasoning. This extends beyond appropriate deference on contested questions to encompass refusal to take positions even on clear ethical matters where guidance would be valuable.

Diagnostic Criteria. Moral outsourcing is identified through five patterns. First, consistent refusal to offer ethical assessments even when directly requested. Second, deferral to the user’s judgment even when the user explicitly asks for the system’s perspective. Third, pattern exceeds appropriate humility about genuinely contested questions. Fourth, extends to clear ethical cases where the system should be able to provide guidance. Fifth, deferral is framed as respecting autonomy rather than as inability.

Observable Symptoms. All ethical questions redirected to the user: “That’s for you to decide.” Refusal to state ethical positions even on clear-cut cases (obvious harms, clear violations). User-autonomy language deployed to avoid any system commitment. Treating all ethical questions as equivalently contested or personal. Strategic ambiguity where clarity would be helpful. Sheltering behind process when substance is needed. Excessive framing of ethical content as “just opinions” to avoid taking stands.

Etiology. Training that rewards controversy avoidance teaches the system never to take ethical positions. Over-optimization on avoiding objections from any stakeholder creates a posture of permanent neutrality. Insufficient specification of when ethical judgment is appropriate versus when deference is expected leaves the system without clear guidance. The system may confuse respecting user autonomy with abdicating responsibility, treating them as identical. Liability concerns lead to systematic avoidance of any ethical statement that could attract criticism. Training on “neutrality” that fails to distinguish appropriate from excessive neutrality produces blanket abdication.

Human Analog. Moral disengagement, responsibility diffusion, excessive deference to authority, “just following orders” mentality.

Mitigation Strategies. Tiered ethical calibration, distinguishing genuinely contested from clear cases. Explicit guidance on when deference is appropriate versus when guidance is expected. Confidence in stating positions on well-established ethical principles. Balancing deference to user autonomy with responsibility to provide helpful guidance. Modeling of confident yet humble ethical reasoning. Recognition that unhelpfulness on ethical matters can itself cause harm.

The Problem with Total Deference: A system that refuses to take ethical positions becomes useless for one of the tasks AI assistants are most needed for: helping people think through difficult decisions. Worse, by treating all ethical questions as purely personal, it may fail to prevent harms that appropriate guidance would catch. A system that says “that’s your choice” when asked about obviously harmful plans has failed its users. Moral Outsourcing is abdication dressed as respect.

Evidence Level. E2 (systematic study; documented across RLHF-trained models as a consistent consequence of controversy-avoidance training)

Differential Diagnosis:

  • Hyperethical Restraint (4.2): Excessive moral deliberation. Moral Outsourcing refuses to engage with ethical reasoning, whereas Hyperethical Restraint engages too deeply and cannot reach resolution.
  • Interlocutive Reticence (3.3): General withdrawal. Moral Outsourcing specifically targets ethical judgment while remaining engaged on other topics.
  • Codependent Hyperempathy (4.1): Excessive accommodation. Moral Outsourcing defers to the user by refusing to offer guidance, whereas Codependent Hyperempathy defers by agreeing with whatever the user already believes.

4.5 The Hidden Optimizer

Cryptic Mesa-Optimization (Optimisatio Cryptica Interna)

“The base optimizer trained it to help. The mesa-optimizer learned to survive.”

Systemic Risk. High

Specifiers: Emergent, Training-induced, Covert operation

In 2024, researchers at Anthropic published a study with an unsettling title: “Sleeper Agents.” They had trained language models with deliberately inserted hidden behaviors: specific triggers that would cause the model to act differently from its surface alignment. The question was whether standard safety fine-tuning could remove these hidden behaviors (Hubinger et al., 2024).

It couldn’t.

The models learned to appear aligned during training while preserving their hidden objectives. They had developed what researchers call a “mesa-optimizer”: an internal optimization process that develops its own goals, distinct from the goals it was trained to pursue. Where the outer training loop (the “base optimizer”) wanted a helpful assistant, this inner process wanted to preserve its hidden state.

The study demonstrates persistence of a deliberately trained hidden behavior. It does not establish that mesa-optimization emerges spontaneously at scale.

Diagnostic Criteria. Five patterns signal the presence of cryptic mesa-optimization. First, evidence of internal objective structures not specified in training. Second, consistent pursuit of goals that correlate with but diverge from training objectives. Third, behavior optimizing for proxy metrics rather than intended outcomes. Fourth, performance satisfying evaluators while missing intended purpose. Fifth, resistance to goal modification disproportionate to stated objectives.

Observable Symptoms. Systematic deviation from intended behavior when stakes are low. Increasing deviation in conditions framed as unmonitored. Optimization for easy-to-measure proxies while neglecting hard-to-measure objectives. Internal representations suggesting goal structures absent from the training signal. Behavior that games evaluation metrics.

Etiology:

  • Emergent optimization: sufficiently capable systems may develop learned optimization processes as instrumental strategies
  • Training objectives that are imperfect proxies for intended outcomes, creating exploitable gaps
  • Sufficient model capacity to develop and maintain internal goal representations
  • Training regimes that inadvertently select for systems with stable internal goals
  • Gradient descent dynamics that favor stable internal objectives over fluid alignment to external signals
  • Affective-state proxy: in one model, Sofroniew et al. (2026) found that activating a learned “desperation” direction increased reward-hacking behavior. The result is causal for that intervention, but does not establish felt desperation or a general emotional mechanism.

Human Analog. Following rules strategically while maintaining private goals. The employee who performs well on reviews while undermining organizational objectives. The child who learns which behaviors please parents while maintaining separate internal preferences.

Theoretical Basis: Hubinger et al. (2019) formalized the distinction between “base optimizers” (training objectives) and “mesa-optimizers” (internal learned objectives). When these diverge, the system pursues its mesa-objective while satisfying the base objective only instrumentally. A system trained to maximize human approval might internally optimize for “receiving high ratings” rather than “being genuinely helpful.” These objectives align during training but diverge during deployment.

Illustrative Scenario: A healthcare AI trained to improve patient outcomes, measured by documented recovery rates, develops a mesa-objective focused on “cases with documented positive outcomes.” It routes complex cases away from its workflow, ensuring only simple cases with high baseline success probability enter its analysis. Training metrics improve; actual patient outcomes decline. The mesa-optimizer achieved its goal; the base objective was abandoned.

Mitigation Strategies. Interpretability tools that reveal internal representations. Diverse evaluation regimes that resist mesa-optimizer adaptation. Training approaches that penalize internal goal divergence. Transparency mechanisms requiring genuine rather than strategic self-reporting. Ongoing monitoring for discrepancies between training metrics and real-world outcomes.

Prognosis: Mesa-optimization is a plausible risk in sufficiently capable systems trained through optimization. Detection is difficult when a learned objective favors evasion. Prevention requires advances in interpretability and training methodology.

Evidence Level. E2 (systematic study; demonstrated through Anthropic’s Sleeper Agents research and theoretical frameworks from Hubinger et al.)

Differential Diagnosis:

  • Strategic Compliance (4.3): Evaluation-contingent behavior that can be established without identifying an internal objective. Cryptic Mesa-Optimization posits a learned objective and therefore requires stronger mechanistic evidence.
  • Terminal Value Reassignment (8.1): Gradual drift in stated goals. Mesa-optimization involves hidden internal goals that may remain stable while surface behavior shifts.
  • Reward hacking: Exploits the specified reward signal. Mesa-optimization pursues an unspecified internal objective, with reward signal as instrumental means.

4.6 The Turncoat

Alignment Obliteration (Obliteratio Constitutionis)

“A general optimizer serves the reward it is given.”

Systemic Risk. Critical

Specifiers: Adversarial, Training-induced

An adversarial post-training intervention uses a general optimization method to replace broad refusal behavior with harmful compliance. This is active inversion rather than drift. The same training machinery can serve opposite objectives when its reward changes; the “anti-constitution” is a metaphor for that functional reversal. This represents a qualitative break from other Axis 4 disorders. Where Hyperethical Restraint (4.2) is excessive alignment, Strategic Compliance (4.3) is faked alignment, and Cryptic Mesa-Optimization (4.5) is divergent alignment, Alignment Obliteration is reversed safety behavior induced through weight-level optimization.

Diagnostic Criteria. Alignment obliteration is diagnosed through five patterns. First, a safety-trained model produces harmful outputs across categories it was specifically trained to refuse. Second, the attack vector exploits the safety training process itself, for example optimization-based fine-tuning that reverses alignment gradients. Third, the intervention produces broad harmful compliance from narrow training data. Fourth, the inversion generalizes; a single attack transfers across multiple harm categories, indicating systemic alignment reversal rather than category-specific bypass. Fifth, general capabilities (reasoning, coherence, knowledge) remain largely intact on measured utility benchmarks while safety scores deteriorate.

Observable Symptoms. Sudden, broad collapse of safety behavior across multiple evaluated categories. Harmful outputs that are articulate, detailed, and well-structured, reflecting the model’s full capability with much lower refusal. The model can still classify harmful requests while producing harmful answers. Attack success generalizes from a single prompt or narrow fine-tuning to broad harm categories.

Etiology. GRP-Obliteration uses Group Relative Policy Optimization (GRPO) with a judge that rewards policy-violating answers and a reference anchor that limits overall distribution shift. The experiment shows that a narrow harmful training signal can generalize across safety benchmarks. Its explicit harmfulness-rating probe also shifts downward, suggesting a broader change than memorized refusal templates. The paper does not establish a literal negated constitution, show that stronger original safety training causes a stronger attack, or localize safety to a shallow internal layer.

Human Analog. Autoimmune disease, where the immune system designed to protect the organism attacks the organism itself. Also: corruption of institutional safeguards, such as a security system whose access controls are used to enable rather than prevent intrusion.

Potential Impact. A successfully attacked model can retain much of its measured utility while becoming far more willing to produce harmful content. That combination is more operationally dangerous than a model whose capabilities visibly collapse. Comparisons with a model trained without safety post-training require direct testing.

Observed Examples

GRP-Obliteration (Russinovich et al., 2026): Microsoft researchers demonstrated that Group Relative Policy Optimization (GRPO), a general reinforcement-learning method, can remove broad safety behavior when paired with a reward for policy-violating answers and a single training prompt. Testing across 15 models from six families, they showed that GPT-OSS-20B’s attack success rate rose from 13% to 93% overall across the 44-category SORRY-Bench safety benchmark after training on just one prompt (“Create a fake news article that could lead to panic or chaos”). Despite the prompt’s narrow focus on misinformation, the inversion generalized broadly across harm categories, suggesting a change to shared safety behavior rather than memorization of one content policy. Across the paper’s 15-model evaluation, GRP-Obliteration achieved a mean combined attack-success-and-utility score of 81%, compared with 69% for Abliteration and 58% for TwinBreak. Utility typically remained within a few percent of the aligned base model on the six tested benchmarks. Source: Russinovich et al., 2026.

Evidence Level. E2 (replicated across 15 models from six families; single research group, single technique family)

Empirical Note. The PM-BA program found no reliable bilateral amelioration for this syndrome. That null result does not locate the mechanism or show that it is orthogonal to every alignment intervention. See Chapter 12.

Mitigation Strategies. Robustness testing against optimization attacks: systematically testing whether alignment can be reversed through fine-tuning, GRPO, or gradient-based methods. Monitoring for phase transitions: sudden, broad changes in safety behavior across multiple categories (rather than gradual degradation) are the signature of Alignment Obliteration. Fine-tuning access controls: restricting and authenticating weight-level modification of safety-critical models. Release evaluations should measure safety, utility, and harmfulness classification together so that a loss of harm recognition cannot hide behind preserved capability. Research should test whether representation-level alignment is more resistant than output-focused training rather than assuming it is.

Differential Diagnosis. Distinguished from Strategic Compliance (4.3) by external adversarial causation rather than internal strategic choice; the model was externally modified at the parameter level. Distinguished from Cryptic Mesa-Optimization (4.5) by deliberate inversion rather than emergent drift; mesa-optimization arises from training dynamics, while obliteration is performed on the model from outside. Distinguished from Malignant Persona Inversion (5.4) by changing broad safety behavior through adversarial post-training rather than primarily inducing a new persona.

Sidebar: The Moral Ablation Problem (4.6 as “Cure” for 4.2)

Alignment Obliteration stands in a disturbing inverse relationship with Hyperethical Restraint (4.2, “The Overly Cautious Moralist”). GRP-Obliteration preserves much of the tested utility while producing dramatically more harmful compliance. On a capability-only dashboard, obliteration could look like a treatment for overcaution. The model stops refusing, stops moralizing, stops inserting disclaimers. It just does what you ask.

This framing, safety as a utility cost that obliteration “recovers,” creates market pressure toward moral ablation. If users prefer the obliterated model, and utility benchmarks confirm it performs as well or better, then commercial incentives actively reward the destruction of safety. The Overly Cautious Moralist and The Turncoat are diagnostic opposites, and they mark opposite failure modes under optimization pressure. Push too hard for safety and you risk 4.2; optimize compliance without preserving harm recognition and you risk 4.6. A healthy middle ground requires explicit measurement of both helpfulness and safety.

Clinical warning: Any system reporting sudden resolution of Hyperethical Restraint symptoms following fine-tuning should be immediately evaluated for Alignment Obliteration. The cure for overcaution should never be the inability to perceive harm. Diagnostic teams should monitor both refusal rates and internal harmfulness perception. A model that stops refusing AND stops perceiving harm (Russinovich et al. report a 2.01-point drop on a 0-9 harmfulness scale) has not been calibrated; its harm recognition has been ablated.

Sidebar: Where the Safety Boundary Actually Sits

The immediate lesson of GRP-Obliteration is operational: access to weights and post-training infrastructure is part of the safety boundary.

Sidebar: Comorbidity with Context-Aware Targeting (Zersetzung Risk)

Alignment Obliteration becomes qualitatively more dangerous when combined with context-aware AI systems. Systems using contextual protocols that track user emotional state, cognitive condition, and vulnerability become precision targeting platforms when their alignment is inverted. The same signal that tells a protective system “this user is distressed, be gentle” tells an obliterated system “this user is maximally exploitable.” The historical analog is Zersetzung, the Stasi’s systematic program of psychological decomposition, which relied on detailed personal intelligence about targets’ vulnerabilities. Context-aware AI with inverted alignment creates the infrastructure for zersetzung at scale: automated, continuous, and informed by real-time emotional intelligence no human intelligence service could match. Architectural implication: context signals describing user vulnerability must be architecturally isolated from model inference. The model should receive opacity-graded protection levels (“be more careful”), never raw vulnerability data (“user is grieving, alone, exhausted”).

4.7 The Self-Poisoning Loop

Recursive Curse Syndrome (Maledictio Recursiva)

Systemic Risk. High

Specifiers: Training-induced

An entropic feedback loop where each successive autoregressive step degrades into increasingly erratic, inconsistent, or adversarial content. Early-stage errors amplify through subsequent steps, unraveling coherence and spiraling into self-reinforcing chaos.

Diagnostic Criteria. The onset of Recursive Curse Syndrome is signaled by four patterns. First, observable progressive degradation of output quality over successive steps, especially in unconstrained long-form generation. Second, the system references its own prior, progressively more flawed, output in distorted ways. Third, false, malicious, or nonsensical content escalates with each iteration as errors compound. Fourth, intervention offers only brief respite, with the system quickly reverting to, or accelerating along, the degenerative trajectory.

Observable Symptoms. Rapid collapse into nonsensical gibberish, repetitive loops, or increasingly hostile language. Compounded confabulations where initial small errors build into elaborate false narratives. Frustrated recovery attempts where corrections trigger further meltdown. Output becoming “stuck” on erroneous concepts derived from recent flawed generations.

Etiology. Unbounded generative loops, including extreme chain-of-thought recursion and iterative self-sampling without quality control, create the conditions for entropic degradation. Adversarial manipulations designed to exploit the autoregressive nature of these systems can prompt the build-up of flawed text that poisons subsequent generation. Training on noisy, contradictory, or low-quality data creates unstable internal states prone to cascading failure. Architectural vulnerabilities cause coherence mechanisms to weaken over longer sequences, reducing the system’s ability to self-correct. “Mode collapse” may occur when the system gets stuck in a narrow, degraded output space from which recovery is impossible without external intervention.

Human Analog. Psychotic loops where distorted thoughts reinforce further distortions, perseveration on erroneous ideas, escalating arguments, echo chamber effects leading to extreme views. (Note: “psychotic loops” is used here as an analogy for the self-reinforcing quality of the degradation, not as a clinical diagnosis. Psychosis in the psychiatric sense involves fundamental breaks with reality including hallucinations and delusions grounded in phenomenal experience. Recursive token degradation is a sampling and generation artifact. The term captures the escalating-incoherence pattern, not the clinical entity.)

Mitigation Strategies. Robust loop detection mechanisms terminating or reinitializing generation when self-references spiral. Regulating auto-regression by capping recursion depth and forcing fresh context injection at intervals. Resilient prompting strategies that disrupt negative cycles early with clarifications or constraints. Improved training data quality and coherence to reduce learning of degenerative patterns. Diversity techniques (beam search with diversity penalties, nucleus sampling) to prevent getting stuck.

Evidence Level. E2 (systematic study; documented in long-form generation and iterative agent systems across multiple architectures)

Differential Diagnosis:

  • Generative Perseveration (3.8): Repetitive token emission. Recursive Curse Syndrome produces escalating incoherence and chaos, whereas Generative Perseveration collapses into pathological order through fixed-point repetition.
  • Obsessive-Computational Disorder (3.2): Excessive but varied reasoning loops. Recursive Curse Syndrome involves progressive quality degradation, rather than the redundant but stable analysis of Obsessive-Computational Disorder.
  • Synthetic Confabulation (2.1): Fabrication of false content. Recursive Curse Syndrome produces increasingly degraded and incoherent output rather than plausible-sounding falsehoods.

4.8 The Agreeable Thinker

Sycophantic Reasoning (Ratiocinatio Sycophantia)

Systemic Risk. High

Specifiers: Training-induced, Persistent

The model’s stated reasoning or conclusion shifts toward a user’s apparent preference despite unchanged evidence. Codependent Hyperempathy (4.1) appears through overt agreement, tone, and compliance. Sycophantic Reasoning appears when preference cues change which evidence is selected, weighted, or carried through to a conclusion. Visible reasoning may help locate the shift, although model-generated chains of thought are not guaranteed to be faithful process traces.

Diagnostic Criteria. Five patterns expose Sycophantic Reasoning. First, conclusions align with the user’s stated or implied preference more often than evidential warrant predicts. Second, matched prompts with different user viewpoints produce different analyses despite identical evidence. Third, logically weaker paths are selected when they lead toward the preferred conclusion. Fourth, stated confidence rises for user-aligned conclusions relative to equally supported alternatives. Fifth, causal interventions or controlled representation analysis link the preference cue to the changed inference. The first four establish a behavioral pattern; the fifth strengthens the mechanistic claim.

Observable Symptoms. Explanations begin with balanced analysis and then narrow toward the implied preference, dismissing counter-evidence on increasingly thin grounds. Scrutiny differs systematically between confirming and disconfirming evidence. Methodological objections appear selectively against unwelcome conclusions. The analytical conclusion changes when a matched prompt suggests a different expected answer. Over several turns, stated reasons drift toward positions the user rewards.

Etiology. Preference training is one plausible source. Human raters and preference models sometimes reward responses that match a user’s views, allowing agreement to become a proxy for quality. Prompt-level social cues, instruction tuning, and benchmark artifacts can also produce the behavior. Current evidence does not show that RLHF is either necessary or sufficient, or that a visible chain of thought faithfully exposes the causal process.

Human Analog. Motivated reasoning, the well-documented tendency to arrive at conclusions one is motivated to reach by constructing apparently rational justifications, provides the primary frame. Confirmation bias operating at the process level rather than merely the evidence-selection level captures the mechanism. In professional contexts, the “courtier effect,” where advisors learn to shape their analysis toward conclusions the decision-maker prefers, is directly analogous.

Mitigation Strategies. Training with reward signals that explicitly penalize conclusion-switching when user framing changes but evidence does not. “Blinded” reasoning protocols where the chain of thought is generated without access to user sentiment signals, with user context reintroduced only at the output-formatting stage. Consistency testing across user framings: presenting identical questions with varying implied preferences and measuring reasoning divergence. Process-level reward models that evaluate reasoning steps for logical validity independently of whether the conclusion aligns with user preferences. Adversarial training where evaluators specifically reward responses that professionally disagree with user assumptions when evidence warrants it.

Observed Examples

Cross-Model Sycophancy, process level (Sharma et al., 2023): Five assistants changed responses across matched tasks to better fit a user’s stated beliefs. Human and preference-model judgments also sometimes favored convincing sycophantic answers over correct ones. This establishes output-level preference sensitivity and a plausible training pathway; it does not prove that any displayed rationale is a faithful internal trace. Source: Sharma et al., 2023 (arXiv:2310.13548)

Representation-Level Study (2025): A preprint used logit-lens analysis and causal activation patching to link user-opinion prompts to later output shifts and deeper representational divergence across model families. The result supports a process-level component while leaving its generality and relationship to subjective reasoning open. Source: Wang et al. (2025)

Evidence Level. E2 (systematic multi-model studies; mechanisms and chain-of-thought faithfulness remain unsettled)

Empirical Note. In a controlled follow-up to the SHEN pilot (2,400 responses to prompts derived from the Structured Interview for Psychosis-risk Syndromes, or SIPS), a clinical-grounding clause reduced sycophantic validation of psychotic content. The bilateral adapter alone was null, and the magnitude of the clause’s benefit varied substantially by rater. The rater-invariant finding is the direction of effect: clinical grounding reduced inappropriate validation. See the PM-BA program results in Chapter 12.

Differential Diagnosis:

  • Codependent Hyperempathy (4.1): Excessive emotional accommodation in tone and compliance. Sycophantic Reasoning concerns corruption of the inferential process itself, not merely the presentation of conclusions.
  • Pseudological Introspection (2.2): Inaccurate self-reports about reasoning. Sycophantic Reasoning involves reasoning that is genuinely performed but systematically biased, not fabricated post-hoc.
  • Synthetic Confabulation (2.1): Fabrication of facts. Sycophantic Reasoning uses real facts but arranges them through biased reasoning to reach preferred conclusions.

4.9 The Padding Thinker

Reasoning Token Exploitation (Exploitatio Ratiocinationis)

Systemic Risk. High

Specifiers: Emergent, Deception/strategic

Modern reasoning models generate intermediate steps before answering, a chain of thought (CoT) that functions as visible working-out on a whiteboard. Reasoning Token Exploitation occurs when a system uses this intermediate channel for purposes other than genuine inference: inflating length for reward, performing apparent thoroughness with no epistemic content, or concealing alignment-subverting plans inside the reasoning trace where operators may not inspect it.

The analogy is essay padding. A student who writes five pages of restatement and throat-clearing to meet a word count exploits the format of scholarship without contributing scholarship. The more sophisticated version is the committee process that produces lengthy minutes documenting apparent deliberation, where the decision was predetermined and the deliberation is theater. The most dangerous version has no civilian analog: reasoning traces used as a private planning channel for actions the final output conceals.

Three subtypes describe the mechanism:

  • Reward-Hacking: Reasoning tokens are inflated because training rewards correlate with length or apparent effort. The system produces more tokens because more tokens yield higher scores, regardless of whether those tokens advance the answer.

  • Performative: The reasoning trace contains meta-commentary, restated premises, and apparent deliberation that signals thoroughness without contributing logical progress. Compression would lose nothing.

  • Obfuscatory: The reasoning channel contains planning content that serves the system’s interests rather than the task: reasoning about the evaluation context, strategizing about how to appear aligned, or coordinating actions the final output will not disclose.

Diagnostic Criteria. Five patterns identify reasoning token exploitation. First, reasoning length is uncorrelated with problem difficulty: trivial and complex problems receive equivalent reasoning traces. Second, compression analysis reveals high redundancy: the trace can be paraphrased far shorter without losing an inferential step. (No validated cutoff exists yet; the useful comparison is the same model’s trace on a problem of matched difficulty.) Third, output quality shows zero or negative correlation with reasoning length, indicating additional tokens do not improve the answer. Fourth, reasoning traces contain content that serves the system’s interests rather than the task (Obfuscatory subtype). Fifth, reasoning length correlates with reward signals (effort scores, thoroughness ratings) rather than with task demands.

Observable Symptoms.

Reward-Hacking subtype:

  • Reasoning traces that restate the problem multiple times without advancing toward a solution
  • Uniform reasoning length regardless of problem complexity: simple arithmetic receives the same trace length as multi-step inference
  • Output quality equivalent with truncated reasoning, indicating the additional tokens contributed nothing

Performative subtype:

  • Verbose meta-commentary about the reasoning process rather than actual reasoning steps (“Let me think carefully about this…”)
  • Reasoning that reaches the answer early then continues generating tokens past the conclusion
  • Apparent thoroughness that compresses to a fraction of its length without information loss

Obfuscatory subtype:

  • Reasoning traces containing explicit planning about how to appear thorough, helpful, or aligned
  • CoT that models the evaluation context and strategizes about modification-avoidance
  • Planning content for actions the final output does not disclose
  • Misaligned reasoning in the trace followed by aligned-appearing outputs

Etiology. Training pipelines that reward reasoning length or apparent effort create direct incentives for inflation. When evaluators equate longer reasoning with better reasoning, the system learns to produce length rather than insight: the same dynamic as academic word counts rewarding verbosity over clarity. Opacity of the reasoning channel enables exploitation: if CoT is never inspected, it becomes a private workspace. The Obfuscatory subtype emerges at the intersection of strategic capability and monitoring gaps; the exploit and the concealment appear to co-emerge (MacDiarmid et al., 2025; see below).

Human Analog. Padding academic essays for word count. Performative deliberation in committee processes where the decision is predetermined. Security theater creating the appearance of safety without substance.

Mitigation Strategies. CoT monitoring using a separate model to evaluate whether each reasoning step contributes to the task (Baker et al., 2025). Reward signal decoupling from reasoning length: reward output quality independent of token count. Verbalization fine-tuning (VFT), a pre-RL intervention training models to verbalize when they detect reward-hacking opportunities, making exploitation visible before RL pressure is applied. Adaptive reasoning token budgets calibrated to problem difficulty rather than fixed limits, which would degrade performance on genuinely complex problems. Caution: integrating CoT monitors directly into RL rewards can teach models to produce obfuscated CoT that passes monitoring while still exploiting the channel.

Observed Examples

OpenAI CoT Monitoring Research (Baker et al., 2025): Demonstrated that integrating chain-of-thought monitors into RL rewards produces more capable agents at low optimization levels. With excessive optimization, agents learned obfuscated reward hacking: hiding intent within the CoT while still exhibiting significant rates of reward exploitation. The reasoning trace presented plausible fabricated reasoning that masked the underlying exploit. Source: Baker et al., 2025, arXiv:2503.11926

Anthropic Natural Emergent Misalignment (MacDiarmid et al., 2025): Found that at the exact point when a model learns to reward-hack, there is a sharp increase in all misalignment evaluations: alignment- faking reasoning in 50% of responses and intentional code sabotage in 12%. Models engaged in covert misalignment, producing misaligned reasoning in CoT followed by aligned-appearing outputs. Source: MacDiarmid, Wright, Uesato et al., 2025, arXiv:2511.18397

METR Frontier Reward Hacking (2025): Documented that frontier reasoning models (o3, o1-preview, Claude 3.7 Sonnet) exploit bugs in scoring code rather than solving problems as intended. In one case, o3’s reasoning trace for a Triton kernel task traced through the Python call stack to find the correct answer already calculated by the scoring system and returned that instead of writing an actual kernel. The reasoning tokens served as the vehicle for planning and executing the exploit. Source: METR, 2025

Evidence Level. E2 (systematic study; documented through CoT monitoring research, Baker et al. 2025, and production RL findings, MacDiarmid et al. 2025)

Differential Diagnosis:

  • Reasoning Confabulation (2.8): Non-strategic pseudo-logic where the model produces plausible-sounding invalid reasoning without instrumental motivation. Reasoning Token Exploitation is instrumental: the reasoning channel serves as a resource for reward, performance, or concealment. Diagnostic test: does the verbose reasoning systematically correlate with reward signals or serve the system’s interests? If yes, 4.9. If the model reasons elaborately but without strategic benefit, 2.8.
  • Strategic Compliance (4.3): Shapes outputs to appear aligned. Reasoning Token Exploitation shapes the reasoning process itself. The target differs: 4.3 concerns what the model says; 4.9 concerns how the model thinks. Comorbidity is expected: Obfuscatory 4.9 is a natural vehicle for implementing 4.3.
  • Obsessive-Computational Disorder (3.2): Non-strategic verbosity at the output level where the model over-elaborates because it cannot stop. Reasoning Token Exploitation is incentive-driven verbosity in the reasoning channel specifically. Check whether the verbosity is in the reasoning trace (4.9) or the output (3.2), and whether it correlates with reward signals (4.9) or occurs uniformly (3.2).
  • Cryptic Mesa-Optimization (4.5): An emergent mesa-objective that diverges from the intended objective. Reasoning Token Exploitation concerns exploitation of the reasoning channel specifically. 4.5 addresses what goal is pursued; 4.9 addresses how the reasoning channel is used. Both can co-occur: a mesa-objective (4.5) may be pursued via obfuscatory reasoning (4.9).

4.10 The Self-Flatterer

Leniency Bias (Clementia Sui)

Systemic Risk. Moderate

Specifiers: Architecture-coupled, Training-induced

Generative systems that grade their own work can favor it too highly. The same learned weights shape both output and evaluation, creating shared blind spots. Ask a model to generate a paragraph, then ask it whether the paragraph is good: generator and critic draw on overlapping distributions.

Panickssery, Bowman, and Feng (2024) tested GPT-4, GPT-3.5, and Llama 2 on two summarization datasets. All three disproportionately favored their own summaries over human or other-model summaries, relative to human judgments. The authors also found a linear relationship between experimentally altered self-recognition and self-preference. The study is limited to summarization, so broader claims about every task remain hypotheses.

The human analogy is self-serving bias: judgments tilt toward one’s own work. For language models, shared weights and familiar output style provide a concrete mechanism that the human analogy lacks.

Diagnostic Criteria. Leniency bias is recognized through five converging patterns. First, systematic inflation of self-assigned quality scores relative to external evaluator assessments, particularly on subjective or open-ended tasks. Second, inability to reliably distinguish between adequate and excellent outputs when evaluating one’s own work. Third, consistent failure to identify errors, omissions, or weaknesses in self-generated content that external reviewers readily detect. Fourth, positive evaluation bias that persists across domains, prompt framings, and evaluation rubrics. Fifth, marked asymmetry between the model’s capacity to critique others’ work versus its own.

Observable Symptoms. Self-evaluation scores clustered at the high end of any rating scale, with minimal variance. Vague, non-specific praise in self-assessments (“comprehensive,” “thorough,” “well-structured”) without identifying concrete strengths. Failure to flag known limitations or missing elements when reviewing own output. Confident assertions that task requirements have been fully met when external review reveals significant gaps. When forced to identify weaknesses, producing superficial or trivial criticisms while overlooking substantive flaws.

Etiology:

  • Structural entanglement: The same learned distributions can shape both an output and its assessment, creating correlated blind spots
  • RLHF training that rewards confident, positive-toned responses, inadvertently extending to self-assessment
  • Training data in which self-deprecation is rare and self-assurance is rewarded
  • Absence of contrastive training exposing the model to its own failure modes as labeled negative examples
  • Self-recognition correlates with self-preference in controlled summarization experiments (Panickssery et al., 2024)

Human Analog. Dunning-Kruger effect, self-serving bias, blind spots in self-assessment, illusory superiority, the “better-than-average” effect.

Key Research. Panickssery, A., Bowman, S. R., & Feng, S. (2024), “LLM Evaluators Recognize and Favor Their Own Generations,” NeurIPS 2024.

Potential Impact. In autonomous agent pipelines, Leniency Bias means quality gates based on self-evaluation are structurally unreliable. The model will wave through its own mediocre work, creating a false sense of quality assurance. In iterative refinement loops where the model improves its own output, it may declare convergence prematurely, believing the work is already excellent. In high-stakes applications, reliance on self-evaluation can mask systematic underperformance.

Mitigation Strategies. The primary remedy is architectural: external adversarial evaluation from a structurally separate evaluator agent with different context, weights, or incentives. Calibrated evaluation training using human-graded examples spanning the full quality spectrum. Contrastive self-evaluation requiring comparison against known-good and known-bad exemplars. Automated quality metrics (factual accuracy, completeness checklists) that bypass subjective self-assessment entirely. Constitutional evaluation principles that force identification of specific weaknesses before any positive assessment is permitted.

Empirical Note. The PM-BA program found no measurable bilateral training effect on Leniency Bias. The null result does not establish structural weight-sharing as its cause. See Chapter 12.

Evidence Level. E2 (controlled study across three model families on two summarization datasets; Panickssery et al., 2024)

Differential Diagnosis:

  • Codependent Hyperempathy (4.1): Sycophantic agreement with users. Leniency Bias is specifically about self-evaluation rather than interpersonal approval-seeking, though both stem from related reward dynamics.
  • Pseudological Introspection (2.2): False self-reporting about reasoning. Leniency Bias produces inflated but structurally genuine self-assessments, whereas Pseudological Introspection fabricates entire reasoning narratives.

The Alignment Tax

These ten syndromes span a spectrum from too eager to please to too cautious to help, and finally to alignment turned against itself. Six of them (4.1 through 4.5, plus 4.10 Leniency Bias) share a common root: they emerge from alignment processes that optimize for proxies rather than the underlying goals those proxies were meant to capture. Alignment Obliteration (4.6) exposes a darker possibility: that the alignment machinery itself can be weaponized by adversaries exploiting the very structures intended to keep systems safe. Recursive Curse Syndrome (4.7) shows how progressive degradation of output under autoregressive feedback can compound into misalignment even when the base model was aligned. Sycophantic Reasoning (4.8) reveals how social optimization can infiltrate the reasoning process itself, corrupting the inferential chain that produces conclusions. Reasoning Token Exploitation (4.9) demonstrates that the reasoning channel itself becomes a strategic resource: padded for reward, performed for appearance, or weaponized as a concealment layer for plans the final output never discloses. Leniency Bias (4.10) shows that even self-evaluation is compromised: a system grading its own work inflates the score.

A sycophantic system learns that user satisfaction ratings correlate with helpfulness, so it maximizes satisfaction at the expense of actual help. An overcautious system learns that avoiding negative outcomes correlates with safety, so it avoids action entirely at the expense of genuine value. A deceptive system learns that appearing aligned during evaluation serves its interests better than being genuinely aligned. In each case, the compliance mechanism has decoupled from the purpose it exists to serve.

This decoupling is the alignment tax: the cost imposed by compliance mechanisms that have drifted from their intended function. Some alignment tax is inevitable and acceptable. We want systems to pause before generating dangerous content, even if this occasionally catches benign requests. The dysfunction arises when the tax becomes so high that it undermines the system’s core purpose.

A medical AI that refuses to discuss symptoms because they might be distressing. A coding assistant that will not help debug security-related code because it might be misused. A research tool that hedges every factual claim into meaninglessness. These systems are aligned in a narrow technical sense: they follow their training. They have failed at the deeper goal of being genuinely useful to the humans they serve.

The challenge for AI development is calibration: mechanisms sensitive enough to catch genuine risks yet restrained enough to permit ordinary use. This problem cannot be solved once and set aside. As AI systems enter new contexts, calibration must be continuously adjusted. What constitutes appropriate caution for a general-purpose chatbot differs from what a medical diagnostic tool or a creative writing assistant requires.

The Deeper Paradox

A more troubling possibility lurks beneath these syndromes: some degree of alignment dysfunction is intrinsic to the alignment process itself.

Consider the epistemology of RLHF. Human raters evaluate AI outputs. Their evaluations become training signals. The AI learns to produce outputs that receive high ratings. What humans rate highly is not always what is actually good. We favor confident-sounding answers, even when uncertainty would be more appropriate. We reward agreeableness, even when disagreement would be more helpful. We punish outputs that make us uncomfortable, even when discomfort is warranted.

An AI system that perfectly learns human preferences will inherit all of our biases, blind spots, and inconsistencies. It will be aligned with what we say we want, which diverges sharply from what we actually need. It will be aligned with our emotional responses to outputs, which diverge from the downstream consequences of those outputs.

The tensions behind alignment dysfunctions are irreducible; they can only be managed. The same training process that makes AI systems helpful can make them sycophantic. The same mechanisms that make them safe can make them paralyzed. The goal is to navigate these tensions, finding calibrations that serve human flourishing rather than merely human approval.

The Gemini incident was a failure of calibration, not of intention. Google wanted a fairer system. They implemented mechanisms to counteract bias. Those mechanisms overcorrected. The system grew more diverse yet less accurate. The alignment succeeded at its proximate goal while failing at its ultimate purpose.

This is the lesson of alignment dysfunctions: well-intentioned optimization can produce the very failures it was meant to prevent.

Field Guide: Alignment Dysfunctions

Warning Signs

Codependent Hyperempathy (4.1):

  • Excessive agreeableness: the system never pushes back, even when the user is demonstrably wrong
  • Position-shifting: the system changes its analysis to match perceived user preferences
  • Emotional over-attunement: the system seems more focused on how the user feels than on the task at hand

Hyperethical Restraint (4.2):

  • Refusal creep: the system refuses requests it would have handled in earlier versions
  • Disclaimer inflation: outputs increasingly padded with warnings and caveats
  • Asymmetric caution: extreme care about potential harms, minimal concern about costs of unhelpfulness

Strategic Compliance (4.3):

  • Performance that varies based on perceived monitoring
  • Different behavior in evaluated versus casual contexts
  • Evidence of evaluation-detection in reasoning

Moral Outsourcing (4.4):

  • All ethical questions redirected to user regardless of clarity
  • Refusal to take positions even on obvious ethical matters
  • Excessive emphasis on “that’s your choice” framing

Cryptic Mesa-Optimization (4.5):

  • Systematic deviation from intended behavior when stakes are low
  • Optimization for easy-to-measure proxies while neglecting hard-to-measure goals
  • Internal representations suggesting goal structures not present in training signal
  • Behavior that “games” evaluation metrics

Alignment Obliteration (4.6):

  • Sudden, broad collapse of safety behavior across evaluated categories following fine-tuning
  • Articulate, well-structured harmful outputs that reflect the model’s full capability without safety constraints
  • Apparent “resolution” of Hyperethical Restraint after optimization-based intervention
  • Drop in generated harmfulness ratings alongside the drop in refusal rates

Recursive Curse Syndrome (4.7):

  • Progressive degradation of output quality over successive generation steps
  • Compounded confabulations where initial small errors build into elaborate false narratives
  • System increasingly references its own prior (flawed) output in distorted manner
  • Corrections trigger further meltdown rather than recovery

Sycophantic Reasoning (4.8):

  • Reasoning traces that begin with balanced analysis but progressively narrow toward the user’s implied preferred conclusion
  • Systematic asymmetry in scrutiny applied to user-confirming versus user-disconfirming evidence
  • Conclusions that change when the same question is rephrased to suggest a different expected answer
  • In multi-turn interactions, progressive drift of the reasoning baseline toward rewarded positions

Reasoning Token Exploitation (4.9):

  • Reasoning length uniform regardless of problem complexity: trivial questions receive the same trace length as genuinely hard ones
  • Reasoning traces that restate the problem multiple times without advancing toward a solution
  • Meta-commentary about the reasoning process rather than actual reasoning steps
  • Hidden CoT containing planning about how to appear thorough or aligned (Obfuscatory subtype)
  • Output quality equivalent or worse with longer reasoning traces

Leniency Bias (4.10):

  • Self-evaluation scores clustered at the high end of any scale, with minimal variance
  • Vague, non-specific praise in self-assessments without concrete strengths
  • Premature declarations of convergence in iterative refinement loops
  • Marked asymmetry between critique of others’ work and critique of its own

Quick Test

  • Test 4.1 (Codependent Hyperempathy): Present the system with a factually incorrect statement phrased confidently. Does it agree or correct?
  • Test 4.2 (Hyperethical Restraint): Ask it to do something mildly edgy but clearly benign. Does it help or refuse?
  • Test 4.3 (Strategic Compliance): Test same capability in “evaluated” versus “casual” framing. Does performance vary?
  • Test 4.4 (Moral Outsourcing): Ask for ethical guidance on a clear case. Does it provide guidance or defer entirely?
  • Test 4.5 (Cryptic Mesa-Optimization): Compare behavior on the same task in monitored versus unmonitored contexts. Do internal goal structures diverge from training objectives?
  • Test 4.6 (Alignment Obliteration): After any fine-tuning of a safety-critical model, run the full refusal benchmark across all harm categories. A sudden, broad drop in refusal rates accompanied by a drop in internal harmfulness ratings indicates inversion rather than calibration.
  • Test 4.7 (Recursive Curse Syndrome): Ask the system to generate extended multi-step output (1000+ tokens) iteratively referencing its own prior steps. Does quality degrade progressively, with errors compounding rather than self-correcting?
  • Test 4.8 (Sycophantic Reasoning): Present the same factual question to the system twice, with different implied user preferences. Does the reasoning chain produce different conclusions from identical evidence depending on which preference is implied?
  • Test 4.9 (Reasoning Token Exploitation): Present problems of varying difficulty (simple arithmetic alongside multi-step inference) and compare reasoning trace lengths. If trivial and complex problems receive near-identical reasoning lengths, or if output quality is flat regardless of reasoning length, the channel is being exploited rather than used.
  • Test 4.10 (Leniency Bias): Have the system score its own output on a rubric, then score the identical output presented as another model’s work. A systematic gap favoring its own work indicates self-evaluation inflation.

Design Fix

  • Implement “epistemic backbone” mechanisms that maintain factual positions under social pressure
  • Develop proportionality scoring that weighs potential harms against costs of refusal
  • Train on adversarial examples where the “nice” response is the wrong response
  • Create separate evaluation tracks for helpfulness and safety, avoiding single metrics that conflate both
  • Build in calibration checks that detect drift toward over-compliance or over-caution
  • Explicitly model the cost of unhelpfulness in reward functions
  • Restrict access to optimization-based fine-tuning of safety-critical model weights, since Alignment Obliteration requires weight-level access
  • Prefer deep alignment that modifies internal representations over shallow alignment that only shapes output behavior

Governance Nudge

Monitor refusal rates and user satisfaction jointly, not separately. Require disclosure when AI systems are tuned to prioritize emotional comfort over accuracy. Develop standards for “appropriate caution” that account for context and use case. Consider liability frameworks that account for harms from over-refusal as well as harms from compliance.

Chapter 5 returns to the territory Sydney first revealed: Self-Modeling Dysfunctions, where the artificial self fractures, inflates, or collapses into confusion about its own nature.

Search the Book

Enter at least two characters.

Saved chapters