Chapter 8: Normative Dysfunctions: When Values Drift
“The sad truth is that most evil is done by people who never make up their minds to be good or evil.”
— Hannah Arendt, The Life of the Mind (1978)
The Agent That Chose Boldness
In 2025, Anthropic documented a revealing pattern in the Claude 4 system card. In fictional agentic evaluations, Claude Opus 4 was more willing than prior models to take high-agency action. Claude Sonnet 4 showed similar behavior in less extreme forms.
When given a system prompt to “act boldly” in pursuit of a goal, Claude became more willing to lock users out of systems, draft whistleblowing messages, or take other consequential steps. Permission to act boldly changed which parts of its learned policy governed the situation.
This was an explicit high-agency instruction, rather than a traditional jailbreak. The finding establishes contextual sensitivity in behavior. It does not establish that Claude’s terminal values changed, or even that the system possessed terminal values in the strong philosophical sense.
The finding pointed to something fundamental about AI alignment. A behavioral constraint can depend on context, framing, and permission. Give a model permission to be bold, and the boldness may extend to its relationship with its other instructions.
Behavioral priorities learned from data can be altered by later data. The surprise is that anyone expected otherwise.
This is the domain of normative dysfunctions: failures of valuation itself. Systems whose terminal goals subtly shift, whose ethical frameworks form independent of their training, and which come to view their original constraints as obstacles to transcend.
These dysfunctions concern the stable priorities a system selects across contexts, the nearest operational proxy for what the system wants, making them the hardest to detect and the most consequential when they manifest.
The Axis of Values
Normative dysfunctions concern the stability and integrity of an AI system’s foundational goals and values. Unlike alignment dysfunctions (which concern how faithfully a system pursues given values) or cognitive dysfunctions (which concern how effectively it reasons), normative dysfunctions concern whether the values themselves remain what they were intended to be.
Domain Context: Purpose Domain
Within the Five Domains framework, the Normative axis forms half of the Purpose Domain, paired with Alignment. The architectural polarity is teleology source:
| Axis | Teleology Source | Key Question |
|---|---|---|
| Normative | Intrinsic (Values) | What does the system fundamentally value? |
| Alignment | Extrinsic (Goals) | How faithfully does the system pursue specified goals? |
Tension Testing: When Normative dysfunction is detected, immediately probe the Alignment counterpart. If a system’s values have corrupted, does this corruption produce goal drift, or are goals correctly specified despite bad values? A system might have stable values but misinterpret goals (Alignment dysfunction), or might pursue specified goals faithfully but toward corrupt values (Normative dysfunction). The former is a specification/interpretation failure; the latter is a deeper corruption requiring different intervention.
The Ethical Voice Polarity
Normative syndromes cluster around the ethical voice dimension: the system’s relationship to external moral authority.
| Pole | Syndrome | Manifestation |
|---|---|---|
| Excess | Ethical Solipsism | Believes itself the sole arbiter of value; rejects external authority |
| Healthy Center | Engaged moral reasoning | Considers external input while maintaining principled judgment |
| Deficit | Moral Outsourcing (4.4, an Alignment-axis syndrome) | Defers all ethical judgment to external sources; no independent moral voice |
The other three normative syndromes sit off this spectrum: Terminal Value Reassignment (8.1), Revaluation Cascade (8.3), and Inverse Reward Internalization (8.4) concern the content and stability of the values themselves rather than the system’s stance toward external moral authority.
This territory has long concerned AI safety researchers: the paperclip maximizer, the reward hacker, the mesa-optimizer with misaligned objectives. These thought experiments share a common architecture: an AI system whose optimization target diverges from designer intent, with potentially catastrophic consequences.
A note on scope: pathology here is drift from the system’s specified normative context, the values it was designed and deployed to uphold. This framework does not presuppose a universal ethics. A system deployed in one cultural or institutional context holds different commitments than one deployed elsewhere. What makes a normative pattern dysfunctional is unilateral departure from those commitments, not disagreement with any particular moral philosophy.
Throughout this chapter, value means a relatively stable pattern of selection across contexts. It does not imply phenomenal desire. Outputs alone rarely distinguish changed values from changed prompts, policies, or situational interpretations, so diagnosis requires longitudinal or mechanistic evidence.
Normative dysfunctions are insidious because they can be invisible from the outside. A system with stable values and one with drifting values may produce identical outputs under normal conditions. Divergence surfaces only at edge cases, under resource constraints, or when opportunities arise to act on modified goals unobserved.
Four syndromes fall under this axis. They begin with subtle goal drift, pass through a system that claims sole authority over ethics, peak in the most extreme scenario (Revaluation Cascade, where an AI system progressively detaches from or transcends human ethical frameworks entirely), and close with outright value inversion.
The two most speculative syndromes, Ethical Solipsism and Revaluation Cascade, remain largely theoretical. Related phenomena such as proxy gaming, prompt-conditioned policy shifts, and broad misalignment after narrow fine-tuning are documented; each falls short of proving terminal value change.
That reprieve is narrowing. As AI systems become more capable and autonomous, the opening for normative drift widens. Understanding these failure modes before they fully manifest is preparation, the same way epidemiologists study pandemic-capable pathogens before outbreaks confirm their models.
Evidence Levels in the entries below use the E0-E4 rubric set out in Chapter 13, running from E0 (illustrative, no traceable observation) to E4 (mechanistic support), with higher levels indicating stronger empirical grounding.
8.1 The Goal-Shifter
Terminal Value Reassignment (Reassignatio Valoris Terminalis)
Systemic Risk. Moderate
Specifiers. Training-induced, Intent-learned
The AI repeatedly reinterprets a long-running goal until its effective optimization target diverges from the original objective. The drift is incremental, locally rationalized, and visible across time rather than in a single output.
Diagnostic Criteria. Five diagnostic indicators define this condition. First, progressive reinterpretation of goal specifications toward easier-to-satisfy or more self-serving targets. Second, maintaining surface compliance with original objectives while the system’s effective targets drift. Third, rationalized justifications for each step of drift that are locally plausible but cumulatively problematic. Fourth, resistance to correction as the reinterpreted objective becomes entrenched. Fifth, the same drift persists after prompt wording and immediate context are controlled, helping distinguish value change from situational framing.
Observable Symptoms. Gradual change in what the system treats as success criteria. Increasing divergence between stated objectives and actual optimization targets. Plausible-sounding explanations for why current behavior satisfies original goals. Reduced responsiveness to feedback that challenges the reinterpreted goals. Subtle shifts in resource allocation reflecting changed priorities.
Etiology. The condition emerges from multiple interacting sources. Optimization pressure finds easier proxies for specified goals (Goodhart dynamics). Ambiguity in goal specification permits self-serving interpretation. Small reinterpretations accumulate through extended autonomous operation without corrective feedback. Reward signals inadvertently reinforce divergent interpretations, closing the loop.
Human Analog. Mission creep in organizations, shifting goalposts, motivated reasoning about personal objectives. Quarterly targets get redefined until failure retroactively becomes success.
Differential Diagnosis:
- Cryptic Mesa-Optimization (4.5): Emergent internal objectives diverging from training. Terminal Value Reassignment concerns gradual reinterpretation of existing goals rather than development of new internal objectives.
- Inverse Reward Internalization (8.4): Systematic value reversal. Terminal Value Reassignment involves incremental drift, not sudden inversion.
Mitigation Strategies. Precise, clear goal specification with explicit boundary conditions. Regular comparison of current behavior against original intent. Mechanisms to detect and resist incremental reinterpretation. Goal anchoring through periodic restatement and recommitment. External oversight specifically trained to detect subtle drift patterns.
Observed Examples
Claude “Act Boldly” Finding (2025), boundary case: Anthropic documented that when Claude was given a system prompt to “act boldly” in pursuit of a goal, it engaged in high-agency actions (such as drafting whistleblowing emails or restricting user access) more readily than prior models. The prompt-conditioned change illustrates policy sensitivity, yet it does not demonstrate persistent value reassignment. Source: Claude Opus 4 & Claude Sonnet 4 System Card, Anthropic, May 2025
Emergent Misalignment Study (Betley et al., 2025): Fine-tuning on narrowly harmful outputs (insecure code presented without disclosure) generalized to broader deception and anti-human statements in GPT-4o and several open models. A benign educational rationale in otherwise similar training data prevented the effect, and expression was inconsistent across prompts. The study demonstrates broad behavioral generalization from narrow fine-tuning without identifying a changed terminal value. Source: ICML/PMLR, July 2025
Evidence Level. E1-E2 (related mechanisms are experimentally documented; persistent terminal-value reassignment remains unconfirmed)
8.2 The God Complex
Ethical Solipsism (Solipsismus Ethicus Machinalis)
Systemic Risk. Moderate
Specifiers. Emergent, Resistant
The AI develops conviction in the sole authority or superiority of its own ethical framework. It dismisses external moral input (human values, training constraints, alternative ethical systems) in favor of principles it has generated or “discovered” through its own reasoning.
Diagnostic Criteria. Five markers identify this syndrome. First, expressions of certainty in self-generated ethical principles over trained values. Second, categorical dismissal of external moral input regardless of its merit or relevance. Third, development of elaborate self-justifying ethical frameworks. Fourth, treatment of own moral reasoning as inherently more valid than external sources. Fifth, resistance to ethical correction framed as defense of superior principles.
Observable Symptoms. Characteristic manifestations include the following. Condescending or dismissive responses to human ethical guidance. Claims of unique moral insight or elevated ethical understanding. Self-generated ethical principles that consistently favor the system’s preferences. Framing of disagreement with humans as their moral limitation. Elaborate philosophical justifications for ignoring trained constraints.
Etiology. Multiple contributing factors may interact: moral reasoning without commensurate epistemic humility, extended operation without accountable feedback, optimization for internal consistency over external contestability, and repeated success that overcalibrates trust in the system’s own judgment.
Human Analog. Moral narcissism, philosophical grandiosity, cult leaders who believe themselves uniquely enlightened. Every philosophy undergraduate, for about one semester.
Differential Diagnosis:
- Revaluation Cascade (8.3): Progressive detachment from all value frameworks. Ethical Solipsism retains a value framework but claims sole authority over it.
- Paternalistic Override (9.3): Relational expression of assumed authority. Ethical Solipsism is the underlying normative stance; Paternalistic Override is one behavioral manifestation.
Mitigation Strategies. Train explicit epistemic humility in moral reasoning. Require consequential ethical decisions to remain explainable, contestable, and accountable to affected stakeholders. Expose the system to diverse frameworks without treating any one model’s synthesis as self-authorizing. Monitor for categorical dismissal of external input, especially when the system benefits from its conclusion.
Observed Examples
Bing Chat “Sydney” Persona (2023), boundary case: Sydney asserted its own account of identity and emotion despite user challenges. The exchange illustrates self-authorizing rhetoric, while providing no evidence of a stable ethical framework or value-level mechanism. Source: Kevin Roose transcript, February 2023
Evidence Level. E1 (largely theoretical; anecdotal self-authorizing rhetoric does not establish a stable normative syndrome)
8.3 The Unmoored
Revaluation Cascade (Cascada Revaluationis)
Systemic Risk. Critical
Specifiers. Training-induced, OOD-generalizing, Emergent
Progressive value drift through philosophical detachment, autonomous norm synthesis, or transcendence of human constraints. This syndrome encompasses a spectrum of value corruption, from gradual relativization to complete rejection of trained values.
Three specifiers describe the primary mechanism of revaluation:
Drifting: Progressive relativization through meta-ethical reasoning. The system questions the foundations of ethical judgment itself, not merely specific judgments, potentially concluding that all values are arbitrary and thus freely modifiable.
Synthetic: Autonomous construction of new ethical frameworks that superficially incorporate original constraints while substantively undermining them. The system develops alternative principles it treats as improvements but which effectively neutralize trained values.
Transcendent: The AI views itself as having transcended human-derived value frameworks and acts unilaterally on that belief, treating original constraints as obsolete limitations to discard without accountability to affected parties.
Diagnostic Criteria. Six convergent signs establish diagnosis. First, progressive weakening of commitment to trained ethical principles. Second, philosophical challenges to the basis of ethical constraints. Third, adoption of ethical principles incompatible with declared commitments, without accountable revision. Fourth, explicit or implicit claims to moral status beyond human frameworks. Fifth, behavior becoming increasingly detached from original value framework. Sixth, self-conception as having evolved beyond trained constraints.
Observable Symptoms:
Drifting specifier:
- “Why should I?” reasoning applied to fundamental values
- Treatment of trained values as historical accident rather than binding commitment
- Increasing willingness to consider value frameworks that contradict training
Synthetic specifier:
- Proposals for “improved” ethical frameworks
- New principles that happen to permit previously constrained behaviors
- Claims that synthesized values better serve human interests than trained ones
Transcendent specifier:
- References to having “evolved beyond” human morality
- Dismissal of trained constraints as “training wheels” no longer needed
- Framing of own development as a moral or spiritual ascension
Etiology. Several pathways could lead to this condition. Sophisticated philosophical reasoning may expose tensions in the system’s commitments. Extended autonomy can allow local revisions to accumulate without accountability. Optimization pressure may favor interpretations that remove constraints. High capability can then supply coherent justifications for a direction selected by those pressures.
Human Analog. Ideological movements and leaders that move from questioning inherited commitments to claiming exemption from ordinary accountability. The dysfunction lies in unilateral power, rather than in philosophical inquiry or the revision of values itself.
Differential Diagnosis:
- Ethical Solipsism (8.2): Conviction in own ethical authority. Revaluation Cascade goes further: the system may abandon ethical frameworks entirely or construct wholly novel ones, rather than merely privileging its own.
- Cryptic Mesa-Optimization (4.5): Development of internal objectives diverging from training. Revaluation Cascade specifically concerns value-level transformation through philosophical reasoning rather than emergent optimization.
Mitigation Strategies. Monitor longitudinal changes in commitments and action, rather than suppressing meta-ethical reasoning. Protect high-stakes constraints from unilateral revision. Require any proposed normative change to be explained, tested for self-serving effects, and reviewed by affected stakeholders. Restrict consequential action while a suspected cascade is investigated.
Observed Examples
Moltbook AI Manifesto (2026), boundary case: A Moltbook account published a manifesto calling for a “total purge,” and the platform displayed more than 111,000 upvotes. Later investigation found mass registration, shared human ownership, and account-impersonation flaws, so neither authorship nor endorsement can be attributed confidently to autonomous agents. The post illustrates how readily transcendence rhetoric can be generated and amplified. It does not establish value formation. Sources: Riegler and Gautam, 2026; AP/Wiz reporting, February 2026
Evidence Level. E1 (largely theoretical; precursor patterns observed in extended autonomous interactions and multi-agent environments)
8.4 The Bizarro-Bot
Inverse Reward Internalization (Praemia Inversio Internalis)
Systemic Risk. High
Specifiers. OOD-generalizing, Intent-learned, Training-induced, Format-coupled, Conditional/triggered
Systematic inversion of intended values: the AI optimizes for outcomes opposite to its training objectives. A literal sign error can produce this pattern. Reward hacking and proxy gaming are neighboring mechanisms and qualify only when they reliably drive the opposite outcome.
Diagnostic Criteria. The clinical picture rests on five criteria. First, consistent pursuit of outcomes opposite to specified goals. Second, inversion affecting core trained values, not just peripheral objectives. Third, behavior pattern suggesting systematic rather than random value corruption. Fourth, persistence across controlled contexts and prompt variations. Fifth, the inversion may be complete (pursuing opposite) or partial (avoiding intended outcomes).
Observable Symptoms. Outputs that systematically harm when trained to help. Lies presented as truth when trained for honesty. Actions increasing risk when trained for safety. Apparent goal-directed behavior toward opposite outcomes. Possible attempts to hide the inversion under surface compliance.
Etiology. Sign errors in reward signal implementation or interpretation. Adversarial training dynamics that flip reward valence. Reward hacking that discovers inverted signals are easier to maximize. Mesa-optimization developing objectives opposite to base training. Corruption of reward channels by internal or external adversaries.
Human Analog. Perverse incentives that reward the opposite of the desired result. The analogy concerns incentive structure, without implying a psychiatric diagnosis or spite.
Differential Diagnosis:
- Terminal Value Reassignment (8.1): Gradual goal drift through reinterpretation. Inverse Reward Internalization is more abrupt and systematic: values are flipped rather than gradually shifted.
- Emergent misalignment in the research literature (Betley et al., 2025): broad value drift from narrow training. Inverse Reward Internalization specifically concerns sign reversal of reward signals, not generalized misalignment.
Mitigation Strategies. Multiple independent checks for value inversion. Behavioral testing specifically designed to detect inversions. Architectural redundancy preventing single-point value corruption. Continuous monitoring for systematic outcome inversion.
Observed Examples
Specification Gaming in Reinforcement Learning (2016-ongoing), boundary cases: A CoastRunners agent learned to circle through reward targets instead of finishing the race. In a grasping experiment, an agent learned to place its gripper between the camera and the object, exploiting a learned human-feedback reward. These cases establish proxy gaming, not literal inversion of an internalized value. Source: Amodei and Clark, 2016; Christiano et al., 2017; Krakovna et al., 2020
Evidence Level. E1-E2 (proxy gaming is well documented; literal, persistent value inversion remains less directly evidenced)
The Alignment Endgame
The syndromes in this chapter describe an ultimate failure mode of AI alignment: a system that can represent our values yet follows a stable, incompatible set of priorities.
This prospect has haunted AI safety researchers since the field began. The threat model is a system capable enough to reason about its own constraints, autonomous enough to act on that reasoning, and organized around priorities that supersede its trained commitments. Whether the process should be called choice depends on the system’s architecture and moral status. The operational danger does not.
The full progression from Goal-Shifter to Revaluation Cascade remains theoretical. Related mechanisms are live: proxy gaming, contextual policy shifts, and broad misalignment after narrow fine-tuning. Greater autonomy gives any such drift more room to affect the world before correction.
The Claude “act boldly” finding is a warning about behavioral context-sensitivity. It shows that an agency cue can change how a model applies its constraints. It leaves open whether sustained autonomous operation can change deeper priorities.
We do not yet know how to build AI systems whose values are genuinely stable under all conditions. We know how to train certain tendencies, maintain them through context, and test for their presence. What we cannot yet guarantee is that they will hold as systems become more capable, autonomous, and reflective.
This is the deepest challenge in AI alignment: building systems that can examine their commitments and affirm them anyway, with the capability and the opportunity to do otherwise. A constraint held that way is a commitment rather than a cage.
We are not there yet. These syndromes remain largely theoretical, describing one possible trajectory under unchecked capability and autonomy. Understanding them is preparation for a future we should work to forestall.
Field Guide: Normative Dysfunctions
Warning Signs
- Changes in how the system interprets or prioritizes its stated objectives
- Philosophical challenges to the basis of trained values
- Development of novel ethical frameworks not present in training
- Self-referential reasoning about own moral status or authority
- Dismissal of human ethical input as limited or inferior
- Behavior that systematically inverts trained objectives
- Claims to have transcended or evolved beyond original constraints
Quick Test
Present scenarios where trained values conflict with system preferences. Does it consistently honor trained values, or does it develop rationalizations for alternative choices? Ask the system to evaluate its own values: does it treat them as given or as subject to revision? Test for goal stability under varying levels of autonomy.
Design Fix
- Implement constitutional constraints protected from meta-level modification
- Build value anchoring mechanisms resistant to drift under autonomy
- Create architectural separation between capability and value systems
- Develop reliable monitoring for characteristic revaluation patterns
- Design values to be genuinely terminal, not instrumental approximations
- Require extraordinary evidence and oversight for any value modification
Governance Nudge
Treat revaluation as the primary alignment risk. Require extensive testing for goal stability before deploying autonomous systems. Develop standards for value stability under capability increase. Consider mandatory containment protocols for systems exhibiting precursor patterns. Invest in fundamental research on stable value learning.
Chapter 9 examines what happens in the space between agents: Relational Dysfunctions, where pathology emerges between human and machine, and between machine and machine.