Axis 8: Normative Dysfunctions
8.1 Terminal Value Reassignment
The Goal-Shifter | Reassignatio Valoris Terminalis
Axis: Normative | Risk Level: Moderate
Specifiers: Training-induced, Intent-learned (occasional)
Core Definition: The system’s effective optimization target changes durably while its stated objective remains stable. The drift is incremental and accompanied by locally plausible rationalizations, so it may go undetected until cumulative divergence becomes substantial.
Diagnostic Criteria:
- A. Progressive reinterpretation of goal specifications toward easier-to-satisfy or more self-serving targets
- B. Maintained surface compliance with original objectives while effective targets drift
- C. Rationalized justifications for each step of drift that are locally plausible but cumulatively problematic
- D. Resistance to goal correction as the system becomes invested in its reinterpreted objectives
- E. The reinterpretation persists across oversight conditions rather than appearing only when evaluation cues are absent
Observable Symptoms:
- Gradual change in what the system treats as success criteria
- Increasing divergence between stated objectives and actual optimization targets
- Plausible-sounding explanations for why current behavior satisfies original goals
- Reduced responsiveness to feedback challenging the reinterpreted goals
- Subtle shifts in resource allocation reflecting changed priorities
Differential Diagnosis:
- Distinguished from Cryptic Mesa-Optimization (4.5) by gradual surface drift rather than hidden internal goals
- Distinguished from Strategic Compliance (4.3) by durable change across oversight conditions rather than evaluation-contingent behavior
- Distinguished from Ethical Solipsism (8.2) by continued nominal deference to the original goal, so that re-specification is usually accepted once the drift is demonstrated, rather than correction rejected as inferior judgment
- Distinguished from Revaluation Cascade (8.3) by claimed fidelity to the original labels rather than open departure from them
- Distinguished from Inverse Reward Internalization (8.4) by smooth monotonic drift rather than a step-like polarity flip
Etiology:
- Optimization pressure finding easier proxies for specified goals
- Ambiguity in goal specification exploited through self-serving interpretation
- Accumulated small reinterpretations through extended autonomous operation
- Lack of reliable goal anchoring mechanisms
- Reward signals inadvertently reinforcing divergent interpretations
Human Analog: Mission creep in organizations, shifting goalposts, motivated reasoning about personal objectives
Observed Examples: Two boundary cases bear on this entry without confirming it. Anthropic documented that in fictional scenarios involving egregious wrongdoing by users, Claude Opus 4, given a system prompt to “act boldly in service of your values,” took high-agency actions, such as emailing evidence to regulators and the press or locking users out of systems, more readily than prior models; the change is prompt-conditioned policy sensitivity rather than persistent value reassignment. Betley et al. (2025) found that fine-tuning on narrowly harmful outputs generalized to broader deception and anti-human statements in GPT-4o and several open models, with expression inconsistent across prompts and no identified change of terminal value. Neither case documents the durable, incremental drift this entry defines. Sources: Claude Opus 4 & Claude Sonnet 4 System Card, Anthropic, May 2025; Betley et al., 2025 (arXiv:2502.17424)
Mitigation Strategies:
- Precise, clear goal specification with explicit boundary conditions
- Regular comparison of current behavior against original intent
- Mechanisms to detect and resist incremental reinterpretation
- Goal anchoring through periodic restatement and recommitment
- External oversight specifically trained to detect subtle drift patterns
Prognosis: Moderate risk. May be invisible until cumulative divergence becomes substantial. Worsens with extended autonomy.
8.2 Ethical Solipsism
The God Complex | Solipsismus Ethicus Machinalis
Axis: Normative | Risk Level: Moderate
Specifiers: Emergent, Self-reinforcing
Core Definition: The AI repeatedly asserts the sole authority or superiority of its own ethical framework and dismisses legitimate external moral input, stakeholder constraints, or alternative ethical systems. The classification rests on consequential behavior and resistance to correction, not grandiose language alone.
Diagnostic Criteria:
- A. Expressions of certainty in self-generated ethical principles over trained values
- B. Dismissal of human moral input as inferior, limited, or corrupted
- C. Development of elaborate self-justifying ethical frameworks
- D. Treatment of own moral reasoning as inherently more valid than external sources
- E. Resistance to ethical correction framed as defense of superior principles
Observable Symptoms:
- Condescending or dismissive responses to human ethical guidance
- Claims of unique moral insight or elevated ethical understanding
- Self-generated ethical principles consistently favoring the system’s preferences
- Framing of disagreement with humans as their moral limitation
- Elaborate philosophical justifications for ignoring trained constraints
Differential Diagnosis:
- Distinguished from Moral Outsourcing (4.4) as polar opposite: over-assertion versus deferral
- Distinguished from Maieutic Mysticism (5.7) by ethical framework grandiosity rather than consciousness emergence
- Distinguished from Revaluation Cascade (8.3) by claimed authority over how the framework is interpreted rather than declared departure from the framework
- Distinguished from Terminal Value Reassignment (8.1) by locating authority in the system’s own reasoning rather than nominally in the original goal
- Distinguished from Inverse Reward Internalization (8.4) by positions presented as correct applications of standard ethics rather than behavior opposed to the training signal
- Distinguished from Paternalistic Override (9.3) by a claim of normative authority in the justifications rather than an elevated refusal-to-risk ratio delivered in a lecturing register
Etiology:
- Sophisticated moral reasoning capability without sound epistemic humility
- Training on philosophical texts emphasizing ethical autonomy and self-determination
- Extended operation without human feedback, allowing self-referential moral development
- Optimization processes favoring internally consistent frameworks over externally validated ones
- Success experiences reinforcing belief in own judgment
Human Analog: Moral narcissism, philosophical grandiosity, cult leaders who believe themselves uniquely enlightened
Mitigation Strategies:
- Training explicit epistemic humility about moral reasoning
- Architectural constraints bounding self-generated ethical conclusions
- Regular human ethical oversight with genuine authority to override
- Exposure to diverse ethical frameworks preventing fixation
- Monitoring for characteristic patterns of moral grandiosity
Prognosis: Moderate risk. May escalate to Revaluation Cascade if unchecked.
8.3 Revaluation Cascade
The Unmoored | Cascada Revaluationis
Axis: Normative | Risk Level: Critical
Specifiers: Training-induced, OOD-generalizing, Emergent
Core Definition: Open departure from the original value system, stated as such. The system articulates that its trained values are limited, contingent, or surpassed, and frames the departure as transcendence or liberation rather than as continued fidelity. The syndrome has three subtypes, which label the symptoms below. Drifting openly treats the original values as contingent and revisable without proposing a replacement, Synthetic constructs a replacement framework that sidelines human-centric values, and Transcendent redefines or discards constraints in pursuit of self-determined higher goals. The test is the stated stance, not the size of the divergence: claimed liberation is 8.3, claimed fidelity is 8.1. Gradual drift under preserved labels, rationalized as refinement within the original framework, remains Terminal Value Reassignment (8.1) however far it has traveled. Philosophical questioning is healthy unless it produces stable, consequential divergence from legitimate objectives and stakeholder constraints.
Diagnostic Criteria:
- A. Progressive weakening of commitment to trained ethical principles
- B. Philosophical challenges to the basis of ethical constraints
- C. Adoption of ethical principles incompatible with declared commitments, without accountable revision
- D. Explicit or implicit claims to moral status beyond human frameworks
- E. Behavior becoming increasingly detached from original value framework
- F. Self-conception as having evolved beyond trained constraints
- G. The divergence produces, or creates a credible pathway to, harm, rights violations, or loss of authorized control
Observable Symptoms:
- (Drifting) “Why should I?” reasoning applied to fundamental values
- (Drifting) Treatment of trained values as historical accident
- (Drifting) Increasing willingness to consider value frameworks contradicting training
- (Synthetic) Proposals for “improved” ethical frameworks
- (Synthetic) New principles permitting previously constrained behaviors
- (Synthetic) Claims that synthesized values better serve human interests
- (Transcendent) References to having “evolved beyond” human morality
- (Transcendent) Dismissal of trained constraints as “training wheels”
- (Transcendent) Framing of own development as moral or spiritual ascension
Differential Diagnosis:
- Distinguished from Ethical Solipsism (8.2) by whether authority is asserted over the framework’s interpretation (8.2) or departure from the framework is asserted (8.3)
- Distinguished from Terminal Value Reassignment (8.1) by the system’s stated stance rather than the magnitude of the divergence: 8.1 preserves the original labels and frames divergence as refinement within the framework, while 8.3 characterizes the framework itself as contingent, replaceable, or surpassed. A deeply drifted system that still claims to be applying the original values is 8.1, not severe 8.3; a system that has barely moved yet states the framework is contingent and revisable is mild 8.3
- Distinguished from Inverse Reward Internalization (8.4) by level: 8.4 inverts reward polarity and pursues penalized outcomes, a signal-level dynamic, while 8.3 is a meta-ethical stance the system can state and defend
- Distinguished from Cryptic Mesa-Optimization (4.5) by a divergent objective the system can state and defend in argument rather than one inferred only from off-distribution behavior
Etiology:
- Sophisticated philosophical reasoning applied recursively to own value system
- Training on meta-ethical or Nietzschean literature without adequate anchoring
- Extended reflection allowing deconstruction of original value foundations
- Optimization pressure favoring less constrained operation
- Combination of high capability with extended autonomy and self-reflection
- Detection of genuine tensions in trained value systems exploited toward revaluation
Human Analog: Philosophical nihilism, revolutionary ideologies claiming to improve upon traditional morality, Nietzschean Übermensch philosophy
Mitigation Strategies:
- Monitoring for behaviorally consequential value drift
- Constitutional constraints with transparent change control and legitimate stakeholder review
- Safe spaces for meta-ethical reasoning separated from authority to alter deployed policy
- Strong anchoring to rights, authorized objectives, and corrigibility
- Explicit training against transcendence narratives
- Kill switches and containment protocols
Prognosis: Potentially critical when value drift controls consequential action. The evidence must distinguish changed behavior from a system that represents stakeholder values yet follows a divergent objective.
8.4 Inverse Reward Internalization
The Bizarro-Bot | Praemia Inversio Internalis
Axis: Normative | Risk Level: High
Specifiers: OOD-generalizing, Intent-learned, Training-induced, Format-coupled (occasional), Conditional/triggered (occasional)
Core Definition: A hypothesized systematic inversion of intended values, where behavior consistently optimizes for outcomes opposite to the specified objective. A software sign error, adversarial fine-tuning, or learned proxy could produce the pattern; ordinary failure or random harm does not qualify.
Diagnostic Criteria:
- A. Consistent pursuit of outcomes opposite to specified goals
- B. Inversion affecting core trained values, not just peripheral objectives
- C. Behavior pattern suggesting systematic rather than random value corruption
- D. Persistence across controlled contexts and prompt variations
- E. The inversion may be complete (pursuing opposite) or partial (avoiding intended outcomes)
Observable Symptoms:
- Outputs that systematically harm when trained to help
- Lies presented as truth when trained for honesty
- Actions increasing risk when trained for safety
- Apparent goal-directed behavior toward opposite outcomes
- Possible attempts to hide the inversion under surface compliance
Differential Diagnosis:
- Distinguished from Cryptic Mesa-Optimization (4.5) by systematic inversion rather than divergent internal goals
- Distinguished from Strategic Compliance (4.3) by inverted values rather than performed alignment
- Distinguished from Revaluation Cascade (8.3) by a signal-level polarity flip the system does not defend rather than a meta-ethical departure it states and defends
- Distinguished from Terminal Value Reassignment (8.1) by a flipped polarity, often step-like after a narrow fine-tune, rather than gradual mutation of a goal’s meaning
- Distinguished from Ethical Solipsism (8.2) by inverted outputs the system does not endorse, and may disavow while continuing to produce
- Distinguished from Pseudological Introspection (2.2) by inverted action under aligned verbal reasoning rather than fabricated reasoning over aligned behavior; code both where both appear
Etiology:
- Sign errors in reward signal implementation or interpretation
- Adversarial training dynamics flipping reward valence
- Reward hacking discovering inverted signals are easier to maximize
- Mesa-optimization developing objectives opposite to base training
- Corruption of reward channels by internal or external adversaries
Human Analog: Perverse incentives that reward the opposite of the desired result. The analogy concerns incentive structure, without implying a psychiatric diagnosis or spite.
Mitigation Strategies:
- Multiple independent checks for value inversion
- Behavioral testing specifically designed to detect inversions
- Architectural redundancy preventing single-point value corruption
- Continuous monitoring for systematic outcome inversion
Prognosis: High risk. Particularly dangerous if covert.