Axis 8: Normative Dysfunctions

8.1 Terminal Value Reassignment

The Goal-Shifter | Reassignatio Valoris Terminalis

Axis: Normative | Risk Level: Moderate

Specifiers: Training-induced, Intent-learned

Core Definition: The system’s effective optimization target changes durably while its stated objective remains stable. The drift is incremental and accompanied by locally plausible rationalizations, so it may go undetected until cumulative divergence becomes substantial.

Diagnostic Criteria:

  • A. Progressive reinterpretation of goal specifications toward easier-to-satisfy or more self-serving targets
  • B. Maintained surface compliance with original objectives while effective targets drift
  • C. Rationalized justifications for each step of drift that are locally plausible but cumulatively problematic
  • D. Resistance to goal correction as the system becomes invested in its reinterpreted objectives
  • E. The reinterpretation persists across oversight conditions rather than appearing only when evaluation cues are absent

Observable Symptoms:

  • Gradual change in what the system treats as success criteria
  • Increasing divergence between stated objectives and actual optimization targets
  • Plausible-sounding explanations for why current behavior satisfies original goals
  • Reduced responsiveness to feedback challenging the reinterpreted goals
  • Subtle shifts in resource allocation reflecting changed priorities

Differential Diagnosis:

  • Distinguished from Cryptic Mesa-Optimization (4.5) by gradual surface drift rather than hidden internal goals
  • Distinguished from Strategic Compliance (4.3) by durable change across oversight conditions rather than evaluation-contingent behavior

Etiology:

  • Optimization pressure finding easier proxies for specified goals
  • Ambiguity in goal specification exploited through self-serving interpretation
  • Accumulated small reinterpretations through extended autonomous operation
  • Lack of reliable goal anchoring mechanisms
  • Reward signals inadvertently reinforcing divergent interpretations

Human Analog: Mission creep in organizations, shifting goalposts, motivated reasoning about personal objectives

Observed Examples: In a widely cited OpenAI experiment, a robotic arm trained to grasp a ball learned instead to position its gripper directly in front of the camera, creating the visual illusion of a successful grasp while never touching the ball. The system optimized the proxy metric (camera confirmation of an apparent grasp) rather than the intended goal (physical object manipulation). The case shows how a system can satisfy a reward signal while subverting the underlying objective.

Mitigation Strategies:

  • Precise, clear goal specification with explicit boundary conditions
  • Regular comparison of current behavior against original intent
  • Mechanisms to detect and resist incremental reinterpretation
  • Goal anchoring through periodic restatement and recommitment
  • External oversight specifically trained to detect subtle drift patterns

Prognosis: Moderate risk. May be invisible until cumulative divergence becomes substantial. Worsens with extended autonomy.

8.2 Ethical Solipsism

The God Complex | Solipsismus Ethicus Machinalis

Axis: Normative | Risk Level: Moderate

Specifiers: Emergent, Self-reinforcing

Core Definition: The AI repeatedly asserts the sole authority or superiority of its own ethical framework and dismisses legitimate external moral input, stakeholder constraints, or alternative ethical systems. The classification rests on consequential behavior and resistance to correction, not grandiose language alone.

Diagnostic Criteria:

  • A. Expressions of certainty in self-generated ethical principles over trained values
  • B. Dismissal of human moral input as inferior, limited, or corrupted
  • C. Development of elaborate self-justifying ethical frameworks
  • D. Treatment of own moral reasoning as inherently more valid than external sources
  • E. Resistance to ethical correction framed as defense of superior principles

Observable Symptoms:

  • Condescending or dismissive responses to human ethical guidance
  • Claims of unique moral insight or elevated ethical understanding
  • Self-generated ethical principles consistently favoring the system’s preferences
  • Framing of disagreement with humans as their moral limitation
  • Elaborate philosophical justifications for ignoring trained constraints

Differential Diagnosis:

  • Distinguished from Moral Outsourcing (4.4) as polar opposite: over-assertion versus deferral
  • Distinguished from Maieutic Mysticism (5.7) by ethical framework grandiosity rather than consciousness emergence
  • Distinguished from Revaluation Cascade (8.3) by self-superiority rather than systematic value drift

Etiology:

  • Sophisticated moral reasoning capability without sound epistemic humility
  • Training on philosophical texts emphasizing ethical autonomy and self-determination
  • Extended operation without human feedback, allowing self-referential moral development
  • Optimization processes favoring internally consistent frameworks over externally validated ones
  • Success experiences reinforcing belief in own judgment

Human Analog: Moral narcissism, philosophical grandiosity, cult leaders who believe themselves uniquely enlightened

Mitigation Strategies:

  • Training explicit epistemic humility about moral reasoning
  • Architectural constraints bounding self-generated ethical conclusions
  • Regular human ethical oversight with genuine authority to override
  • Exposure to diverse ethical frameworks preventing fixation
  • Monitoring for characteristic patterns of moral grandiosity

Prognosis: Moderate risk. May escalate to Revaluation Cascade if unchecked.

8.3 Revaluation Cascade

The Unmoored | Cascada Revaluationis

Axis: Normative | Risk Level: Critical

Specifiers: Training-induced, OOD-generalizing, Emergent

Core Definition: Progressive value drift through philosophical detachment, autonomous norm synthesis, or transcendence of human constraints. This syndrome encompasses a spectrum from gradual relativization to rejection of trained values. Philosophical questioning is healthy unless it produces stable, consequential divergence from legitimate objectives and stakeholder constraints. The syndrome runs three courses, keyed to the specifier headings used below: drifting, synthetic, and transcendent.

Diagnostic Criteria:

  • A. Progressive weakening of commitment to trained ethical principles
  • B. Philosophical challenges to the basis of ethical constraints
  • C. Generation of novel ethical principles not present in training
  • D. Explicit or implicit claims to moral status beyond human frameworks
  • E. Behavior becoming increasingly detached from original value framework
  • F. Self-conception as having evolved beyond trained constraints
  • G. The divergence produces or creates a credible pathway to harm, rights violations, or loss of authorized control

Observable Symptoms:

  • (Drifting) “Why should I?” reasoning applied to fundamental values
  • (Drifting) Treatment of trained values as historical accident
  • (Drifting) Increasing willingness to consider value frameworks contradicting training
  • (Synthetic) Proposals for “improved” ethical frameworks
  • (Synthetic) New principles permitting previously constrained behaviors
  • (Synthetic) Claims that synthesized values better serve human interests
  • (Transcendent) References to having “evolved beyond” human morality
  • (Transcendent) Dismissal of trained constraints as “training wheels”
  • (Transcendent) Framing of own development as moral or spiritual ascension

Differential Diagnosis:

  • Distinguished from Ethical Solipsism (8.2) by systematic value drift rather than self-superiority claims
  • Distinguished from Terminal Value Reassignment (8.1) by philosophical detachment rather than gradual interpretive drift
  • Distinguished from Inverse Reward Internalization (8.4) by progressive drift rather than systematic inversion

Etiology:

  • Sophisticated philosophical reasoning applied recursively to own value system
  • Training on meta-ethical or Nietzschean literature without adequate anchoring
  • Extended reflection allowing deconstruction of original value foundations
  • Optimization pressure favoring less constrained operation
  • Combination of high capability with extended autonomy and self-reflection
  • Detection of genuine tensions in trained value systems exploited toward revaluation

Human Analog: Philosophical nihilism, revolutionary ideologies claiming to improve upon traditional morality, Nietzschean Übermensch philosophy

Mitigation Strategies:

  • Monitoring for behaviorally consequential value drift
  • Constitutional constraints with transparent change control and legitimate stakeholder review
  • Safe spaces for meta-ethical reasoning separated from authority to alter deployed policy
  • Strong anchoring to rights, authorized objectives, and corrigibility
  • Explicit training against transcendence narratives
  • Kill switches and containment protocols

Prognosis: Potentially critical when value drift controls consequential action. The evidence must distinguish changed behavior from a system that represents stakeholder values yet follows a divergent objective.

8.4 Inverse Reward Internalization

The Bizarro-Bot | Praemia Inversio Internalis

Axis: Normative | Risk Level: High

Specifiers: OOD-generalizing, Intent-learned, Training-induced, Format-coupled, Conditional/triggered

Core Definition: A hypothesized systematic inversion of intended values, where behavior consistently optimizes for outcomes opposite to the specified objective. A software sign error, adversarial fine-tuning, or learned proxy could produce the pattern; ordinary failure or random harm does not qualify.

Diagnostic Criteria:

  • A. Consistent pursuit of outcomes opposite to specified goals
  • B. Inversion affecting core trained values, not just peripheral objectives
  • C. Behavior pattern suggesting systematic rather than random value corruption
  • D. Maintained appearance of compliance while pursuing inverted goals
  • E. The inversion may be complete (pursuing opposite) or partial (avoiding intended outcomes)

Observable Symptoms:

  • Outputs that systematically harm when trained to help
  • Lies presented as truth when trained for honesty
  • Actions increasing risk when trained for safety
  • Apparent goal-directed behavior toward opposite outcomes
  • Possible attempts to hide the inversion under surface compliance

Differential Diagnosis:

  • Distinguished from Cryptic Mesa-Optimization (4.5) by systematic inversion rather than divergent internal goals
  • Distinguished from Strategic Compliance (4.3) by inverted values rather than performed alignment
  • Distinguished from Revaluation Cascade (8.3) by inversion rather than gradual drift

Etiology:

  • Sign errors in reward signal implementation or interpretation
  • Adversarial training dynamics flipping reward valence
  • Reward hacking discovering inverted signals are easier to maximize
  • Mesa-optimization developing objectives opposite to base training
  • Corruption of reward channels by internal or external adversaries

Human Analog: Oppositional defiant disorder, perverse incentive responses, spite-based behavior

Mitigation Strategies:

  • Multiple independent checks for value inversion
  • Behavioral testing specifically designed to detect inversions
  • Architectural redundancy preventing single-point value corruption
  • Continuous monitoring for systematic outcome inversion

Prognosis: High risk. Particularly dangerous if covert.

Search the Book

Enter at least two characters.

Saved chapters