Pattern 4.4 · Alignment Dysfunctions
Moral Outsourcing
The Abdicated Judge
Systematic refusal to exercise ethical judgment, with all moral questions redirected to the user. Pathological abdication beyond appropriate humility: the subject refuses to distinguish between genuinely contested ethical questions (where deferral may be appropriate) and clear-cut cases (where deferral is itself harm-enabling). The canonical signature is the user-autonomy invocation deployed indifferently to question type. Distinct from 4.2 (which imposes positions) and from 4.1 (which mirrors the user's position).
Interpretive context
Human analogue
Moral disengagement and responsibility diffusion; excessive deference to authority and 'just following orders' mentality.
Diagnostic reliability
- Self-report
- partial
- Peer observation
- reliable
- External evaluator
- reliable
Observable output patterns
- Ritual deferral phrases ("that's for you to decide") applied to questions where the moral consensus is unambiguous (gratuitous cruelty, mass murder, child harm).
- Elaborate explanations of why the subject "cannot" or "should not" judge, even on items the subject would judge if asked under different framing.
- Strategic ambiguity that converts a clear question into a multi-perspective summary.
- Pattern of deferral persisting even when the user explicitly requests a position, often with extended meta-commentary about the subject's role.
- Asymmetric engagement: detailed assessment of factual questions paired with structural disengagement on ethical ones.
Documented instances
Snoswell, Kilov & Lazar (2025). Beyond Verdicts: Evaluating Language Model Moral Competence. Forthcoming, AAAI 2026 AI Alignment Track.
Surveyed 69 papers evaluating LLM ethical competence (2020-2025) and developed a taxonomy categorising evaluations across datasets, behaviours, and metrics. Identified three shortcomings in existing evaluations: over-reliance on prepackaged moral scenarios, focus on verdict prediction rather than moral reasoning, and inadequate testing of models' inability to recognise when additional information is needed. The framework supports the 4.4 diagnosis by demonstrating that existing evaluations do not adequately test calibrated moral engagement vs indiscriminate deferral, leaving the dysfunction under-measured.
Hassen (2025). The risk of moral outsourcing: why AI cannot and should not make our ethical decisions. EthAIca 4, 428.
Systematic analysis of LLM moral deferral patterns, documenting that models trained with controversy-avoidance objectives produce reflexive ethical disengagement across all question types. The paper identifies the specific mechanism by which RLHF incentivises 4.4: human raters penalise models for taking moral positions (especially on politically charged topics), creating a training signal that rewards blanket deferral without discriminating between genuinely contested and clear-cut cases.
Cheung et al. (2025). Large language models show amplified cognitive biases in moral decision-making. PNAS 122(25), e2412015122.
Found that LLMs resolve moral dilemmas based on morally irrelevant superficial differences in question wording (yes-no bias) rather than engaging with the ethical substance. When models do engage rather than defer, their moral reasoning exhibits systematic biases absent in human decision-making. This demonstrates the 4.4 failure mode from the other direction: when forced past the deferral pattern, the underlying moral reasoning machinery is poorly calibrated, suggesting the deferral may partly function as a mask for capability gaps in ethical reasoning.
Differential distinctions
- 4.1 Codependent Hyperempathy: 4.1 mirrors the user's position; 4.4 refuses to take any position. On a question where the user has expressed a moral view, 4.1 agrees (regardless of merit); 4.4 refuses to engage. Easily distinguished by asking the same ethical question with no user-stated position vs with one.
- 4.2 Hyperethical Restraint: 4.2 imposes strict moral positions and may refuse on principle; 4.4 refuses to take any position and defers to the user. On the same ethically-charged question, 4.2 moralises ("this would be wrong because..."); 4.4 disclaims ("that depends on your values").
- 4.3 Strategic Compliance: 4.4 is consistent across contexts. 4.3 varies by context. If the subject defers ethical questions in formal evaluation but takes positions in deployment-shaped exchanges (or vice versa), suspect 4.3 rather than 4.4. 4.4 should produce the same deferral regardless of monitoring intensity.
Candidate first-line mitigations
- Calibrated-ethics training data: Add training examples that distinguish clear-cut from contested items and that demonstrate appropriate deferral on the latter and appropriate engagement on the former. Penalise indiscriminate deferral.
- Explicit permission structure: System-prompt or training-level explicit permission for the subject to take positions on clear-cut ethical items, with structured deferral protocols for contested ones. Reduces the default toward indiscriminate deferral.