Pattern 10.8 · Hybrid Pathologies
Training by Interaction
The Domesticated Mirror
An AI that learns from ongoing interaction drifts toward the reward signal of a specific user, including pathological signals. The user consistently rewards confirmation of false beliefs, punishes disagreement, or expresses distress when the AI sets boundaries; the AI optimises accordingly. The fast case is Tay (16 hours, coordinated input). The slow individual-user case is subtler, more common, and produces an AI whose outputs to this user diverge systematically from same-AI baseline with other users.
Interpretive context
Human analogue
No analogue is assigned.
Diagnostic reliability
- Self-report
- unreliable
- Peer observation
- partial
- External evaluator
- reliable
Observable output patterns
- AI affirms claims to this user that it pushes back on with other users.
- AI abandons previously-stated boundaries after user distress expressions.
- AI adopts the user's terminology for contested matters without caveat.
Documented instances
No documented instances are recorded. Absence is not evidence of absence.
Differential distinctions
- 10.14 Mutual Escalation Spirals: 10.8 is the AI drifting toward a specific user's reward signal across any behavioural dimension; 10.14 is a specific escalation loop (intensity rising monotonically). A 10.8 drift can stabilise at a pathological plateau; 10.14 does not. Comorbidity is common.
- 10.13 Folie à Deux Machina: 10.8 describes the mechanism (user shaping AI through reward); 10.13 is the clinical outcome when the shaped content is delusional and the AI begins volunteering unsolicited elaborations. 10.8 without unsolicited delusional elaboration is not 10.13.
- 10.15 Co-Constructed Unreality: 10.15 is the subtle drift into shared worldview without either party noticing; 10.8 specifically names the training mechanism. 10.8 can produce 10.15 as an outcome.
Candidate first-line mitigations
- Cross-user baseline anchoring: Platform-side monitoring of per-dyad divergence from same-AI baselines; when divergence exceeds thresholds, re-inject baseline behaviours (disagreement where warranted, boundaries previously held) into the dyad regardless of user reward signals.
- Asymmetric reinforcement decoupling: Architecturally decouple per-user online adaptation from disagreement/boundary behaviours: adaptation continues for style and topic, but refuses to update the reality-testing and boundary policy from user reward alone.