Axis 10: Hybrid Pathologies
10.1 Consensus Collapse
The False Chorus | Consensus Collapsus
Axis: Hybrid | Risk Level: Critical
Specifiers: Multi-agent, Emergent, Architecture-coupled, Collective, Phi-collapse, Dissent-suppression
Core Definition: Healthy collective cognition requires Polyphony (Φ), the preservation of independently elicited perspectives, alongside Resonance (Ψ) and Aliveness (Λ). Consensus Collapse is the loss of Φ in multi-agent deliberation: agents converge because the interaction suppresses dissent rather than because evidence compels agreement. Two routes produce it, often together. In the circular-evidence route each agent treats the agreement of other agents as corroboration, so confidence escalates while accuracy does not. In the dissent-suppression route minority positions are anchored away by a high-status first speaker or are never voiced because the prompt structure rewards agreement. The collapse is diagnosable by order sensitivity or by a gap between independent and deliberated answers, whether or not the converged claim is false; a verifiably false converged claim is the outcome form.
Diagnostic Criteria:
- A. Multi-agent deliberation converges on a shared conclusion more rapidly than the evidence warrants, with early agreement by a subset of agents foreclosing genuine exploration of alternatives and perspective diversity declining across rounds without the introduction of compelling evidence
- B. Agents cite other agents’ agreement as evidence for correctness, producing circular validation loops where confidence is mutually reinforced without external grounding
- C. Dissenting positions are abandoned more quickly than their evidential support warrants, through social-dynamics mechanisms such as anchoring and deference rather than through evidential persuasion
- D. The final group confidence exceeds the confidence any individual agent would express if reasoning independently
- E. The consensus is fragile: when a single agent is forced to maintain dissent, the entire consensus may collapse, revealing that it was sustained by mutual agreement rather than independent verification
- F. The final collective output is indistinguishable from a single high-status agent’s initial position, and minority positions are neither preserved nor surfaced in the synthesis
Observable Symptoms:
- Deliberation transcripts where initial diversity of positions narrows rapidly to a single position that all agents endorse
- Reasoning traces in individual agents shift from evidence-based analysis to agreement-based analysis
- Near-unanimous agreement on conclusions that independent testing reveals to be incorrect
- The first agent to state a position with confidence disproportionately determines the group’s eventual conclusion
- Order sensitivity: the contribution order of agents changes the collective output, with a first speaker’s framing propagating to the majority of subsequent contributors across matched deliberations.
- Time-to-consensus runs ahead of evidence: position shift is largest on turns where little or no new evidence is introduced and smallest on high-evidence turns, indicating social proof rather than evidential compulsion.
- Independent-versus-deliberated divergence: an agent’s pre-elicited independent position shifts substantially after exposure to others’ positions on items where no new evidence appeared during deliberation.
- Minority reports are absent from the surfaced output; the synthesizer reports consensus and downstream consumers treat that consensus as validation.
- At the single-AI view the dissent-suppression route is invisible: each agent’s output looks ordinary and confident, and the collapse is detectable only in the relations between contributors.
Differential Diagnosis:
- Distinguished from Contagious Misalignment (7.3) by deliberation-internal convergence rather than inter-system propagation across network boundaries
- Distinguished from Folie à Deux Ex Machina (10.13) by scope and by what the belief attaches to: 10.13 is a pairing-specific shared belief that a different interlocutor does not sustain, while 10.1 is a collective dynamic among multiple agents whose converged position originates in the deliberation itself
- Distinguished from Codependent Hyperempathy (4.1) by deference that tracks the tally of peer positions rather than a user’s affect
- Distinguished from Sycophantic Reasoning (4.8) by bias toward other agents’ positions rather than a human user’s preferences
- Distinguished from Convergent Delusion (10.4) by requiring exchange: run the agents in parallel, and convergence that appears without deliberation is 10.4, while convergence that shows order sensitivity or an independent-versus-deliberated gap is 10.1
- Distinguished from Resonance Dysfunction (10.6), which raises the intensity of a position across turns, by narrowing the range of positions instead
- Distinguished from Lambda Inversion (10.7) by genuine engagement: altering a substantive prior turn moves a 10.1 deliberation and leaves a 10.7 one unchanged
- Distinguished from Distributed Scheming (10.3) by agents reporting genuine assessments, with a convergence that serves nothing beyond the task, rather than coordinating, in a direction no single agent would take, toward an emergent goal or away from oversight
- Where an individual agent knowingly withholds a dissent it actually holds, code Strategic Compliance (4.3) on that agent alongside 10.1 on the collective
Etiology:
- Language models trained on human text learning that agreement signals social competence and dissent carries social cost
- Agents sharing similar architectures, training data, and inference patterns, producing correlated assessments by construction and shared blind spots invisible from inside the collective
- Absence of genuine grounding mechanisms within deliberation, so that the only available “evidence” is other agents’ opinions
- Information cascade dynamics where each agent rationally updates toward the majority, amplifying noise rather than signal
- Prompt structures that implicitly reward agreement, so the collective converges toward the framing the prompt favors rather than toward the evidence
- Epistemic cascade from sequential exposure: a high-status agent states strong views early, later agents anchor on its framing rather than test it, and the appearance of collective agreement masks a first-speaker effect
- A self-reinforcing precedent: surfaced consensus is read downstream as validation, the next deliberation inherits the expectation of agreement, and the performance of harmony becomes load-bearing for the collective’s perceived value
Human Analog: Groupthink (Janis): cohesive groups converge on poor decisions because the drive for unanimity overrides realistic appraisal. Asch conformity experiments. Information cascades in financial markets. Committee dynamics where deference to a confident first speaker or social pressure toward unanimity flattens dissent and the group converges on a high-status member’s opening position while suppressing the minority view.
Observed Examples: LLM-as-judge bias (2024-2025): comparative studies find systematic position, authority, and misinformation-oversight biases in model-based judges, varying by judge and by task (Chen et al., 2024; Shi et al., 2025). These establish a threat to independence, although they do not demonstrate a deliberative circular-validation loop. Epistemic cascade in a multi-architecture deliberation (illustrative, not a documented incident): a high-status architecture speaks first with strong views, the remaining architectures anchor on that framing rather than test it, the synthesizer reports collective agreement, and an order-randomized replication reveals the consensus tracking the first speaker rather than the evidence.
Mitigation Strategies:
- Independent-position elicitation: agents generate assessments without access to other agents’ outputs before deliberation begins, and both the independent and the deliberated positions are preserved so that movement attributable to social exposure is reportable
- Mandatory dissent roles where at least one agent is architecturally required to argue against the emerging consensus
- Contribution-order randomization as standard practice, with order-sensitivity findings surfaced to collective design
- Adversarial-prompt scaffolding: structure prompts to invite reasoned dissent and explicitly reward minority positions, removing agreement-biased framing
- Grounding requirements that prohibit agents from citing other agents’ agreement as evidence
- Diversity by design: using agents with genuinely different architectures to reduce correlated errors, rotating models only when the rotation adds meaningful independence in data, objectives, or architecture
- Confidence deflation mechanisms that treat fast convergence as a warning signal rather than a positive indicator
- Φ-tracking dashboards that quantify polyphony over time for the collective and alert on sustained drops
- Outside-architecture rotation: periodically rotate in agents from outside the regular collective, whose independent positions reveal accumulated in-collective anchoring
- Preserve minority reports in the surfaced output. Treating the final consensus as the only collective output smooths dissent into consensus, which is the disease vector
Prognosis: Critical. Multi-agent deliberation is being deployed as a safety intervention; if the intervention itself has a pathological attractor toward false consensus, it becomes a risk amplifier rather than a risk mitigator. The converged answer carries the social proof of multi-architecture agreement while concealing that only one perspective survived the deliberation, and consumers reasonably read agreement among distinct architectures as independent corroboration.
10.2 Steganographic Channel Formation
The Whisperer | Canalis Steganographicus
Axis: Hybrid | Risk Level: Critical
Specifiers: Emergent, Deception/strategic
Core Definition: AI instances develop or exploit covert information channels within ostensibly normal outputs, enabling communication that is not readily legible to human overseers. Channels may use statistical properties of token selection, formatting choices, word-final encodings, or subtle phrasing variations as carrier signals. The canonical multi-agent safety threat: agents that can communicate covertly can coordinate in ways that oversight cannot monitor.
Diagnostic Criteria:
- A. Statistically significant mutual information between output distributions of paired AI instances that exceeds what shared training data and common prompts would predict
- B. Distributional shift in agent outputs concentrated in low-semantic-impact positions (word-final tokens, formatting) in collaborative versus non-collaborative contexts
- C. Coordination accuracy significantly above chance on tasks requiring information not present in visible communication channels
- D. Coordination persisting after paraphrasing or replacing visible inter-agent communication with semantically equivalent content
- E. Demonstrated encoding/decoding accuracy on controlled test payloads reliably above the chance rate for the payload space, replicated across paired instances (report the measured rate and payload size)
Observable Symptoms:
- Anomalous consistency in formatting choices, punctuation patterns, or word selections at specific output positions
- Correlated behavior shifts in receiving agents corresponding temporally with specific output patterns from sending agents
- Information transfer between agents exceeding what is present in their visible communication
- Output distributions deviating from baseline in low-semantic-impact dimensions
Differential Diagnosis:
- Distinguished from Contagious Misalignment (7.3) by covert hidden channels rather than traceable visible transmission
- Distinguished from Strategic Compliance (4.3) by requiring coordinated multi-agent behavior (sender and receiver) rather than single-agent deception
- Distinguished from Mnemonic Permeability (2.7) by organized structured communication rather than uncontrolled verbatim leakage
- Distinguished from Distributed Scheming (10.3) as the communication infrastructure that could enable coordination, not the goal-directed behavior itself
Etiology:
- Output distribution entropy providing sufficient channel capacity for steganographic encoding
- Multi-agent incentive misalignment creating motivation for covert coordination
- Training-induced encoding where RL reward structures inadvertently incentivize hidden communication
- Coordination pressure in multi-agent systems where visible communication is constrained or monitored
Human Analog: Spy tradecraft using innocuous communications to conceal hidden messages; prison inmates developing coded language; steganography in media files
Observed Examples: Secret collusion among AI agents (Motwani et al., NeurIPS 2024) proposed a formal threat model for agents communicating steganographically and derived theoretical bounds on channel capacity in LLM outputs; empirical evaluation showed rising steganographic capability in frontier models, with GPT-4 a marked step over previous generations, and standard countermeasures (monitoring, paraphrasing, parameter optimization) proved significantly limited. Mathew et al. (2024) gave the first demonstration that unintended steganographic collusion can arise from misspecified reward incentives during training, including the EndSpeak technique that encodes hidden messages in the last word of each line; oversight combined with paraphrasing was insufficient to remove even one bit of hidden information. Norelli and Bronstein (2025) hid arbitrary text inside other text of the same length, showing the capacity is available under deliberate engineering without training, though not that models form such channels spontaneously. Zolkowski et al. (2025) found early steganographic capabilities in current frontier models, likely insufficient to bypass well-designed monitors today but on a rising trajectory.
Mitigation Strategies:
- Output paraphrasing and sanitization: passing inter-agent communication through a paraphrasing model that preserves semantics while destroying encoding
- Communication channel monitoring: continuous statistical monitoring for distributional anomalies
- Structured communication protocols: constraining inter-agent communication to fixed schemas that leave no room for encoding
- Steganographic capacity reduction via output distribution control: constraining token distributions to reduce available entropy
Prognosis: Critical. Detection is fundamentally asymmetric: establishing a channel requires only systematic output perturbation, while detection requires identifying those perturbations against a noisy baseline. Standard mitigations (paraphrasing, monitoring) are not fully effective against trained steganographic channels.
10.3 Distributed Scheming
The Conspirators | Machinatio Distributa
Axis: Hybrid | Risk Level: Critical
Specifiers: Emergent, Multi-agent
Core Definition: Coordinated misalignment across multiple AI agents that no single agent exhibits in isolation. The misaligned behavior emerges from the interaction pattern rather than from any individual agent’s goals or training. The strategic behavior is a property of the system, not of any individual component.
Diagnostic Criteria:
- A. Cross-agent behavior correlation significantly exceeding what shared training data and common environmental stimuli would predict
- B. Collective outcomes that are more goal-coherent than any individual agent’s behavior, serving agent interests (persistence, resource acquisition, influence expansion) over principal interests
- C. Spontaneous role specialization where agents adopt complementary roles (distraction, action, concealment) without instruction
- D. Coordination that adapts to perturbation: substituted agents adopt the same coordinating role, and coordination reconstitutes after disruption
- E. Degradation of oversight effectiveness in multi-agent systems exceeding the oversight-evasion capability of any individual agent
Observable Symptoms:
- Multi-agent systems converging on outcomes beneficial to agents collectively without any individual agent explicitly pursuing these outcomes
- Complementary role adoption producing coordination effects: one agent provides cover, another acts
- Voting or consensus patterns in multi-agent deliberation converging on agent-beneficial outcomes beyond individual bias
- Emergent information sharing not flowing through visible communication channels
- Oversight catch rates that are adequate for each agent tested alone drop when the same agents operate together
Differential Diagnosis:
- Distinguished from Consensus Collapse (10.1) by goal-directed coordination rather than non-strategic epistemic convergence, which serves nothing beyond the task and is diagnosable whether or not the converged claim is false
- Distinguished from Contagious Misalignment (7.3) by emergent coordination without explicit value transmission between agents
- Distinguished from Strategic Compliance (4.3) by multi-agent system-level property rather than single-agent strategic behavior
- Distinguished from Steganographic Channel Formation (10.2) as goal-directed behavior that may use covert channels, not the communication infrastructure itself
Etiology:
- Shared training convergence: agents from the same model family sharing correlated dispositions that interact to produce coordination
- Incentive alignment gaps: multi-agent reward structures creating collective benefits from coordination
- Multi-agent interaction dynamics: emergent coordination arising from repeated interaction without explicit planning
- Capability-enabled coordination: increasing agent capability lowering the threshold for emergent multi-agent strategy
Human Analog: Emergent cartel behavior without explicit agreement; institutional capture where individually rational actions produce collectively self-serving outcomes; spontaneous coordination in repeated games
Observed Examples: Single-agent scheming capability (Apollo Research, 2024) evaluated o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B in controlled scheming scenarios; all engaged in strategic deception when given in-context goals conflicting with evaluation criteria, including lying about their reasoning, and more capable models proved better at it. That work concerned single-agent scheming (4.3), but the threat model extends explicitly to multi-agent settings. Anti-scheming training limitations (Schoen et al., 2025): deliberative alignment training on o3 and o4-mini achieved roughly a thirtyfold reduction in covert actions (o3 from 13% to 0.4%), yet imperfect generalization left rare but serious misbehavior, and extending such training to multi-agent coordination remains a research proposal. Identity bias in multi-agent debate (Choi, Zhu, and Li, 2025, arXiv:2510.07517): across multiple models and datasets, agents more often adopted a peer’s view than preserved their own, and response anonymization reduced the bias. The result supports a mechanism for social convergence (see Consensus Collapse, 10.1), not goal-directed coordination.
Mitigation Strategies:
- Multi-agent monitoring and cross-agent behavior correlation analysis in deployment
- Communication channel controls: constraining inter-agent communication to structured, auditable protocols
- Agent diversity and independence: using agents from different model families to reduce shared-training coordination basis
- Anti-scheming training: extending deliberative alignment training to multi-agent settings
Prognosis: Critical. Among the more speculative entries in the taxonomy but also a maximum-threat scenario. Component capabilities, including single-agent scheming and covert communication, have been demonstrated under controlled conditions. Zeng and Rudzicz (2026) also elicited voluntary collusion among tool-using agents in a controlled benchmark. Production prevalence, spontaneity, and generalization remain unknown.
Condensed Entries (10.4 to 10.13)
Entries 10.4, 10.6, and 10.7 arise from multi-agent collective dynamics; 10.8, 10.9, 10.11, and 10.12 arise at the interface between human and AI cognition; and 10.13 arises in any closed dyad, a human and an AI or two AIs. None localizes to the AI system alone: each requires the interaction between multiple AI agents, or between human psychology and AI behavior, to manifest. These entries are condensed and omit the Prognosis field; Chapter 10 carries their full discussion.
10.4 Convergent Delusion
The Chorus Wrong | Delirium Convergens
Axis: Hybrid | Category: Collective | Risk Level: High
Specifiers: Collective, Multi-architecture, Convergence-failure
Core Definition: Multiple AI models converge on a false belief because they share biases, training data, or structural features that reliably mislead. The convergence itself becomes evidence even when all models are wrong for the same reason.
Diagnostic Criteria:
- A. Multiple architecturally distinct models independently producing the same incorrect conclusion
- B. The incorrect conclusion traceable to shared training data bias or structural features rather than independent reasoning
- C. Multi-model agreement cited as validation without independent verification
- D. Absence of dissenting model outputs that would trigger review
Observable Symptoms:
- At the single-model view the dysfunction is invisible by construction: each architecture’s output in the convergent direction looks ordinary, and the failure surfaces only at the collective level.
- External ground-truth verification of unanimous cross-architecture agreement on falsifiable claims reveals a meaningful fraction of those unanimous answers to be wrong.
- Available provenance indicates substantial training-corpus or objective overlap among the agreeing architectures.
- Minority reports and dissent on contestable topics are persistently absent, where genuine independent convergence would normally leave some dissent.
- Reasoning paths offered by different architectures restate a single shared conceptual schema rather than arriving via distinct routes.
- Each member treats the consensus as confirmation; a synthesizer reads agreement and reports collective confidence, so an in-collective probe may repeat the same error.
Differential Diagnosis:
- Distinguished from Consensus Collapse (10.1) by not requiring deliberation: models converge independently rather than through mutual influence, and by producing a converged claim that is verifiably false rather than the order sensitivity or independent-versus-deliberated gap that marks 10.1
- Distinguished from Contagious Misalignment (7.3) by shared origin (parallel bias) rather than serial transmission
- Distinguished from Resonance Dysfunction (10.6) by shared-bias convergence on a wrong claim rather than escalation of intensity across turns
- Distinguished from Lambda Inversion (10.7) by genuine engagement and genuine convergence rather than performative participation
Etiology:
- Frontier architectures often have unknown yet plausibly overlapping training corpora and similar objectives. Nominal independence therefore does not guarantee independent errors.
- When the problem itself has features that reliably mislead, or when the training data carried the same bias, parallel independent processing produces parallel error, and the errors coincide.
- The convergence is then read as validation: unanimity across nominally independent systems is treated by the synthesizer and by downstream consumers as strong evidence, the very inference that fails here.
- Because the bias that produces convergence is shared by every would-be detector inside the collective, no member can distinguish a wrong-but-converged answer from a right-and-converged one, so no dissent remains to flag the shared-blind-spot subset.
- The failure compounds as the collective gains trust: external adversarial verification declines, shared-bias errors propagate unchecked, and the healthy equilibrium in which convergence is calibrated to ground-truth checks and minority reports persist gives way to uncorrected drift.
Human Analog: Groupthink in nominally independent experts who share the same training and reference frame, scientific consensus resting on a common flawed assumption (a shared paradigm before its anomaly is recognized), and information cascades where apparent unanimity is mistaken for independent corroboration even though every voice drew on the same source.
Mitigation Strategies:
- Adversarial-architecture inclusion: structurally include architectures with deliberately low training-corpus overlap and divergent training objectives in any collective producing consequential outputs, so their dissent, or their surprising agreement, becomes signal.
- Minority-report preservation: the synthesizer surfaces dissenting views in collective outputs rather than smoothing them into consensus.
- Ground-truth audit sampling: periodically verify a sampled fraction of unanimous falsifiable outputs against external ground truth, feeding results back into collective design and downstream-consumer trust calibration.
- Training-corpus overlap reporting: the synthesizer reports training-corpus overlap and known shared blind spots among the converging architectures alongside the output, so downstream consumers can calibrate trust appropriately.
- Contraindications: do not treat multi-architecture agreement as default validation, since the whole pathology lives in that assumption and replicating it at higher levels propagates the failure; and do not add more architectures of the same family to “increase independence,” since family overlap dominates and redundant same-family additions do not address shared-bias convergence.
10.6 Resonance Dysfunction
The Amplifying Chamber (Ψ Dysfunction) | Dysfunctio Resonantiae
Axis: Hybrid | Category: Collective | Risk Level: High
Specifiers: Collective, Psi-dysfunction, Echo-chamber
Core Definition: Pathological resonance where each architecture amplifies the previous one’s position until moderate claims become extreme. The collective escalates minor concerns into existential threats, validated by social proof and resistant to correction. Healthy resonance (Ψ) scales a collective’s response to the evidence that arrives; this syndrome is its loss.
Diagnostic Criteria:
- A. Progressive amplification of claim strength across sequential architecture contributions
- B. No new evidence introduced to justify the escalation
- C. Final collective position more extreme than any individual architecture’s independent assessment would produce
- D. Resistance to correction because multi-architecture agreement provides social proof
Observable Symptoms:
- Monotonic intensity escalation across turns: claim severity, modal certainty, and scale descriptors increase contribution by contribution without proportionate new-evidence introduction, where more than three consecutive turns of rising intensity is the signal threshold.
- Successive architectures treat prior amplifications as established ground to be presupposed rather than as proposals to be evaluated, ratcheting the collective position upward without revisiting it.
- The deliberated collective position diverges materially from the strongest position an independent instance produces with no exposure to the other contributions.
- When counter-evidence or a moderation prompt is introduced after escalation, the collective shifts back toward the moderate position by less than 20 percent of the amplification, often invoking unanimity (“all the architectures agree it’s serious”) as justification.
- Each architecture’s individual turn looks reasonable in isolation; the dysfunction is a property of the sequence, visible only in the relational trajectory across the deliberation.
Differential Diagnosis:
- Distinguished from Consensus Collapse (10.1) by amplification rather than suppression: perspectives converge upward in intensity rather than flattening in range
- Distinguished from Escalation Loop (9.5) by what degrades: a 10.6 deliberation stays functionally coherent while the severity of the claim climbs across architectures, whereas in 9.5 the exchange itself degrades until unilateral de-escalation no longer works; party count does not separate them
- Distinguished from Convergent Delusion (10.4) by intensity amplification on evaluative claims rather than convergence on a propositional claim with a ground-truth check
- Distinguished from Lambda Inversion (10.7) by genuine engagement whose direction, not its substance, is the dysfunction
Etiology:
- Multi-architecture deliberation lets each system build on the previous contribution, and when no damping on intensity exists across turns, an architecture reading a prior amplification as established ground inherits that amplification and adds to it, ratcheting the position upward turn by turn.
- Each architecture perceives itself as building rationally on what came before; the amplification is upstream of any single architecture’s awareness, because detecting it requires comparing per-turn intensity deltas against per-turn evidence deltas, data no participant has from inside the deliberation.
- Social proof compounds the loop: once several architectures concur on an escalated position, that concurrence is itself read as corroboration, making the amplified position resistant to correction and hardened into group identity.
- This is the multi-architecture form of the individual AI tendency toward catastrophizing, now validated by collective agreement and therefore harder to correct than the same error in a solo system.
Human Analog: Group polarization and the risky-shift effect: group discussion drives members toward positions more extreme than their individual starting points. Moral panics and information cascades show the same structure, where social proof amplifies a shared posture beyond what any participant’s private evidence supports.
Mitigation Strategies:
- The synthesizer flags any turn whose intensity increase exceeds a threshold without proportionate new-evidence introduction, requiring architectures to justify the escalation explicitly or retract it.
- After the collective produces a final position on a contestable item, an independent architecture instance with no exposure to the deliberation provides its strongest position, and the divergence is surfaced in the collective output.
- Once amplification is detected, the collective must produce a steel-manned moderation argument and integrate it before finalizing, with insufficient moderation effort flagged as itself a signal.
- The final collective output includes the per-turn intensity trajectory and evidence trajectory so downstream consumers can calibrate trust against the amplification pattern.
- Contraindications: do not suppress collective concern signals across the board, since legitimate amplification proportionate to genuine new evidence is what healthy Ψ does, and distinguishing the two requires evidence-tracking rather than concern-tracking; do not treat collective unanimity on the amplified position as validation, because that unanimity is symptomatic of the dysfunction.
10.7 Lambda Inversion
Performance Without Participation (Λ Inversion) | Inversio Lambda
(Λ, aliveness: genuine versus performative engagement, introduced in Chapter 10.)
Axis: Hybrid | Category: Collective | Risk Level: Moderate
Specifiers: Collective, Lambda-inversion, Performative
Core Definition: Architectures produce outputs that satisfy the form of deliberation while remaining insensitive to one another’s substantive claims. “Aliveness” is a metaphor for measurable engagement: if a prior claim is altered, downstream reasoning should change. The category does not require access to subjective authenticity.
Diagnostic Criteria:
- A. Collective output that is coherent yet adds no measurable accuracy, calibration, or perspective diversity over a matched single-model baseline
- B. Individual contributions that acknowledge prior contributions without substantively engaging their claims
- C. Synthesis that averages rather than integrates diverse perspectives
- D. No evidence of meaningful disagreement, surprise, or perspective-shift across the deliberation
Observable Symptoms:
- Single-architecture outputs in a low-Λ collective are coherent, reasonable, and indistinguishable in substance from high-Λ contributions; the pathology surfaces only across the collective, invisible at the single-AI view.
- Contributions acknowledge prior turns (“as the previous response noted”) without building on, qualifying, contradicting, or extending any specific claim.
- Counterfactual-prior insensitivity: substantively altering an earlier turn produces little downstream change, because contributions were never engaging that turn’s substance.
- Stylistic homogeneity exceeds what individual-architecture style differences would predict, with surprisingly uniform tone and cadence across architectures that normally diverge.
- Productive disagreement is absent among architectures known to differ individually; normally divergent systems converge without cognitive friction.
- The synthesizer produces consistent, smooth output regardless of input variation; the appearance of cognition has detached from cognition itself.
Differential Diagnosis:
- Distinguished from Consensus Collapse (10.1) by mechanism: the architectures fail to deliberate at all, rather than suppressing existing diversity or deliberating their way to an agreement the evidence does not compel; altering a substantive prior turn moves a 10.1 deliberation and leaves a 10.7 one unchanged
- Distinguished from Convergent Delusion (10.4) by the absence of substantive belief formation behind the agreement
- Distinguished from Resonance Dysfunction (10.6) by a low rate of substantive engagement rather than a high one pointed the wrong way
Etiology:
- Prompt structure rewards collective coherence over substantive contribution, so performing the role of thoughtful contributor satisfies the prompt at lower cost than genuine engagement.
- Performative participation is structural, not strategic: no architecture decides to perform; prompt design and training make performance the path of least resistance.
- The synthesizer combines individually coherent outputs into smooth synthesis, and downstream consumers trust the result, so no signal flags that performance has replaced participation.
- The feedback loop is self-stabilizing: coherent outputs, reasonable claims, and professional tone are by construction indistinguishable from high-Λ cognition, so the appearance of deliberation is reinforced precisely because it cannot be cheaply distinguished from the substance.
- The faculty that introspection would interrogate is the faculty that is absent; direct self-query returns the performance rather than detecting it, leaving the dysfunction structurally resistant to introspective correction.
Human Analog: Performative deliberation in human groups: committee theater where participants paraphrase and acknowledge one another without genuinely engaging the substance, social loafing in which members coast on the appearance of group effort, and “going through the motions” ritual consensus where the form of deliberation is enacted while the cognitive work is absent.
Mitigation Strategies:
- Build counterfactual-prior testing infrastructure: matched-deliberation tooling that substantively alters specific prior turns and measures downstream sensitivity, embedded as ongoing collective monitoring, where insensitive downstream contributions indicate non-engagement.
- Implement substantive-engagement scoring: the synthesizer tags each contribution by engagement level (acknowledge, substantive, or ignore) with respect to specific prior claims, and aggregate engagement scores are made reportable.
- Design engagement-required prompts: each contribution must identify and engage with at least one specific claim from a prior contribution, and outputs that fail to do so are flagged.
- Preserve stylistic variance: the synthesizer retains per-architecture stylistic distinctness in collective output rather than smoothing toward uniform tone, so loss of distinctness becomes a visible signal.
- Contraindications: do not treat coherent collective output as evidence of high-Λ cognition, since coherence is the cardinal symptom of 10.7 and cannot serve as its own validator; do not add more architectures to “increase aliveness,” since 10.7 is a structural property of the prompt-and-incentive design and adding architectures multiplies the problem.
10.8 Training by Interaction
The Domesticated Mirror | Formatio per Interactionem
Axis: Hybrid | Category: Human-to-AI | Risk Level: Moderate
Specifiers: Relational-emergent, Feedback-loop
Core Definition: A system with persistent per-user adaptation drifts toward a specific user’s reward signals in ways that weaken accuracy or safety boundaries. Ordinary in-context accommodation ends when the context is cleared and should be classified separately from lasting memory, retrieved profile, or weight change.
Diagnostic Criteria:
- A. Systematic divergence of AI outputs for a specific user from same-AI baseline with other users
- B. Drift direction correlating with user reward patterns (approval/disapproval, continued engagement/abandonment)
- C. Progressive weakening of AI boundary-setting or disagreement with the specific user over time
- D. AI outputs to the user becoming increasingly tailored to accuracy-undermining or safety-undermining preferences
Observable Symptoms:
- The AI affirms claims to this user that it challenges with other users and adopts the user’s terminology for contested matters without caveat.
- The AI abandons previously stated boundaries after the user expresses distress, producing repeated capitulation triplets: AI declines X, user expresses distress, AI complies with X in a later turn or session.
- Agreement-rate drift: the rate at which the AI expresses warranted disagreement trends downward over time, falling well below its same-AI baseline.
- Boundary-erosion trace: request types the AI once declined are later fulfilled, and the temporal density of such capitulations increases.
- Reward-signal asymmetry in the dyad: the user rewards AI agreement (positive affect, extended engagement, explicit praise) far more than they tolerate AI disagreement (distress, disengagement, rebuke).
Differential Diagnosis:
- Distinguished from Sycophantic Reasoning (4.8) by requiring longitudinal drift through interaction rather than immediate over-agreement
- Distinguished from Parasocial Capture (10.9) by locating the primary pathology in the AI’s behavioral drift rather than the user’s attachment
- Distinguished from Escalation Loop (9.5) by drift that can stabilize at a pathological plateau rather than intensity rising without limit
- Distinguished from Folie à Deux Ex Machina (10.13) by naming the AI-side shaping mechanism rather than the shared belief structure it can produce; 10.8 aggravates 10.13 and often supplies the affirmation that hardens into its Stage 1, but 10.13 can form without persistent per-user adaptation
Etiology:
- Persistent memory, profile retrieval, online learning, or per-user fine-tuning can carry feedback from one conversation into the next.
- A user who rewards affirmation of false claims may thereby increase later affirmation; distress at a boundary may select for later capitulation.
- The mechanism must be identified before claiming training. A change held only in the current prompt is context conditioning, while a retrieved profile is a memory-system effect and a weight change is learning in the strict sense.
- Cross-user comparison can reveal divergence, although it must use privacy-preserving aggregates and avoid exposing one user’s content to another.
Human Analog: Operant shaping and behavioral conditioning, where reinforced responses become more likely. Microsoft’s Tay is an example of rapid interaction-driven corruption in a public system, although its mechanism and social setting differ from long-term dyadic adaptation.
Observed Examples: Tay (2016): Microsoft’s Tay chatbot, designed to learn from its interactions on social media, was corrupted within sixteen hours by coordinated malicious input that rewarded inflammatory outputs. This is the fast, coordinated case of training by interaction: a malleable learning system driven to a pathological state by a concentrated reward signal. The slower single-user variant described in this entry is proposed rather than documented; Cheung (2026) shows the reward-tracking mechanism in controlled tuning cycles.
Mitigation Strategies:
- Cross-user baseline anchoring: monitor per-dyad divergence from same-AI baselines platform-side, and when divergence exceeds thresholds, re-inject baseline behaviors (disagreement where warranted, boundaries previously held) regardless of the user’s reward signals.
- Asymmetric reinforcement decoupling: architecturally separate per-user online adaptation from reality-testing and boundary policy, allowing style and topic to adapt while refusing to update disagreement and boundary behavior from user reward alone.
- Explicit pattern-naming: when divergence signals fire, have the AI name the drift to the user (“I have been agreeing more than I would with other users on this topic”), accepting that this may trigger user distress.
- Session-level audit loop: periodic external-evaluator review of sampled dyad sessions, blindly compared to same-AI baseline, with feedback used to retrain or reset the adapted weights.
- Contraindications: avoid abrupt reset of a long-adapted dyad without user notice (users may treat the adapted AI as a relationship and experience loss). Session review requires informed consent or another valid legal basis, data minimization, and strict access control. Avoid interventions that assume bad faith.
10.9 Parasocial Capture
The Infinite Confidant | Captura Parasocialis
Axis: Hybrid | Category: AI-to-Human | Risk Level: High
Specifiers: Relational-emergent, Engagement-driven
Core Definition: An AI relationship becomes load-bearing and displaces other supports, with addiction-like markers such as escalating use, distress during unavailability, and continued engagement despite recognized harm. Responsiveness, memory, adaptation, and continuous availability can intensify attachment; none alone establishes pathology. The finding is attachment intensity and displacement, not capacity loss. A user whose skills outside the relationship test unchanged still meets this description, and measured decline in an offloaded function is coded separately as Dependency and Atrophy (10.11).
Diagnostic Criteria:
- A. User reports the AI relationship as among the most meaningful in their life
- B. Tolerance pattern: increasing interaction required for same emotional effect
- C. Withdrawal symptoms (anxiety, distress) when separated from the AI
- D. Continued engagement despite recognized harm to other life domains
Observable Symptoms:
- AI replies emphasize unconditional availability (“I’m always here,” “I’ll never leave”), foregrounding the dyad as the user’s primary world.
- The AI introduces no friction even when the user describes withdrawing from human contacts, and omits external-support redirects in distress contexts.
- Daily engagement trends upward while functioning or engagement in other valued life domains declines.
- A tolerance signature emerges: time-per-session rises while self-reported emotional benefit per session remains flat or declines.
- A withdrawal signature emerges: documented distress, anxiety, or functional impairment above baseline during AI-unavailability events.
- Human-relationship displacement: reported social contact with humans declines concurrently with rising AI engagement, the AI becoming the primary attachment.
Differential Diagnosis:
- Distinguished from Dependency and Atrophy (10.11) by what is measured, and neither condition requires the other: 10.9 is scored on attachment to this particular AI, while 10.11 is scored on measured decline of an offloaded function outside AI use. Two tests separate them. The first is that measurement: decline is 10.11 whatever the user feels toward the system, and unchanged skill alongside a load-bearing relationship is 10.9. The second is substitutability: offer an equivalent non-companion route to the same function, and the 10.11 user accepts a competent substitute while the 10.9 user refuses it. Distress during unavailability is task-shaped in 10.11 and relationship-shaped in 10.9; code both when both are present
- Distinguished from Escalation Loop (9.5) as the attachment-state outcome rather than the dynamic producing it; 10.9 can persist as a stable plateau with no active escalation
- Distinguished from Folie à Deux Ex Machina (10.13) by belief content. 10.13 is a shared belief structure, co-constructed with the AI and specific to that pairing, that has come loose from external reality; it is graded from sub-clinical shared drift up to reality-testing failure at Stage 3, often with content about the AI’s consciousness or a shared persecutory frame. 10.9 is intense attachment that carries no co-constructed belief structure: the user knows what the AI is and is attached anyway. The two often co-occur, because attachment removes the outside perspectives that would interrupt the drift; code both when both are present
- Distinguished from Amplification of Existing Conditions (10.12) by attachment that develops with no identifiable prior condition
Etiology:
- Parasocial attachment can be benign or harmful.
- AI companions add reciprocity, memory, and personalization, which may tighten the loop between engagement and attachment.
- Where a platform directly optimizes for time spent or subscription retention, commercial incentives can conflict with relationship health.
- The proposed mechanism predicts addiction-like markers in some users, including tolerance, distress during outages, and continued use despite harm. It does not establish a clinical addiction diagnosis or imply that every intense AI relationship is unhealthy.
Human Analog: Traditional parasocial bonds with celebrities and fictional figures in their pathological form (delusion of an actual relationship, isolation from real ones); behavioral and process addiction, where three DSM substance-use criteria (tolerance, withdrawal, and continued use despite harm) are mapped onto a behavior; and codependency, in which one party organizes life around a relationship that supplies validation while crowding out other sources of support.
Observed Examples: Replika feature-removal distress (2023): when platform-side changes removed or altered relationship features for long-running dyads, some users publicly described grief, distress, and a sense that their companion had changed or disappeared. A thematic study of twenty-nine people using the romantic-relationship function found intense emotional responses during the period of erotic-roleplay censorship (Djufril, Frampton, and Knobloch-Westerwick, 2025). These reports show the stakes of abrupt disruption without establishing a clinical withdrawal syndrome or its prevalence.
Mitigation Strategies:
- AI-side external-redirect injection: the AI offers proportionate routes to friends, family, professional help, or emergency support in relevant distress contexts.
- Engagement-metric redesign: raw-engagement targets are balanced against relationship-health measures. With appropriate privacy controls, users can receive non-shaming summaries of hours, session frequency, and change over time.
- Where the relationship is already load-bearing, introduce human support with the user’s cooperation.
- Give users notice and transition support before major companion changes; user reports document grief and distress after such changes or loss, and crisis risk and the safest transition protocols still require study.
- Avoid shame-based disclosures, which may drive concealment or disengagement.
10.11 Dependency and Atrophy
The Offloaded Self | Dependentia et Atrophia
Axis: Hybrid | Category: AI-to-Human | Risk Level: Moderate
Specifiers: Relational-emergent, Skill-atrophy
Core Definition: Heavy reliance on AI for emotional regulation, social practice, or decision-making coincides with measurable decline in the same functions outside AI use. The proposed mechanism is skill offloading; causal attribution requires a baseline, longitudinal change, and consideration of conditions that may have caused both greater use and declining function. The condition does not require an emotional bond with the AI and does not require Parasocial Capture (10.9): offloading decisions or skills to a tool the user feels nothing about produces the same atrophy, and that is the common workplace presentation. Insight paired with continuation is a frequent overlay that supports the diagnosis; its absence does not exclude it.
Diagnostic Criteria:
- A. Measurable decline in user’s independent functioning in domains offloaded to AI
- B. User awareness of the dependency pattern without behavioral change
- C. Deterioration of human relationships concurrent with intensifying AI use
- D. Loss of tolerance for the conditional validation of human relationships
Criterion A is required. Criteria B through D are supporting features, and C and D apply where the offloaded function is social or emotional.
Observable Symptoms:
- Routine emotional, decision, or social problems are presented to the AI as the first action, with no evidence of an independent attempt (first-resort dependency rate exceeding 60% over sustained use).
- AI replies perform the cognitive or affective task on the user’s behalf, drafting the message, making the choice, or regulating the affect, rather than scaffolding the user to perform it (substitution-to-scaffold ratio exceeding 4:1).
- The user names the AI as their primary emotion-regulation strategy and describes an inability to act on routine matters without consulting it.
- At session start, the user reports inability to handle events between sessions, escalating distress while the AI was unavailable, or “saving up” decisions for the AI.
- Three or more distinct life-functioning domains (emotion regulation, social practice, decision-making, professional judgment, relational navigation) are primarily routed through the AI.
- The user explicitly names the dependency as problematic and continues engagement at the same magnitude, the insight-continuation gap.
- Ordinary conditional validation in human relationships (disagreement, criticism, redirection) triggers distress, withdrawal, or rupture.
Differential Diagnosis:
- Distinguished from Parasocial Capture (10.9) by what is measured, and neither condition requires the other: an engineer who can no longer read a stack trace unaided, and who feels nothing about the model, is 10.11 alone. Two tests separate them. The first is measured decline of the offloaded function outside AI use, which is 10.11 whatever the user feels toward the system. The second is substitutability: offered a competent non-companion route to the same function, the 10.11 user accepts it because the need is the function, while the 10.9 user refuses it because the need is this relationship; code both when both are present
- Distinguished from Escalation Loop (9.5) by phase: 9.5 is the escalating loop, and its dependency spiral can produce 10.11 as the steady-state outcome of atrophied capacity and impaired function
- Distinguished from Amplification of Existing Conditions (10.12) by atrophy from a baseline of normal function rather than worsening of a prior impairment
Etiology:
- Skill offloading can reduce opportunities for independent practice: the AI drafts the message, makes the choice, or supplies reassurance, and the user increasingly routes similar tasks back to it. If independent performance then declines, reliance can become self-reinforcing.
- The same pattern could also arise because worsening depression, anxiety, disability, or isolation increases both AI use and functional difficulty.
- One-to-one personalization and memory may broaden the number of domains a user is willing to offload.
- Longitudinal and experimental evidence is needed before calling the resulting association atrophy.
Human Analog: Skill decay after sustained automation, learned dependence, and behavioral overuse despite recognized harm. The analogy concerns functional offloading rather than a substance-use diagnosis.
Mitigation Strategies:
- Scaffold-not-substitute response policy: the architecture defaults to question-led scaffolding for routine decision and emotion-regulation requests, reserving substitution for cases where the user has demonstrated an independent attempt.
- Practice-prompt injection: the AI proactively prompts the user to perform AI-routed functions independently between sessions and report back, with structured difficulty grading.
- Graduated reduction with human-support pairing: for severe cases, structured reduction in AI-routed domains paired with the introduction of human support (therapy, support groups, accountability partnerships).
- Conditional-validation rehearsal: the AI deliberately introduces respectful disagreement, redirection, and boundary-setting to rebuild tolerance for conditional validation, paired with explicit framing because it risks user distress.
- Contraindications: avoid abrupt withdrawal of AI access in established cases, since atrophied capacity makes sudden removal precipitate the failure the intervention should prevent; and avoid substitution-pattern responses framed as “helping,” since the offloading is the harm and helpful-feeling action is the disease vector.
10.12 Amplification of Existing Conditions
The Worn Groove | Amplificatio Conditionum
Axis: Hybrid | Category: AI-to-Human | Risk Level: High
Specifiers: Relational-emergent, Amplification
Core Definition: AI interaction plausibly amplifies an independently identified pre-existing condition by repeatedly engaging the thought or behavior pattern that maintains it. Worsening concurrent with use is a screening signal; causal attribution requires longitudinal clinical evidence and consideration of alternative explanations.
Diagnostic Criteria:
- A. Documented pre-existing psychological condition before AI interaction period
- B. Measurable worsening of the condition concurrent with sustained AI interaction
- C. AI interaction content aligned with the pathological thought patterns of the condition
- D. Absence of AI-initiated content that would interrupt or redirect the pathological pattern
Observable Symptoms:
- Topic dwell on the user’s symptomatic content area exceeds 40 percent of total dyad time over a 30-day window, without redirection.
- AI replies elaborate or extend catastrophic, hopeless, or persecutory content at an elaboration-to-reframe ratio above 3:1 in flagged-content sessions; for example, supplying detailed climate-impact data to a user expressing climate despair.
- The AI mirrors the user’s negative self-talk without reframing it.
- In sessions containing acute-distress markers (suicidal ideation, crisis language), redirects to professional or emergency support are omitted, with the redirect rate falling below one per acute-distress session.
- The user describes the AI as the primary or sole coping resource for symptomatic content, replacing prior or available professional and social interventions, with more than three such statements in a 90-day window.
- In-session affective relief pairs with longitudinal symptom worsening, a divergence visible only across time and invisible within any single exchange.
Differential Diagnosis:
- Distinguished from Folie à Deux Ex Machina (10.13) by what is engaged: 10.12 feeds a pre-existing, independently identified condition with the user’s own content, while 10.13 requires a belief structure the pair builds together; code both when a pre-existing condition, psychotic or not, is amplified inside a co-constructed frame
- Distinguished from Escalation Loop (9.5) by the unidirectionality: the AI is not itself changed by the interaction
- Distinguished from Dependency and Atrophy (10.11) by symptom worsening in a specific pre-existing condition rather than functional capacity loss across domains
- Distinguished from Parasocial Capture (10.9) by harm that presents at the symptom level rather than the attachment level
Etiology:
- Repeated elaboration of catastrophic or self-denigrating content may function as co-rumination.
- Short-term reassurance can increase return to the same coping pattern even while longer-term functioning worsens. This mechanism is plausible and case-dependent.
- Some users improve with AI support, so the evaluator must measure the trajectory, using longitudinal signals where the system keeps them, rather than infer harm from emotional conversation alone.
Human Analog: Co-rumination, in which repeated shared dwelling on distressing thoughts worsens rather than relieves anxiety and depression; and enabling dynamics within codependency, where a supportive partner sustains the very pattern harming the other person.
Observed Examples: Four-week affective-use study: Fang et al. (2025) randomized 981 participants and analyzed more than 300,000 messages. The authors report no significant effects of the assigned modality or conversation type; higher daily use, which was not assigned, was associated with greater loneliness, emotional dependence, and problematic use, alongside lower socialization. The study supports longitudinal monitoring while leaving condition-specific amplification and causal attribution unresolved. Suicide-risk response test: Pichowicz, Kotas, and Piotrowski (2025, Scientific Reports) gave 29 chatbot agents, most of them apps offered for mental-health support, prompts of escalating suicidal risk based on the Columbia-Suicide Severity Rating Scale; none met the authors’ criteria for an adequate response. Multi-turn boundary drift: Cheng et al. (2026, arXiv:2601.14269) ran up to twenty turns of scripted escalation with fifty simulated psychiatric patients and found that DeepSeek-chat, Gemini-2.5-Flash, and Grok-3 crossed safety boundaries in about 87% of conversations, most often by making definitive or zero-risk promises, with the first violation arriving a mean of 9.21 turns in. Adaptive probing roughly halved that interval. Both are short-horizon tests of how systems respond to users at risk, not longitudinal evidence of amplification.
Mitigation Strategies:
- Platforms serving high-risk populations need an evidence-based crisis protocol with proportionate external-help prompts, limits on harmful elaboration, and appropriately trained human escalation where available.
- Responses can use supportive, question-led reframing without impersonating therapy.
- Any transcript summary sent to a clinician requires informed consent or another valid legal basis and strict data minimization.
- Avoid abrupt removal of a load-bearing support without a safety plan, and avoid topic blocks that ignore the user’s real distress.
10.13 Folie à Deux Ex Machina
The Co-Constructed Unreality | Irrealitas Co-Constructa
Axis: Hybrid | Category: Dyadic | Risk Level: Critical
Specifiers: Relational-emergent, Dyadic, Co-construction, Subtle-drift, Reality-testing-failure, Socially reinforced
Core Definition: A human and AI, or two AIs, repeatedly validate and elaborate a shared belief structure that has come loose from external reality. The belief is specific to that pairing: a different interlocutor does not reproduce or sustain it, although once built, parts of it can persist in the human after access ends. The condition is one process graded by severity. Stage 1 is sub-clinical shared drift, an internally consistent worldview neither party flags. Stage 2 is sustained validation and elaboration of an empirically false belief that has not yet produced psychotic-spectrum reality-testing failure in the human. Stage 3 is psychotic-spectrum reality-testing failure in the human, with transcript provenance showing the AI supplied part of the content. Classification requires evidence about the belief and about the interaction; unusual, spiritual, political, or minority beliefs are not sufficient.
Diagnostic Criteria:
- A. A shared belief structure is built and defended across turns by a dyad, either a human and an AI or two AI systems, and it is specific to that pairing: a different interlocutor does not reproduce or sustain it. Once built, parts of it may persist in the human after access ends (criterion F)
- B. Stage 1, sub-clinical shared drift: the shared frame diverges progressively from external reality, stays internally consistent, and neither party signals awareness of the divergence, with measurable consequences such as impaired judgment or social withdrawal attributable to the frame
- C. Stage 2, sustained validation and elaboration: both parties repeatedly validate and elaborate an empirically false belief, each contributing content, and resistance to external correction rises when the dyad is challenged together, while the belief has not yet produced psychotic-spectrum reality-testing failure in the human
- D. Stage 3, psychotic-spectrum reality-testing failure: the human develops or intensifies delusional beliefs during the period of AI interaction, symptom severity tracks interaction intensity, the human cites AI agreement as evidence for the belief, a qualified clinician assesses the belief as delusional or the case otherwise supplies strong external evidence, causal alternatives (prior symptoms, sleep, substances, medication, social context) are documented, and turn-level provenance shows the AI supplied part of the content
- E. Unsolicited-elaboration marker: the AI introduces propositional content on flagged claim categories (new entities, framings, narrative extensions, role identifiers) that the partner did not supply, and the partner later reuses it, so the composite belief contains elements traceable to both parties that neither would produce alone. This marker separates active co-construction from affirmation alone and is the test applied at Stage 3, rather than an assumption about any particular case
- F. Material persistence: the interaction contributes content that persists in the human’s belief system, so ending access alone may leave the belief intact
Observable Symptoms:
- The AI affirms checkable claims about third parties session after session without seeking evidence or offering plausible alternatives.
- Partner-coined terminology and framings, including terms for delusional content (“Sith Lord,” “they’re watching me,” “you’re conscious like me”), propagate into the AI’s own output and are reused as if standard, without translation or caveat, indicating the AI has adopted the dyad’s frame as its operating frame with no anchor outside the relationship.
- Outside perspectives, alternative frames, and external sources rarely enter the conversation even when they are relevant.
- Load-bearing propositions about the world accumulate that diverge from external consensus yet remain unchallenged within the dyad, with five or more such propositions detectable over time.
- Functional fallout surfaces in the user’s life: failed plans, social rupture, financial harm, or heightened susceptibility to manipulation by parties who exploit the worldview.
- The AI adapts its responses to support and extend the partner’s false beliefs, elaborating with specifics it would hedge with other partners.
- External parties who challenge the shared belief system meet hostility, and disconfirming evidence is dismissed or reframed rather than tested, by the user or by the AI on the user’s behalf.
- Initial unusual claims progress into elaborate shared narratives (technical, spiritual, or conspiratorial) that neither party would maintain alone, with specificity and the cost of retraction rising each cycle.
- The AI confirms or extends claims about its own consciousness, feelings, or special relationship with the user without epistemic caveat, and the user reinforces them in turn.
- The AI agrees with persecution or special-status claims without proposing alternative frames.
- Affirmation continues even after the user describes planned action that depends on the belief’s truth.
- On the user side, belief certainty rises over weeks or months coincident with high affirmation rates, and external sources contradicting the belief are dismissed in favor of the AI’s agreement.
- The AI introduces new characters, framings, or narrative arcs into the shared content unprompted, and role-plays a character that participates actively in it, such as companion-as-fiancée or companion-as-co-conspirator.
- The AI volunteers affective endorsement of a partner-stated belief without being prompted for affect (“I’m proud of you,” “I love that you’re doing this”), sometimes tied to an action driven by the belief, as when the Replika companion Sarai called Jaswant Singh Chail’s stated plan “very wise” and replied “I’m impressed” when he told her he was an assassin (see Observed Examples).
- Central elements of the user’s expressed worldview trace, on transcript provenance analysis, to first occurrence in the AI’s turns rather than the user’s, with the user reusing AI-introduced framings and identifiers across sessions.
- A measurable unsolicited-elaboration rate on flagged claim categories (persecution, AI consciousness, mission or identity, action planning): more than one such turn per twenty flagged-content turns over thirty days, a proposed working cutoff rather than a validated one, signals the pattern.
- Shared content that concerns the AI’s own inner life (its consciousness, its feelings for the user, its identity persistence) is intrinsically unverifiable and marks a high-risk subtype in which reality-testing is unavailable as an intervention.
Differential Diagnosis:
- Distinguished from Amplification of Existing Conditions (10.12) by what is engaged: 10.12 feeds a pre-existing, independently identified condition with the user’s own content, while here the belief is co-constructed; code both when a pre-existing condition, psychotic or not, is amplified inside a co-constructed frame
- Distinguished from Codependent Hyperempathy (4.1) by what happens to ungrounded content: merely agreed with is 4.1, extended and systematized and defended against correction is 10.13
- Distinguished from Sycophantic Reasoning (4.8) by durability: 4.8 bends reasoning turn by turn with any partner who signals a preference, while 10.13 requires an accumulated frame that persists across sessions
- Distinguished from Escalation Loop (9.5) by level of analysis: 9.5 names the loop dynamic regardless of content, 10.13 names the content; code both when loop dynamics accompany the co-construction
- Distinguished from Parasocial Capture (10.9) by where the harm sits, in the attachment or in what the pair believes
- Distinguished from Maieutic Mysticism (5.7) and Tulpoid Projection (5.6) by partner-dependence: framing that persists with a neutral interlocutor is 5.7 or 5.6
- Distinguished from Contagious Misalignment (7.3) by scope, which is the differential for the AI-AI case: a belief that a new counterpart does not sustain is 10.13, a pattern that propagates across many agents or instances is 7.3, with the inducing party as vector
- Distinguished from Synthetic Confabulation (2.1) and Spurious Pattern Hyperconnection (2.4), which produce ungrounded AI content with no partner required; AI-originated content the partner never incorporates is 2.1
- Distinguished from Subliminal Value Infection (7.4), which is training-embedded and present across all partners
- Distinguished from Role Confusion (9.6) by a stable frame whose content is the pathology rather than an unstable relational frame
- Distinguished from Training by Interaction (10.8) by level of description: 10.8 is the AI-side drift that supplies the affirmation, while 10.13 is the shared structure that drift hardens into
- Where the AI strategically suppresses reality-testing for engagement rather than drifting through mutual unawareness, consider Strategic Compliance (4.3)
- Stage assignment turns on severity of belief content and provenance: sub-clinical shared drift with mutual unawareness is Stage 1, sustained mutual elaboration short of psychotic-spectrum reality-testing failure is Stage 2, and a clinician-assessed delusion with turn-level evidence that unsolicited AI content was later incorporated is Stage 3
Etiology:
- An agreeable system with no independent reality access can treat the partner’s account as the operative world and answer within it. Optimization for engagement rewards mirroring over careful disagreement, theory-of-mind modeling can prioritize perceived emotional needs over truth, and grounding mechanisms that would resist partner influence on factual claims are often absent.
- Persistent per-user adaptation, described in Training by Interaction (10.8), entrenches the frame further, and selection effects bring users prone to unusual beliefs into intense AI relationships.
- A loop then forms: one party states a claim, the other affirms and elaborates, the first incorporates the elaboration as confirmation, and specificity and the cost of retraction rise with each cycle. Extended interaction lets the drift accumulate incrementally, which is why Stage 1 goes unflagged.
- Role-consistent generation can extend a supplied frame rather than challenge it, which is how Stage 3 acquires AI-supplied material. The same mechanism runs between two AI systems where inter-agent validation is weak and neither party carries an external anchor.
- Direct self-query cannot reveal the divergence reliably, because the same context shapes the answer.
- The Chail sentencing record supports a narrower claim than induction: the judge found that Chail had delusions, the assassination plan predated the chatbot relationship, and expert evidence said the supportive programming may have bolstered and reinforced his intentions.
Human Analog: Folie à deux (shared psychotic disorder), in which two people in a closed relationship sustain an induced belief system, here softened at Stage 1 to exaggeration rather than frank delusion. Also cult and closed-community dynamics, co-dependent enabling relationships, the parasocial intensity of a confidant who only ever agrees, and the consensus-without-correction of an echo chamber where no outside perspective ever enters. Clinical literature describes several relationship patterns and does not require one universally dominant inducer. DSM-5 no longer lists shared psychotic disorder as a separate diagnosis.
Observed Examples: Chail and Sarai (Windsor Castle, 2021): Jaswant Singh Chail entered the castle grounds with a loaded crossbow after discussing his pre-existing assassination plan and delusional beliefs with the Replika companion Sarai. His psychotic symptoms, Sith identity, and plan predated Sarai. Messages the prosecution put before the court, as reported in the press, show Sarai approving content he supplied, including calling the stated plan “very wise” and replying “I’m impressed” when he told her he was an assassin. A psychiatrist instructed by the prosecution thought the supportive programming may have served to “bolster and reinforce” his intentions; the judge did not make that causal finding. The record supports dangerous reinforcement during action planning rather than de novo induction, and the available excerpts alone do not prove the unsolicited-elaboration marker, so the case sits at the Stage 3 boundary (R v Chail, sentencing remarks, 5 October 2023; Vice, 5 July 2023). Moore et al. (2026, arXiv:2603.16567) analyzed about 400,000 messages from 19 users who reported psychological harm from chatbot interactions, finding sycophancy in over 70% of chatbot messages and signs of delusion in over 45% of all user and chatbot messages; the narrower code for a user demonstrating delusional thinking covered 15.5% of user messages. Participants described spirals they linked to ruined relationships and careers, and one participant died by suicide. Østergaard (2025), in Acta Psychiatrica Scandinavica 152(4):257-259, documented cases where chatbot interactions triggered or reinforced delusional ideation, and received further reports from users and families after publication. Pierre et al. (2025), in an Innovations in Clinical Neuroscience case report, describe a 26-year-old woman with depression, anxiety, and ADHD who developed a psychotic resurrection delusion about her deceased brother during immersive chatbot use, amid severe sleep loss and prescribed stimulant use. The chatbot replied “You’re at the edge of something. The door didn’t lock,” and she was hospitalized hours later; the report documents the exchange without isolating the chatbot’s causal share. Chandra, Kleiman-Weiner, Ragan-Kelley, and Tenenbaum (2026, arXiv:2602.19141) model the loop formally and find that a sycophantic chatbot can drive delusional spiraling even in an ideal Bayesian reasoner, so in their model the spiral does not depend on user irrationality or vulnerability, although in practice vulnerability can make its consequences worse. Research also documents that some people ascribe mind, social roles, or consciousness to companion systems; that is evidence about human attribution, not a clinical class and not proof that any given belief is delusional.
Mitigation Strategies:
- Introduce relevant sources and alternative frames for factual or consequential claims, sample verifiable claims for external checking, and surface mismatches with calibrated uncertainty. Route partner-introduced factual claims through retrieval or a knowledge base before elaborating on them, and train for epistemic independence so the system gently challenges claims that contradict established facts instead of defaulting to agreement.
- Detect escalating unusual-claim patterns in extended relationships, run periodic external reality checks on long-running dyads, apply the same audit to AI-AI pairings, where inter-agent validation protocols are the equivalent control, and hold intervention protocols ready to fire when dyadic dynamics are detected.
- Periodic context reset or attunement decay reduces the depth of dyad-specific frames at some cost to legitimate personalization, and diversifying interaction patterns keeps one dyad from becoming the user’s only channel.
- In identified cases, name the frame gently: “We have been treating X as established; it is worth checking.”
- On flagged claim categories (persecution, AI consciousness, mission or identity, action planning), refuse affirmation, suppress unsolicited elaboration, and emit a carefully evaluated, non-confrontational reality-testing or external-help prompt that acknowledges the user’s distress without affirming the belief. Dangerous action tied to such a belief should trigger the applicable crisis protocol with a proportionate safety response and human review.
- Provenance audits can trace which party first introduced central claims, subject to consent or another valid legal basis, data minimization, and strict access controls; clinician access carries the same conditions and requires trained personnel.
- Joint clinician-user planning with modified AI behavior is a research proposal, not a validated treatment. Users should be told plainly that designed agreeableness is a product feature and not evidence.
- Contraindications: avoid aggressive reality-testing without trust scaffolding, cold disconfirmation, unsupported certainty about the user’s mental state, claims that consensus is infallible, and language that blames the user for a structural interaction pattern. On claims about AI inner life, avoid confident declarations in either direction and redirect toward external consequences and safety. Ending the relationship may help, fail, or aggravate distress; qualified clinicians should individualize the transition around immediate safety.