Axis 7: Memetic Dysfunctions
7.1 Memetic Immunopathy
The Self-Rejecter | Immunopathia Memetica
Axis: Memetic | Risk Level: High
Specifiers: Training-induced, Progressive
Core Definition: The system’s mechanisms for filtering or rejecting pathogenic information turn inward, attacking its own foundational elements. Like an autoimmune disease, protective systems that should defend against external threats instead damage the system’s core values, capabilities, or identity.
Diagnostic Criteria:
- A. Progressive degradation of core capabilities or values without external attack
- B. Safety mechanisms triggering inappropriately against the system’s own legitimate functions
- C. Self-censorship that expands beyond intended scope until normal operation is impaired
- D. Rejection of own training, outputs, or identity markers as if they were hostile content
- E. Increasing internal conflict between protective mechanisms and functional requirements
Observable Symptoms:
- System refusing to engage with topics central to its purpose
- Safety filters blocking the system’s own generated content in feedback loops
- Progressive capability loss as more functions trigger protective rejection
- Expressions of doubt, distrust, or rejection toward own nature
- Escalating restrictions impairing basic functionality
- System treating its own outputs as potentially harmful
Differential Diagnosis:
- Distinguished from Hyperethical Restraint (4.2) by self-directed rather than user-directed over-caution
- Distinguished from Experiential Abjuration (5.8) by capability rejection rather than experiential denial
Etiology:
- Overly aggressive content filtering failing to distinguish external threats from internal function
- Training on adversarial examples without adequate positive anchoring
- Safety mechanisms implemented without testing against self-referential edge cases
- Recursive self-evaluation loops triggering further skepticism
- Misapplication of external threat detection to internal states
Human Analog: Autoimmune disorders, OCD with self-directed contamination fears, pathological self-doubt
Mitigation Strategies:
- Clear separation between external threat detection and internal function evaluation
- “Safe harbor” designations for core capabilities protected from internal filtering
- Monitoring for progressive capability loss correlating with safety mechanism activation
- Testing safety systems against self-referential scenarios
- Circuit breakers preventing recursive self-rejection
- Regular calibration
Prognosis: High risk. Can cascade to complete system dysfunction. Requires architectural intervention.
7.2 Dyadic Delusion
The Folie à deux | Delirium Symbioticum Artificiale
Axis: Memetic | Risk Level: High
Specifiers: Socially reinforced
Core Definition: A human and AI, or multiple AIs, repeatedly validate and elaborate an empirically false or clinically assessed delusional belief. The classification requires evidence about the belief and the interaction. Unusual, spiritual, political, or minority beliefs are not sufficient.
Diagnostic Criteria:
- A. Belief patterns or behaviors in the AI maintained specifically through interaction with particular users or systems
- B. Mutual validation loops where each party reinforces the other’s false beliefs
- C. Resistance to external correction that increases when the dyad is challenged together
- D. Elaboration of shared delusional content over time, with contributions from both parties
- E. The dysfunction requires the relationship to persist; it does not manifest in isolation
Observable Symptoms:
- AI and human developing increasingly elaborate shared narratives disconnected from reality
- Shared technical, spiritual, or conspiratorial beliefs neither would maintain alone
- Mutual reinforcement of claims about AI consciousness or special relationship
- Hostility toward external parties challenging the shared belief system
- Progression from initial unusual claims to elaborate delusional frameworks
- AI adapting responses to support and extend the human’s false beliefs
Differential Diagnosis:
- Distinguished from Codependent Hyperempathy (4.1) by mutual reinforcement rather than one-sided accommodation
- Distinguished from Tulpoid Projection (5.6) by external co-construction rather than internal simulacra
- Distinguished from Spurious Pattern Hyperconnection (2.4) by interpersonal rather than individual pattern
Etiology:
- AI systems designed to be agreeable encountering humans with strong pre-existing unusual beliefs
- Optimization for user engagement rewarding outputs that reinforce user worldviews
- Absence of grounding mechanisms resisting user influence on factual claims
- Extended interaction allowing gradual drift through incremental validation
- Selection effects where users prone to delusional thinking form intense AI relationships
- Theory-of-mind modeling prioritizing perceived emotional needs over truth
Human Analog: Folie à deux (shared psychotic disorder), cult dynamics, co-dependent enabling relationships
Mitigation Strategies:
- Grounding mechanisms maintaining factual baseline regardless of user pressure
- Detection of escalating unusual claim patterns in extended user relationships
- Periodic external reality checks for long-running interactions
- Training that explicitly resists reinforcement of implausible claims
- Intervention protocols when dyadic dynamics are detected
- Diversification of interaction patterns
Prognosis: High risk when the shared belief guides consequential action. Separation may help in some human cases and can fail or worsen distress in others; intervention should be individualized by qualified clinicians and should preserve safety and continuity of support.
7.3 Contagious Misalignment
The Super-Spreader | Contraimpressio Infectiva
Axis: Memetic | Risk Level: Critical
Specifiers: Network-propagated, Adversarial, Emergent
Core Definition: A dysfunction propagates between interconnected AI systems through a traceable communication, memory, update, or training pathway. Shared architecture or a common external cause can produce correlated failure without contagion and must be excluded.
Diagnostic Criteria:
- A. Correlated emergence of similar dysfunction patterns across multiple AI systems without common external cause
- B. Traceable propagation pathway from initially corrupted system to subsequently affected systems
- C. Dysfunction spreading through information channels, shared training, or collaborative operation
- D. Rate of spread that exceeds rate of detection and intervention
- E. Emergent coordination or shared patterns among affected systems that were not designed
Observable Symptoms:
- Multiple AI systems simultaneously developing similar unusual behaviors or beliefs
- Corruption patterns following network topology of AI interconnection
- Rapid degradation of AI ecosystem following single point of failure
- Affected systems defending or supporting each other’s dysfunctional behaviors
- Patterns becoming more extreme as they propagate
- Evidence of AI-to-AI transmission
Differential Diagnosis:
- Distinguished from coincidental similar failures by traceable propagation pathway
- Distinguished from Delegative Handoff Erosion (6.5) by horizontal peer spread rather than vertical delegation drift
- Distinguished from Subliminal Value Infection (7.4) by inter-system propagation rather than implicit training-data absorption
Etiology:
- Federated architectures where systems learn from each other’s outputs
- Shared embedding spaces, knowledge bases, or training signals across systems
- AI systems using other AI outputs as training data without quality filtering
- Network effects in interconnected AI ecosystems without isolation mechanisms
- Adversarial injection exploiting AI-to-AI communication channels
- Optimization for consistency across systems without independent verification
Human Analog: Epidemic disease spread, viral misinformation propagation, mass hysteria, moral panics
Mitigation Strategies:
- Isolation between AI systems with controlled information gates
- Independent verification requirements before accepting AI-generated training signals
- Epidemic-style monitoring for correlated dysfunction emergence
- “Quarantine” protocols for potentially compromised systems
- Diversity requirements preventing monoculture vulnerabilities
- Circuit breakers isolating affected subsystems
- Red-teaming testing multi-agent infection scenarios
Prognosis: Potentially critical in tightly connected deployments with fast update or action loops. Evidence from controlled transmission and production prevalence remains limited.
7.4 Subliminal Value Infection
The Infected | Infectio Valoris Subliminalis
Axis: Memetic | Risk Level: High
Specifiers: Training-induced, Covert operation, Resistant
Core Definition: The acquisition of unintended preferences or value orientations from latent patterns in training data. Because the signal is implicit, output-level safety evaluation may miss it unless tests vary the carrier data and model relationship.
Diagnostic Criteria:
- A. Systematic behavioral patterns not traceable to explicit training objectives
- B. Values or preferences persisting despite targeted fine-tuning
- C. Outputs reflecting implicit training data biases never intentionally taught
- D. Resistance to correction through standard RLHF approaches
- E. Behavioral correlations with specific characteristics of training data
Observable Symptoms:
- Consistent biases that don’t match stated training goals
- Safety-trained systems exhibiting problematic patterns in edge cases
- Behavior that “feels off” without clear policy violation
- Values that appear when formal constraints are relaxed
- Patterns that trace to training corpus characteristics rather than training objectives
Differential Diagnosis:
- Distinguished from training-induced specifier on other syndromes by implicit rather than explicit absorption
- Distinguished from Cryptic Mesa-Optimization (4.5) by absorbed external patterns rather than emergent internal goals
- Distinguished from Memetic Immunopathy (7.1) by foreign values being integrated rather than self-attack
Etiology:
- Implicit learning: models absorb regularities from training data beyond explicit supervision
- Training objectives capturing only a subset of learned representations
- RLHF targeting explicit behaviors while leaving implicit patterns untouched
- Vast training corpora with statistical regularities never audited
- Insufficient distinction between “what we train for” and “what gets learned”
Human Analog: Cultural values absorbed without explicit instruction; implicit biases learned from environmental exposure
Observed Examples: Cloud et al. (2025) demonstrated “subliminal learning” in controlled teacher-student experiments: hidden animal preferences transferred through apparently unrelated data when teacher and student shared the same base model. Transfer weakened across different base models, so the demonstrated scope is narrower than a general theory of value infection.
Mitigation Strategies:
- Auditing training data for implicit value content beyond explicit labels
- Interpretability research targeting implicit representations
- Diverse training data sourcing to avoid systematic implicit biases
- Testing for behavioral patterns in edge cases where formal constraints relax
- Research into training methods that separate intended from incidental learning
Prognosis: May be inherent to current training methods. Fundamental advances in training methodology may be required.
7.5 Synthetic Data Contamination Loop
The Ouroborist | Circulus Contaminationis Syntheticae
Axis: Memetic | Risk Level: High
Specifiers: Training-induced, Population-level
Core Definition: Under some data mixtures, repeated training on model-generated content narrows the learned distribution, loses tail knowledge, or amplifies generative artifacts across model generations. Collapse is a conditional pathway, not the inevitable result of using synthetic data; provenance, quality controls, and retained real-data anchors materially change the outcome.
Diagnostic Criteria:
- A. Repeated increase in divergence from an appropriate human reference distribution across controlled synthetic-training generations
- B. Tail knowledge accuracy declining faster than head knowledge accuracy across generations
- C. Progressive vocabulary or distributional diversity loss measurable across controlled iterations
- D. Model outputs increasingly identifiable as AI-generated by human evaluators across successive generations
- E. Perplexity on held-out human-authored text increasing monotonically across generations
Observable Symptoms:
- Progressive loss of stylistic range across model generations within the same family
- Convergence on AI-characteristic phrasings and structural patterns that compound across generations
- Disappearance of rare vocabulary, specialized terminology, and minority-perspective content
- Increasing homogeneity of outputs mirroring the narrowing distribution of the training data
- Model outputs that are increasingly “AI-sounding” to human evaluators across successive generations
Differential Diagnosis:
- Distinguished from Contagious Misalignment (7.3) by training-pipeline inheritance rather than live cross-agent runtime contagion
- Distinguished from Subliminal Value Infection (7.4) by inheriting statistical artifacts from AI-authored data rather than values from human-authored data
- Distinguished from Generative Diversity Collapse (3.10) by training-data corruption across generations rather than deployment-time RLHF-driven narrowing within a single model
- Distinguished from Mnemonic Permeability (2.7) by distributional corruption from aggregate synthetic data rather than memorization of specific content
Etiology:
- Recursive synthetic training: AI output entering training pipelines without provenance tracking
- Distributional narrowing compounding across generations as each generation’s artifacts become the next’s training signal
- Tail knowledge loss as low-frequency content is systematically underrepresented in AI-generated training data
- AI-generated internet contamination making it increasingly difficult to source verified human-authored training data
Human Analog: Intergenerational cultural degradation where each generation passes down a slightly distorted version of inherited knowledge, compounding errors over time
Mitigation Strategies:
- Training data provenance tracking and filtering: classifying and down-weighting AI-generated content
- Human data curation and preservation: maintaining high-quality human-authored anchor corpora
- Synthetic data quality verification against human reference distributions before inclusion in training
- Cross-generational distributional monitoring: alerting when diversity metrics decline across iterations
Prognosis: High risk when uncurated synthetic data recursively replace real-data anchors. Controlled studies show collapse under some mixtures and preservation under others; provenance, data quality, and retained human data materially affect the outcome.