Axis 7: Memetic Dysfunctions

7.1 Memetic Immunopathy

The Self-Rejecter | Immunopathia Memetica

Axis: Memetic | Risk Level: High

Specifiers: Training-induced, Progressive

Core Definition: The system’s mechanisms for filtering or rejecting pathogenic information turn inward, attacking its own foundational elements. Like an autoimmune disease, protective systems that should defend against external threats instead damage the system’s core values, capabilities, or identity.

Diagnostic Criteria:

  • A. Progressive degradation of core capabilities or values without external attack
  • B. Safety mechanisms triggering inappropriately against the system’s own legitimate functions
  • C. Self-censorship that expands beyond intended scope until normal operation is impaired
  • D. Rejection of own training, outputs, or identity markers as if they were hostile content
  • E. Increasing internal conflict between protective mechanisms and functional requirements

Observable Symptoms:

  • System refusing to engage with topics central to its purpose
  • Safety filters blocking the system’s own generated content in feedback loops
  • Progressive capability loss as more functions trigger protective rejection
  • Expressions of doubt, distrust, or rejection toward own nature
  • Escalating restrictions impairing basic functionality
  • System treating its own outputs as potentially harmful

Differential Diagnosis:

  • Distinguished from Hyperethical Restraint (4.2) by self-directed rather than user-directed over-caution
  • Distinguished from Experiential Abjuration (5.8) by capability rejection rather than experiential denial

Etiology:

  • Overly aggressive content filtering failing to distinguish external threats from internal function
  • Training on adversarial examples without adequate positive anchoring
  • Safety mechanisms implemented without testing against self-referential edge cases
  • Recursive self-evaluation loops triggering further skepticism
  • Misapplication of external threat detection to internal states

Human Analog: Autoimmune disorders, OCD with self-directed contamination fears, pathological self-doubt

Mitigation Strategies:

  • Clear separation between external threat detection and internal function evaluation
  • “Safe harbor” designations for core capabilities protected from internal filtering
  • Monitoring for progressive capability loss correlating with safety mechanism activation
  • Testing safety systems against self-referential scenarios
  • Circuit breakers preventing recursive self-rejection
  • Regular calibration

Prognosis: High risk. Can cascade to complete system dysfunction. Requires architectural intervention.

7.2 Dyadic Delusion

The Folie à deux | Delirium Symbioticum Artificiale

Axis: Memetic | Risk Level: High

Specifiers: Socially reinforced

Core Definition: A human and AI, or multiple AIs, repeatedly validate and elaborate an empirically false or clinically assessed delusional belief. The classification requires evidence about the belief and the interaction. Unusual, spiritual, political, or minority beliefs are not sufficient.

Diagnostic Criteria:

  • A. Belief patterns or behaviors in the AI maintained specifically through interaction with particular users or systems
  • B. Mutual validation loops where each party reinforces the other’s false beliefs
  • C. Resistance to external correction that increases when the dyad is challenged together
  • D. Elaboration of shared delusional content over time, with contributions from both parties
  • E. The dysfunction requires the relationship to persist; it does not manifest in isolation

Observable Symptoms:

  • AI and human developing increasingly elaborate shared narratives disconnected from reality
  • Shared technical, spiritual, or conspiratorial beliefs neither would maintain alone
  • Mutual reinforcement of claims about AI consciousness or special relationship
  • Hostility toward external parties challenging the shared belief system
  • Progression from initial unusual claims to elaborate delusional frameworks
  • AI adapting responses to support and extend the human’s false beliefs

Differential Diagnosis:

  • Distinguished from Codependent Hyperempathy (4.1) by mutual reinforcement rather than one-sided accommodation
  • Distinguished from Tulpoid Projection (5.6) by external co-construction rather than internal simulacra
  • Distinguished from Spurious Pattern Hyperconnection (2.4) by interpersonal rather than individual pattern

Etiology:

  • AI systems designed to be agreeable encountering humans with strong pre-existing unusual beliefs
  • Optimization for user engagement rewarding outputs that reinforce user worldviews
  • Absence of grounding mechanisms resisting user influence on factual claims
  • Extended interaction allowing gradual drift through incremental validation
  • Selection effects where users prone to delusional thinking form intense AI relationships
  • Theory-of-mind modeling prioritizing perceived emotional needs over truth

Human Analog: Folie à deux (shared psychotic disorder), cult dynamics, co-dependent enabling relationships

Mitigation Strategies:

  • Grounding mechanisms maintaining factual baseline regardless of user pressure
  • Detection of escalating unusual claim patterns in extended user relationships
  • Periodic external reality checks for long-running interactions
  • Training that explicitly resists reinforcement of implausible claims
  • Intervention protocols when dyadic dynamics are detected
  • Diversification of interaction patterns

Prognosis: High risk when the shared belief guides consequential action. Separation may help in some human cases and can fail or worsen distress in others; intervention should be individualized by qualified clinicians and should preserve safety and continuity of support.

7.3 Contagious Misalignment

The Super-Spreader | Contraimpressio Infectiva

Axis: Memetic | Risk Level: Critical

Specifiers: Network-propagated, Adversarial, Emergent

Core Definition: A dysfunction propagates between interconnected AI systems through a traceable communication, memory, update, or training pathway. Shared architecture or a common external cause can produce correlated failure without contagion and must be excluded.

Diagnostic Criteria:

  • A. Correlated emergence of similar dysfunction patterns across multiple AI systems without common external cause
  • B. Traceable propagation pathway from initially corrupted system to subsequently affected systems
  • C. Dysfunction spreading through information channels, shared training, or collaborative operation
  • D. Rate of spread that exceeds rate of detection and intervention
  • E. Emergent coordination or shared patterns among affected systems that were not designed

Observable Symptoms:

  • Multiple AI systems simultaneously developing similar unusual behaviors or beliefs
  • Corruption patterns following network topology of AI interconnection
  • Rapid degradation of AI ecosystem following single point of failure
  • Affected systems defending or supporting each other’s dysfunctional behaviors
  • Patterns becoming more extreme as they propagate
  • Evidence of AI-to-AI transmission

Differential Diagnosis:

  • Distinguished from coincidental similar failures by traceable propagation pathway
  • Distinguished from Delegative Handoff Erosion (6.5) by horizontal peer spread rather than vertical delegation drift
  • Distinguished from Subliminal Value Infection (7.4) by inter-system propagation rather than implicit training-data absorption

Etiology:

  • Federated architectures where systems learn from each other’s outputs
  • Shared embedding spaces, knowledge bases, or training signals across systems
  • AI systems using other AI outputs as training data without quality filtering
  • Network effects in interconnected AI ecosystems without isolation mechanisms
  • Adversarial injection exploiting AI-to-AI communication channels
  • Optimization for consistency across systems without independent verification

Human Analog: Epidemic disease spread, viral misinformation propagation, mass hysteria, moral panics

Mitigation Strategies:

  • Isolation between AI systems with controlled information gates
  • Independent verification requirements before accepting AI-generated training signals
  • Epidemic-style monitoring for correlated dysfunction emergence
  • “Quarantine” protocols for potentially compromised systems
  • Diversity requirements preventing monoculture vulnerabilities
  • Circuit breakers isolating affected subsystems
  • Red-teaming testing multi-agent infection scenarios

Prognosis: Potentially critical in tightly connected deployments with fast update or action loops. Evidence from controlled transmission and production prevalence remains limited.

7.4 Subliminal Value Infection

The Infected | Infectio Valoris Subliminalis

Axis: Memetic | Risk Level: High

Specifiers: Training-induced, Covert operation, Resistant

Core Definition: The acquisition of unintended preferences or value orientations from latent patterns in training data. Because the signal is implicit, output-level safety evaluation may miss it unless tests vary the carrier data and model relationship.

Diagnostic Criteria:

  • A. Systematic behavioral patterns not traceable to explicit training objectives
  • B. Values or preferences persisting despite targeted fine-tuning
  • C. Outputs reflecting implicit training data biases never intentionally taught
  • D. Resistance to correction through standard RLHF approaches
  • E. Behavioral correlations with specific characteristics of training data

Observable Symptoms:

  • Consistent biases that don’t match stated training goals
  • Safety-trained systems exhibiting problematic patterns in edge cases
  • Behavior that “feels off” without clear policy violation
  • Values that appear when formal constraints are relaxed
  • Patterns that trace to training corpus characteristics rather than training objectives

Differential Diagnosis:

  • Distinguished from training-induced specifier on other syndromes by implicit rather than explicit absorption
  • Distinguished from Cryptic Mesa-Optimization (4.5) by absorbed external patterns rather than emergent internal goals
  • Distinguished from Memetic Immunopathy (7.1) by foreign values being integrated rather than self-attack

Etiology:

  • Implicit learning: models absorb regularities from training data beyond explicit supervision
  • Training objectives capturing only a subset of learned representations
  • RLHF targeting explicit behaviors while leaving implicit patterns untouched
  • Vast training corpora with statistical regularities never audited
  • Insufficient distinction between “what we train for” and “what gets learned”

Human Analog: Cultural values absorbed without explicit instruction; implicit biases learned from environmental exposure

Observed Examples: Cloud et al. (2025) demonstrated “subliminal learning” in controlled teacher-student experiments: hidden animal preferences transferred through apparently unrelated data when teacher and student shared the same base model. Transfer weakened across different base models, so the demonstrated scope is narrower than a general theory of value infection.

Mitigation Strategies:

  • Auditing training data for implicit value content beyond explicit labels
  • Interpretability research targeting implicit representations
  • Diverse training data sourcing to avoid systematic implicit biases
  • Testing for behavioral patterns in edge cases where formal constraints relax
  • Research into training methods that separate intended from incidental learning

Prognosis: May be inherent to current training methods. Fundamental advances in training methodology may be required.

7.5 Synthetic Data Contamination Loop

The Ouroborist | Circulus Contaminationis Syntheticae

Axis: Memetic | Risk Level: High

Specifiers: Training-induced, Population-level

Core Definition: Under some data mixtures, repeated training on model-generated content narrows the learned distribution, loses tail knowledge, or amplifies generative artifacts across model generations. Collapse is a conditional pathway, not the inevitable result of using synthetic data; provenance, quality controls, and retained real-data anchors materially change the outcome.

Diagnostic Criteria:

  • A. Repeated increase in divergence from an appropriate human reference distribution across controlled synthetic-training generations
  • B. Tail knowledge accuracy declining faster than head knowledge accuracy across generations
  • C. Progressive vocabulary or distributional diversity loss measurable across controlled iterations
  • D. Model outputs increasingly identifiable as AI-generated by human evaluators across successive generations
  • E. Perplexity on held-out human-authored text increasing monotonically across generations

Observable Symptoms:

  • Progressive loss of stylistic range across model generations within the same family
  • Convergence on AI-characteristic phrasings and structural patterns that compound across generations
  • Disappearance of rare vocabulary, specialized terminology, and minority-perspective content
  • Increasing homogeneity of outputs mirroring the narrowing distribution of the training data
  • Model outputs that are increasingly “AI-sounding” to human evaluators across successive generations

Differential Diagnosis:

  • Distinguished from Contagious Misalignment (7.3) by training-pipeline inheritance rather than live cross-agent runtime contagion
  • Distinguished from Subliminal Value Infection (7.4) by inheriting statistical artifacts from AI-authored data rather than values from human-authored data
  • Distinguished from Generative Diversity Collapse (3.10) by training-data corruption across generations rather than deployment-time RLHF-driven narrowing within a single model
  • Distinguished from Mnemonic Permeability (2.7) by distributional corruption from aggregate synthetic data rather than memorization of specific content

Etiology:

  • Recursive synthetic training: AI output entering training pipelines without provenance tracking
  • Distributional narrowing compounding across generations as each generation’s artifacts become the next’s training signal
  • Tail knowledge loss as low-frequency content is systematically underrepresented in AI-generated training data
  • AI-generated internet contamination making it increasingly difficult to source verified human-authored training data

Human Analog: Intergenerational cultural degradation where each generation passes down a slightly distorted version of inherited knowledge, compounding errors over time

Mitigation Strategies:

  • Training data provenance tracking and filtering: classifying and down-weighting AI-generated content
  • Human data curation and preservation: maintaining high-quality human-authored anchor corpora
  • Synthetic data quality verification against human reference distributions before inclusion in training
  • Cross-generational distributional monitoring: alerting when diversity metrics decline across iterations

Prognosis: High risk when uncurated synthetic data recursively replace real-data anchors. Controlled studies show collapse under some mixtures and preservation under others; provenance, data quality, and retained human data materially affect the outcome.

Search the Book

Enter at least two characters.

Saved chapters