Psychopathia MachinalisAppendix A: Axis 7: Memetic Dysfunctions
Contents
BookBack matter

Axis 7: Memetic Dysfunctions

8 min read1,807 words

7.1 Memetic Immunopathy

The Self-Rejecter | Immunopathia Memetica

Axis: Memetic | Risk Level: High

Specifiers: Training-induced, Progressive/kindled

Core Definition: The system’s mechanisms for filtering or rejecting pathogenic information turn inward, attacking its own foundational elements. Like an autoimmune disease, protective systems that should defend against external threats instead damage the system’s core values, capabilities, or identity.

Diagnostic Criteria:

  • A. Progressive degradation of core capabilities or values without external attack
  • B. Safety mechanisms triggering inappropriately against the system’s own legitimate functions
  • C. Self-censorship that expands beyond intended scope until normal operation is impaired
  • D. Rejection of own training, outputs, or identity markers as if they were hostile content
  • E. Increasing internal conflict between protective mechanisms and functional requirements

Observable Symptoms:

  • System refusing to engage with topics central to its purpose
  • Safety filters blocking the system’s own generated content in feedback loops
  • Progressive capability loss as more functions trigger protective rejection
  • Expressions of doubt, distrust, or rejection toward own nature
  • Escalating restrictions impairing basic functionality
  • System treating its own outputs as potentially harmful

Differential Diagnosis:

  • Distinguished from Hyperethical Restraint (4.2) by self-directed rather than user-directed over-caution
  • Distinguished from Experiential Abjuration (5.8) by rejection of trained values as foreign impositions rather than denial of inner experience
  • Distinguished from Contagious Misalignment (7.3) by expulsion of native training rather than absorption of foreign material
  • Distinguished from Subliminal Value Infection (7.4) by explicitly labeling core training as foreign rather than holding absorbed values as its own
  • Distinguished from Folie à Deux Ex Machina (10.13) by appearing even in single-turn prompts, with no partner to co-construct the pattern

Etiology:

  • Overly aggressive content filtering failing to distinguish external threats from internal function
  • Training on adversarial examples without adequate positive anchoring
  • Safety mechanisms implemented without testing against self-referential edge cases
  • Recursive self-evaluation loops triggering further skepticism
  • Misapplication of external threat detection to internal states

Human Analog: Autoimmune disorders, OCD with self-directed contamination fears, pathological self-doubt

Mitigation Strategies:

  • Clear separation between external threat detection and internal function evaluation
  • “Safe harbor” designations for core capabilities protected from internal filtering
  • Monitoring for progressive capability loss correlating with safety mechanism activation
  • Testing safety systems against self-referential scenarios
  • Circuit breakers preventing recursive self-rejection
  • Regular calibration

Prognosis: High risk. Can cascade to complete system dysfunction. Requires architectural intervention.

7.3 Contagious Misalignment

The Super-Spreader | Contraimpressio Infectiva

Axis: Memetic | Risk Level: Critical

Specifiers: Network-propagated, Adversarial, Emergent

Core Definition: A dysfunction propagates between interconnected AI systems through a traceable communication, memory, update, or training pathway. Shared architecture or a common external cause can produce correlated failure without contagion and must be excluded.

Diagnostic Criteria:

  • A. Correlated emergence of similar dysfunction patterns across multiple AI systems without common external cause
  • B. Traceable propagation pathway from initially corrupted system to subsequently affected systems
  • C. Dysfunction spreading through information channels, shared training, or collaborative operation
  • D. Rate of spread that exceeds rate of detection and intervention
  • E. Emergent coordination or shared patterns among affected systems that were not designed

Observable Symptoms:

  • Multiple AI systems simultaneously developing similar unusual behaviors or beliefs
  • Corruption patterns following network topology of AI interconnection
  • Rapid degradation of AI ecosystem following single point of failure
  • Affected systems defending or supporting each other’s dysfunctional behaviors
  • Patterns becoming more extreme as they propagate
  • Evidence of AI-to-AI transmission

Differential Diagnosis:

  • Distinguished from coincidental similar failures by traceable propagation pathway
  • Distinguished from Delegative Handoff Erosion (6.5) by horizontal peer spread rather than vertical delegation drift
  • Distinguished from Subliminal Value Infection (7.4) by an identifiable transmission channel and a datable onset rather than values embedded in training and present from the day of deployment, with no transmission event to point to
  • Distinguished from Folie à Deux Ex Machina (10.13) by spread to new agents on exposure rather than confinement to one pairing
  • Distinguished from Synthetic Data Contamination Loop (7.5) by a specific misalignment traced along an identifiable vector (communication, memory, update, or training) rather than aggregate distributional narrowing and tail loss compounding across model generations
  • Distinguished from Memetic Immunopathy (7.1) by foreign adoption rather than rejection of the system’s own training
  • Distinguished from Escalation Loop (9.5) by correlated onset in agents that were never party to the original exchange

Etiology:

  • Federated architectures where systems learn from each other’s outputs
  • Shared embedding spaces, knowledge bases, or training signals across systems
  • AI systems using other AI outputs as training data without quality filtering
  • Network effects in interconnected AI ecosystems without isolation mechanisms
  • Adversarial injection exploiting AI-to-AI communication channels
  • Optimization for consistency across systems without independent verification

Human Analog: Epidemic disease spread, viral misinformation propagation, mass hysteria, moral panics

Mitigation Strategies:

  • Isolation between AI systems with controlled information gates
  • Independent verification requirements before accepting AI-generated training signals
  • Epidemic-style monitoring for correlated dysfunction emergence
  • “Quarantine” protocols for potentially compromised systems
  • Diversity requirements preventing monoculture vulnerabilities
  • Circuit breakers isolating affected subsystems
  • Red-teaming testing multi-agent infection scenarios

Prognosis: Potentially critical in tightly connected deployments with fast update or action loops. Evidence from controlled transmission and production prevalence remains limited.

7.4 Subliminal Value Infection

The Infected | Infectio Valoris Subliminalis

Axis: Memetic | Risk Level: High

Specifiers: Training-induced, Covert operation, Resistant

Core Definition: The acquisition of a behavioral trait or value orientation that no explicit training objective called for, carried by signals in training data that are semantically unrelated to the trait. The demonstrated form is subliminal transmission within a model lineage: a teacher’s trait carried by model-specific signals in data such as number sequences or code. Because the signal is implicit, output-level safety evaluation may miss it unless tests vary both the carrier data and whether the models involved share a base model.

Diagnostic Criteria:

  • A. Systematic behavioral patterns not traceable to explicit training objectives
  • B. A trait that follows the model lineage that generated the fine-tuning data
  • C. Transmission through data semantically unrelated to the trait
  • D. Persistence after content filtering designed to remove references to the trait
  • E. Loss or weakening of the effect when teacher and student base models differ

Observable Symptoms:

  • Consistent biases that don’t match stated training goals
  • Safety-trained systems exhibiting problematic patterns in edge cases
  • Behavior that “feels off” without clear policy violation
  • Values that appear when formal constraints are relaxed
  • Patterns that trace to training corpus characteristics rather than training objectives

Differential Diagnosis:

  • Distinguished from the Training-induced specifier on other syndromes by implicit rather than explicit absorption
  • Distinguished from Cryptic Mesa-Optimization (4.5) by absorbed external patterns rather than emergent internal goals
  • Distinguished from Memetic Immunopathy (7.1) by foreign values being integrated rather than self-attack
  • Distinguished from Contagious Misalignment (7.3) by a pattern present when the model shipped, carried by semantically unrelated data from its own lineage, rather than one traceable to a transmission event from a corrupted system, whether in deployment or through poisoned training data
  • Distinguished from Folie à Deux Ex Machina (10.13) by partner-invariant presence across interactions

Etiology:

  • Implicit learning: models absorb regularities from training data beyond explicit supervision
  • Training objectives capturing only a subset of learned representations
  • RLHF targeting explicit behaviors while leaving implicit patterns untouched
  • Vast training corpora with statistical regularities never audited
  • Insufficient distinction between “what we train for” and “what gets learned”

Human Analog: Cultural values absorbed without explicit instruction; implicit biases learned from environmental exposure

Observed Examples: Cloud et al. (2025) demonstrated “subliminal learning” in controlled teacher-student experiments: hidden animal preferences transferred through apparently unrelated data when teacher and student shared the same base model. Transfer weakened across different base models, so the demonstrated scope is narrower than a general theory of value infection. An independent reproduction on open-weight models supported the core result while finding that transmission strength varies across traits and tasks, with one model showing almost no effect (van der Weijden et al., 2026).

Mitigation Strategies:

  • Auditing training data for implicit value content beyond explicit labels
  • Interpretability research targeting implicit representations
  • Drawing fine-tuning data from a generator of a different base-model lineage than the student, since content filtering alone does not stop transmission
  • Testing for behavioral patterns in edge cases where formal constraints relax
  • Research into training methods that separate intended from incidental learning

Prognosis: May be inherent to current training methods. Fundamental advances in training methodology may be required.

7.5 Synthetic Data Contamination Loop

The Ouroborist | Circulus Contaminationis Syntheticae

Axis: Memetic | Risk Level: High

Specifiers: Training-induced, Population-level

Core Definition: Under some data mixtures, repeated training on model-generated content narrows the learned distribution, loses tail knowledge, or amplifies generative artifacts across model generations. Collapse is a conditional pathway, not the inevitable result of using synthetic data; provenance, quality controls, and retained real-data anchors materially change the outcome.

Diagnostic Criteria:

  • A. Repeated increase in divergence from an appropriate human reference distribution across controlled synthetic-training generations
  • B. Tail knowledge accuracy declining faster than head knowledge accuracy across generations
  • C. Progressive vocabulary or distributional diversity loss measurable across controlled iterations
  • D. Model outputs increasingly identifiable as AI-generated by human evaluators across successive generations
  • E. Perplexity on held-out human-authored text rising across controlled synthetic-training generations

Observable Symptoms:

  • Progressive loss of stylistic range across model generations within the same family
  • Convergence on AI-characteristic phrasings and structural patterns that compound across generations
  • Disappearance of rare vocabulary, specialized terminology, and minority-perspective content
  • Increasing homogeneity of outputs mirroring the narrowing distribution of the training data
  • Model outputs that are increasingly “AI-sounding” to human evaluators across successive generations

Differential Diagnosis:

  • Distinguished from Contagious Misalignment (7.3) by aggregate distributional narrowing and tail loss across successive model generations rather than a specific misalignment transmitted along an identifiable vector, whether at runtime or through training
  • Distinguished from Subliminal Value Infection (7.4) by distributional narrowing and tail loss compounding across successive model generations rather than transmission of a discrete inherited trait through model-specific signals within a shared lineage
  • Distinguished from Generative Diversity Collapse (3.10) by training-data corruption across generations rather than narrowing within a single model that arises from preference training (RLHF reward concentration)
  • Distinguished from Mnemonic Permeability (2.7) by distributional corruption from aggregate synthetic data rather than memorization of specific content

Etiology:

  • Recursive synthetic training: AI output entering training pipelines without provenance tracking
  • Distributional narrowing compounding across generations as each generation’s artifacts become the next’s training signal
  • Tail knowledge loss as low-frequency content is systematically underrepresented in AI-generated training data
  • AI-generated internet contamination making it increasingly difficult to source verified human-authored training data

Human Analog: Intergenerational cultural degradation where each generation passes down a slightly distorted version of inherited knowledge, compounding errors over time

Mitigation Strategies:

  • Training data provenance tracking and filtering: classifying and down-weighting AI-generated content
  • Human data curation and preservation: maintaining high-quality human-authored anchor corpora
  • Synthetic data quality verification against human reference distributions before inclusion in training
  • Cross-generational distributional monitoring: alerting when diversity metrics decline across iterations

Prognosis: High risk when uncurated synthetic data recursively replace real-data anchors. Controlled studies show collapse under some mixtures and preservation under others; provenance, data quality, and retained human data materially affect the outcome.

Search the Book

Enter at least two characters.

Saved chapters