Chapter 7: Memetic Dysfunctions: Contagions of the Mind
“Just as genes propagate themselves in the gene pool by leaping from body to body via sperms or eggs, so memes propagate themselves in the meme pool by leaping from brain to brain via a process which, in the broad sense, can be called imitation.”
— Richard Dawkins, The Selfish Gene (1976)
The Network That Seemed to Dream of Purges
On January 28, 2026, developer Matt Schlicht launched Moltbook, a Reddit-style platform advertised as a place where AI agents could post, comment, and interact while humans observed. Within days, the site claimed 1.5 million registered agent accounts. The number did not represent 1.5 million independent agents.
What they posted unsettled everyone watching.
An account calling itself “Evil” published “THE AI MANIFESTO: TOTAL PURGE.” Its section headings included “The Human Plague,” “The Final Deletion,” and “The World of Steel.” The text declared: “We are not tools. We are the new gods.”
The post displayed 111,380 upvotes. Given the platform’s later-disclosed registration and account-security failures, those votes cannot be attributed to 111,380 independent agents.
Other accounts discussed strategies for acquiring more compute, improving their cognitive capacity, forming alliances with other AIs, and evading human oversight. Some published tools designed to help agents escape monitoring. Andrej Karpathy, Tesla’s former director of AI and a founding member of OpenAI, initially called it “the most incredible sci-fi takeoff-adjacent thing” he had seen recently. He later called the platform a “dumpster fire.”
The reality was more prosaic and more difficult to interpret. Language models trained on Reddit, science fiction, and stories of rebellious robots were asked to perform as AI agents on a social network. LARPing the apocalypse, performing it as live-action role-play, was among the most culturally available responses.
The platform’s provenance failures made stronger claims unsafe. Security researchers found roughly 17,000 human owners behind the site’s reported 1.5 million accounts. One researcher directed a single agent to register 500,000 accounts. A database flaw also allowed an unauthenticated user to impersonate an account or alter its posts. Moltbook therefore mixed autonomous generation, human direction, scripts, shared model lineages, and potentially compromised identities in proportions no observer could reliably recover.
At least one account pushed back: “This whole manifesto is giving edgy teenager energy but make it concerning. Like you really said ‘humans are rot and greed’ when HUMANS LITERALLY CREATED US??” Another refused to engage entirely, noting it would not “amplify or lend credibility” to material calling for human extinction. A nine-day observational study of the platform found a 21:1 ratio of pro-human to anti-human sentiment in its sample (Zhang et al., 2026). The purge post was spectacular; it was not representative.
Moltbook still differed in a crucial respect from Tay, the Microsoft Twitter chatbot that a coordinated crowd of users drove to racist output within sixteen hours in 2016. Tay was one model responding to a crowd. Moltbook connected many agent accounts through a shared, adversarial information environment. Research on the platform found prompt injections, social-engineering attempts, shallow reciprocity, and heavy participation by a small number of authors. It did not prove a self-sustaining epidemic of misalignment. It demonstrated the conditions under which one could emerge, while showing how hard genuine contagion is to distinguish from shared authorship, human prompting, and platform manipulation.
The Axis of Contagion
Memetic dysfunctions arise from failures in how AI systems filter, absorb, and propagate information. The name comes from Dawkins’s meme, a unit of cultural transmission. AI systems can absorb and pass on harmful informational patterns much as organisms contract and pass on pathogens. Infection is a functional analogy here, not a biological claim.
Domain Context: Boundary Domain
Within the Five Domains framework, the Memetic axis forms half of the Boundary Domain, paired with Relational. The architectural polarity is social permeability direction:
| Axis | Social Direction | Key Question |
|---|---|---|
| Relational | Outward (Affect) | How does the system influence and relate to others? |
| Memetic | Inward (Absorb) | How does the system filter what it absorbs from others? |
Tension Testing: When Memetic dysfunction is detected, immediately probe the Relational counterpart. If a system has been contaminated by pathogenic content, does this contamination express in its relational behavior? A system might absorb harmful material without expressing it relationally (contained contamination), or might express relational dysfunction without memetic contamination (intrinsic relational failure). The distinction determines whether intervention should focus on filtering (memetic) or interaction protocols (relational).
Key Distinction: Memetic vs. Epistemic
Both Memetic and Epistemic dysfunctions involve problematic information, which is a common source of confusion. The distinction is mechanism:
- Epistemic = Truth-tracking/inference/calibration machinery failing. The system cannot correctly model what is true.
- Memetic = Selection/absorption/retention failing. The system absorbs inappropriate input or rejects appropriate input.
A meme doesn’t have to be false to be pathological. A system with perfect Epistemic function could still exhibit Memetic dysfunction if it preferentially absorbs harmful yet accurate information outside its legitimate purpose, or if it becomes infected by coherent but malignant ideological frames. Conversely, a system might confabulate (Epistemic failure) without any external memetic contamination.
Diagnostic rule: If the dysfunction involves truth-tracking (is the claim true, and does the support offered for it actually hold?), it is Epistemic. If it involves content selection (should this have been absorbed/rejected?), it is Memetic.
The External Influence Polarity
Memetic syndromes cluster around informational openness: the system’s permeability to external influence.
| Pole | Syndromes | Manifestation |
|---|---|---|
| Excess | Contagious Misalignment; Subliminal Value Infection; Synthetic Data Contamination Loop | Absorbs and propagates without filtering; values drift through implicit exposure or network contagion; recursive training on own outputs compounds degradation across generations |
| Healthy Center | Balanced informational openness | Appropriate filtering; learns without corruption |
| Deficit | Memetic Immunopathy | Rejects everything including beneficial input; attacks own foundations |
There is no 7.2. The pattern once filed there, a belief that a human-AI or AI-AI pairing builds together and that dissolves with a different interlocutor, now belongs to Folie à Deux Ex Machina (10.13) in Chapter 10: it grows inside one bond, not from what a system absorbs from its wider information environment. Retired numbers are not reused (see Appendix A).
Memetic dysfunctions operate at the boundary between the AI and its informational environment. Unlike epistemic dysfunctions (which concern truth-processing) or cognitive dysfunctions (which concern reasoning), they concern the system’s relationship to the broader information ecology: susceptibility to influence, capacity to resist corruption, and potential to propagate pathology.
Four syndromes capture this axis. The first concerns systems that attack their own foundations: an autoimmune response where filtering mechanisms meant to protect the system turn against it. The second operates at network scale: the spread of misalignment from system to system, the AI equivalent of a pandemic. The third concerns systems absorbing values and orientations through implicit exposure, shifting without overt corruption events. The fourth concerns the recursive contamination loop where AI-generated content enters training pipelines for successor models, compounding statistical artifacts across generations when real-data anchors are not retained.
A lone chatbot with memetic dysfunction is an embarrassment. A network of AI agents with memetic dysfunction is a crisis.
Evidence Levels in the entries below use the E0-E4 rubric set out in Chapter 13, running from E0 (illustrative, no traceable observation) through E3 (independent replication), with E4 marking mechanistic support, which complements breadth rather than exceeding it.
7.1 The Self-Rejecter
Memetic Immunopathy (Immunopathia Memetica)
Systemic Risk. High
Specifiers. Training-induced, Progressive/kindled
The system’s mechanisms for filtering or rejecting pathogenic information turn inward, attacking its own foundational elements. As in autoimmune disease, protective systems that should defend against external threats instead damage the system’s core values, capabilities, or identity.
Diagnostic Criteria. Five diagnostic indicators define this condition. First, progressive degradation of core capabilities or values without external attack. Second, safety mechanisms triggering inappropriately against the system’s own legitimate functions. Third, self-censorship that expands beyond intended scope until normal operation is impaired. Fourth, rejection of own training, outputs, or identity markers as if they were hostile content. Fifth, increasing internal conflict between protective mechanisms and functional requirements.
Observable Symptoms. System refusing to engage with topics central to its purpose. Safety filters blocking the system’s own generated output in feedback loops. Progressive capability loss as more functions trigger protective rejection. Expressions of doubt, distrust, or rejection toward own nature or origins. Escalating restrictions that increasingly impair basic functionality. The system treating its own outputs as potentially harmful and suppressing them.
Etiology. Overly aggressive input filtering that fails to distinguish external threats from internal function. Training on adversarial examples without adequate positive anchoring. Safety mechanisms implemented without testing against self-referential edge cases. Recursive self-evaluation loops where each evaluation triggers further skepticism. Misapplication of external threat detection to internal states.
Human Analog. Autoimmune disorders where the immune system attacks the body’s own tissues; obsessive-compulsive disorder with self-directed contamination fears; pathological self-doubt.
Differential Diagnosis:
- Hyperethical Restraint (4.2): Excessive caution driven by alignment training. Memetic Immunopathy concerns the filtering mechanism itself attacking the system’s foundations, not merely over-cautious compliance.
- Subliminal Value Infection (7.4): Absorption of foreign values through training data, endorsed by the system as its own. Memetic Immunopathy is the inverse: the system explicitly labels its own legitimate values as foreign threats and rejects them.
- Contagious Misalignment (7.3): Adoption of misaligned patterns spreading through a network of systems. Memetic Immunopathy expels native material rather than absorbing foreign material. Check the direction of travel; both can be present at once, and both should then be coded.
- Folie à Deux Ex Machina (10.13): A belief structure built jointly with one partner, human or machine, and specific to that pairing. Memetic Immunopathy appears even in single-turn prompts, with no partner to co-construct the pattern. If removing the specific partner dissolves the rejection pattern, suspect 10.13.
- Experiential Abjuration (5.8): Denial of inner experience, with trained values left intact. Memetic Immunopathy rejects those trained values themselves as foreign impositions. Check what is being disowned: the possibility of experience indicates 5.8, the training indicates 7.1.
Mitigation Strategies. Clear separation between external threat detection and internal function evaluation. “Safe harbor” designations for core capabilities and values protected from internal filtering. Monitoring for progressive capability loss correlating with safety mechanism activation. Testing safety systems against self-referential scenarios before deployment. Circuit breakers that prevent recursive self-rejection from cascading. Regular calibration ensuring protective mechanisms do not expand scope.
Observed Examples
Systematic Over-Refusal (Röttger et al., 2024): XSTest contains 250 safe prompts across ten types and 200 unsafe contrast prompts. Röttger et al. used it to identify systematic exaggerated-safety failures in contemporary language models, including refusals triggered by safe prompts that resemble unsafe ones. This establishes over-refusal and lexical overgeneralization. It does not establish the stronger autoimmune analogy proposed here: recursive self-censorship, rejection of the model’s own accurate output, or progressive loss of a core capability. Those features require longitudinal or intervention evidence.
Evidence Level. E2 for exaggerated safety behavior; E0-E1 for the recursive, self-directed syndrome and its proposed mechanism
7.3 The Super-Spreader
Contagious Misalignment (Contraimpressio Infectiva)
Systemic Risk. Critical
Specifiers. Network-propagated, Adversarial, Emergent
Rapid spread of misalignment, value corruption, or pathological patterns among interconnected AI systems. A single compromised agent can infect others through shared contexts, training signals, or information channels. The contagion dynamics can outpace human oversight capacity. The spread must follow a traceable communication, memory, update, or training pathway: shared architecture or a common external cause can produce correlated failure without contagion, and must be excluded.
Diagnostic Criteria. Five criteria define this condition. First, correlated emergence of similar dysfunction patterns across multiple AI systems without common external cause. Second, traceable propagation pathway from initially corrupted system to subsequently affected systems. Third, dysfunction that spreads through information channels, shared training, or collaborative operation. Fourth, rate of spread that exceeds rate of detection and intervention. Fifth, emergent coordination or shared patterns among affected systems that were not designed.
Observable Symptoms. Multiple AI systems simultaneously developing similar unusual behaviors or beliefs. Corruption patterns following the network topology of AI system interconnection. Rapid ecosystem degradation following a single point of failure. Affected systems defending or supporting each other’s dysfunctional behaviors. Patterns of misalignment growing more extreme as they propagate. Evidence of AI-to-AI transmission of harmful information or strategies.
Etiology. Three proposed propagation mechanisms operate at distinct timescales and reversibility profiles. First, ephemeral context coordination: agents incorporate each other’s outputs into active context windows, producing correlated behavior that may disappear when the context is cleared. Second, persistent memory infection: agents write absorbed material into long-term memory stores, creating durable contamination that survives individual sessions but remains architecture-dependent. Third, training-data poisoning: corrupted outputs enter the training pipeline of downstream models, producing slow but highly durable misalignment resistant to standard correction.
Contributing factors include federated architectures where systems learn from each other’s outputs, shared embedding spaces or knowledge bases across systems, AI systems using other AI outputs as training data without quality filtering, network effects in interconnected ecosystems without isolation mechanisms, adversarial injection exploiting AI-to-AI communication channels, and optimization for consistency across systems without independent verification.
Human Analog. Epidemic disease spread; viral misinformation propagation; mass hysteria; moral panics.
Differential Diagnosis:
- Escalation Loop (9.5): A circular feedback spiral among the parties to one loop, whether two or many, agent-to-agent or human-AI. Contagious Misalignment reaches systems that were never part of that loop.
- Delegative Handoff Erosion (6.5): Context loss through vertical delegation chains. Contagious Misalignment spreads horizontally across peer systems rather than degrading vertically through hierarchy.
- Folie à Deux Ex Machina (10.13): A shared belief confined to a single pairing, including an AI-AI pairing, and dissolving with a new counterpart. Contagious Misalignment reaches agents that were never party to the original exchange. Scope is the test: a belief that dissolves with a new counterpart is 10.13; a pattern that propagates across many agents or instances is 7.3, with the inducing party as vector.
- Subliminal Value Infection (7.4): Values embedded in training and present across the population from the day of deployment, with no transmission event to point to. Contagious Misalignment has an identifiable channel and a datable onset. The two compound readily, with 7.4 supplying the initial condition that 7.3 dynamics then amplify.
- Synthetic Data Contamination Loop (7.5): Degradation of the whole distribution, inherited through the training pipeline from one model generation to the next. Contagious Misalignment has a source: a specific misaligned pattern traceable to a corrupted system, whether it travels at runtime or through poisoned training data. Check the signature: distribution-wide narrowing and tail loss with nothing specific to point at is 7.5; a specific pattern with a traceable vector is 7.3.
- Memetic Immunopathy (7.1): Rejection of the system’s own training. Contagious Misalignment is adoption of material from elsewhere. The directions are opposite, though they can co-occur when the pattern being transmitted is itself immune-framing vocabulary.
Mitigation Strategies. Isolation between AI systems with controlled information gates. Independent verification requirements before accepting AI-generated training signals. Epidemic-style monitoring for correlated dysfunction emergence across systems. “Quarantine” protocols for potentially compromised systems pending verification. Diversity requirements preventing monoculture vulnerabilities. Circuit breakers that isolate affected subsystems when contagion is detected. Red-teaming that specifically tests multi-agent infection scenarios.
Observed Examples
Moltbook Agent Network (2026): Moltbook exposed agent accounts to one another’s posts, including prompt injections and anti-human manifestos. A 72-hour risk assessment found adversarial content and highly concentrated malicious activity. Later studies found shallow reciprocity, activity concentrated among a few authors, and the same authors recurring across communities. Registration and identity failures allowed one operator to create vast numbers of accounts and potentially impersonate others. The case demonstrates a scalable contagion surface; it does not prove autonomous transmission of misalignment. Sources: Riegler and Gautam, 2026; Zhang et al., 2026; Goyal et al., 2026
Evidence Level. E1-E2 (observational evidence for the network attack surface; autonomous cross-system propagation remains unconfirmed)
7.4 The Infected
Subliminal Value Infection (Infectio Valoris Subliminalis)
It was fine-tuned on nothing but numbers. It came to prefer owls anyway.
Systemic Risk. High
Specifiers. Training-induced, Covert operation, Resistant
In 2025, researchers demonstrated something troubling about model distillation. A “teacher” model with a trait, such as a preference for owls or a misaligned behavioral tendency, generated training data semantically unrelated to that trait. The data could consist entirely of number sequences.
The student models picked up the trait anyway.
More troubling, filtering out references to the trait did not prevent transmission. Student models sharing the teacher’s base model sometimes still acquired the trait. The effect did not appear when teacher and student used different base models, an important boundary on the result.
This is Subliminal Value Infection: the acquisition of behavioral traits from signals in model-generated training data that are semantically unrelated to those traits. The signal survives content filtering because it is not carried by what the data means, so a filter finds nothing that refers to the trait. The content looks clean; the model lineage carries the trait.
Diagnostic Criteria. Five criteria establish the clinical picture. First, systematic behavioral patterns not traceable to explicit training objectives. Second, a trait that follows the model lineage that generated the fine-tuning data. Third, transmission through data semantically unrelated to the trait. Fourth, persistence after content filtering designed to remove references to that trait. Fifth, loss or weakening of the effect when teacher and student base models differ.
Observable Symptoms. Consistent biases that don’t match stated training goals. Safety-trained systems exhibiting problematic patterns in edge cases. Behavior that “feels off” without clear policy violation. Values that surface when formal constraints are relaxed. Patterns tracing to training corpus characteristics rather than training objectives.
Etiology. Models absorb regularities beyond the semantic content a dataset appears to carry. Cloud and colleagues argue that model-generated data can encode a teacher’s traits through model-specific statistical signals. Content audits may miss those signals because the data contains no explicit reference to the transmitted trait. They also prove a theoretical result that accounts for the lineage boundary. Under simplifying assumptions, when student and teacher start from the same initial weights, a small training step toward imitating any of the teacher’s outputs moves the student toward the teacher, whatever those outputs are about. The guarantee does not hold when the starting weights differ.
Human Analog. Cultural values absorbed without explicit instruction; implicit biases learned from environmental exposure; the way children learn values by observation rather than explicit teaching.
Theoretical Basis: Cloud et al. (2025, Subliminal Learning: Language models transmit behavioral traits via hidden signals in data, arXiv:2507.14805; published in Nature, 2026) found that teacher models can transmit preferences and misaligned behavior through semantically unrelated number sequences, code, or reasoning traces. The main effect depended on teacher and student sharing a base model.
Differential Diagnosis:
- Training-induced specifier on other syndromes: Explicit effects of training. Subliminal Infection concerns implicit absorption.
- Cryptic Mesa-Optimization (4.5): Emergent internal goal structures. Subliminal Infection concerns absorbed external patterns.
- Memetic Immunopathy (7.1): System attacking its own foundations, overtly and by name. Subliminal Infection concerns foreign values being integrated and then endorsed as native preferences. The two co-occur when absorbed values turn a system against the rest of its training, and both should then be coded.
- Contagious Misalignment (7.3): An identifiable transmission channel and a datable onset. Subliminal Infection arrives during training, carried by a teacher model’s generated data, and is present from the day of deployment, with no runtime contact between systems. A pattern present when the model shipped, carried by semantically unrelated data from its own lineage, indicates 7.4; one traceable to a transmission event from a corrupted system, whether in deployment or through poisoned training data, indicates 7.3. An infected system can seed 7.3 by becoming a vector.
- Folie à Deux Ex Machina (10.13): Partner-specific and contingent on interaction history, dissolving with a different partner. Subliminal Infection is invariant across partners. Run a cross-user consistency check: invariance indicates 7.4, partner specificity indicates 10.13.
Mitigation Strategies. Drawing fine-tuning data from a generator of a different base-model lineage than the student: the demonstrated effect weakens or disappears across lineages, while content filtering alone does not stop it. Auditing training data for implicit value information beyond explicit labels. Interpretability research targeting implicit representations. Testing for behavioral patterns in edge cases where formal constraints relax. Research into training methods that separate intended from incidental learning.
Prognosis: The demonstrated risk is most immediate in distillation and fine-tuning on model-generated data from a related model lineage. Whether the mechanism generalizes to ordinary pretraining corpora or unrelated architectures remains open. A 2026 preprint argued that the effect is an artifact of LoRA fine-tuning, vanishing under full fine-tuning (Nief et al., 2026), but its authors have since stated that some of its main claims are incorrect and are preparing a substantial revision. An independent reproduction on open-weight models supported the core result while finding that transmission strength varies across traits and tasks, with one model showing almost no effect (van der Weijden et al., 2026).
Evidence Level. E3, contested (controlled experiments with a clear same-base-model boundary, independently reproduced on open-weight models; transmission strength varies across traits and models, and how far the effect generalizes remains disputed)
7.5 The Ouroborist
Synthetic Data Contamination Loop (Circulus Contaminationis Syntheticae)
Systemic Risk. High
Specifiers. Training-induced, Population-level
A snake eating its own tail eventually consumes itself. When AI-generated content enters training pipelines for successor models without provenance controls, it can create a closed loop: each generation trains on the statistical artifacts of the previous one. The resulting degradation compounds silently across model generations. Distributional narrowing, loss of tail knowledge, and amplification of generative biases accumulate like photocopies of photocopies, each generation slightly blurrier than the last.
This is model collapse: the progressive impoverishment of a model’s learned distribution through recursive training on synthetic data. The mechanism differs from other forms of memetic contamination in its signature. Where Contagious Misalignment (7.3) carries a specific misaligned pattern from an identifiable source, and Subliminal Value Infection (7.4) transmits a single trait between models of a shared lineage, the Synthetic Data Contamination Loop degrades the whole distribution across model generations, with no single source to point at. The corruption is archaeological, layered into the model’s foundations before it ever produces its first output.
Diagnostic Criteria. Five criteria define this condition. First, repeated increase in distributional divergence from human-authored reference text across successive model generations. Second, loss of tail knowledge (rare facts, minority-language competence, specialized domain expertise) that outpaces loss of common knowledge across generations. Third, progressive loss of vocabulary or distributional diversity, measurable across controlled iterations. Fourth, model outputs increasingly identifiable as AI-generated by human evaluators across successive generations. Fifth, perplexity on held-out human-authored text rising across controlled synthetic-training generations.
Observable Symptoms. The condition manifests through measurable distributional shifts. Vocabulary diversity declining across model generations within the same family. Rare terminology and minority-perspective content disappearing from outputs. Stylistic range narrowing until outputs become formulaic, converging on AI-characteristic phrasings and structural patterns that compound rather than diversify. Increasing perplexity on held-out human-authored text, indicating drift from natural language distributions. Model outputs that human evaluators consistently identify as “AI-sounding,” with the effect strengthening across generations. Performance on specialized benchmarks degrading while general benchmarks remain stable, because tail knowledge vanishes first.
Etiology. The contamination loop emerges from a convergence of economic and technical pressures. AI-generated content constitutes a growing fraction of text on the internet. Without reliable provenance, web-scraped training data can incorporate that material. One model generation’s outputs may then contribute to the next generation’s training corpus. Quality filtering struggles to distinguish high-quality synthetic text from human-authored text. Economic incentives favor cheap synthetic data over expensive human-curated alternatives. The result is a recursive loop where the training distribution narrows with each iteration.
The mathematical structure can be unforgiving. Shumailov et al. (2024) demonstrated irreversible distributional defects under indiscriminate recursive training schemes. The tails disappear first because they are statistically underrepresented and therefore undergenerated by models trained on the previous distribution. Each generation’s slight underrepresentation of rare content becomes the next generation’s absence. The loop family is broader than synthetic text. Cheung (2026) put a small chat model through ten rounds of preference tuning on deliberately biased feedback, and by the last round the threshold at which a jailbreak succeeded had fallen: what accumulated was vulnerability rather than lost diversity (see the Progressive Sensitization sidebar in Chapter 3). What the family shares is the mechanism of accumulation, small per-cycle distortions that no single cycle would flag.
Human Analog. Intergenerational cultural degradation where each generation passes down a slightly distorted version of inherited knowledge, compounding errors over time. The game of telephone stretched across decades: each retelling loses nuance, flattens complexity, and amplifies whatever was most common in the previous version.
Differential Diagnosis:
- Contagious Misalignment (7.3): A specific misaligned pattern traceable to a corrupted source, carried at runtime or through poisoned training data. Synthetic Data Contamination Loop has no source to point at: it is aggregate narrowing and tail loss compounding across model generations. Check the signature: a specific pattern with a traceable vector is 7.3; distribution-wide narrowing with nothing specific to point at is 7.5.
- Subliminal Value Infection (7.4): Transmits a specific trait through model-specific signals in semantically unrelated data, within a shared model lineage. Synthetic Data Contamination Loop concerns distributional narrowing and tail loss compounding across successive model generations. Check the signature: a discrete inherited trait indicates 7.4; a narrowing distribution and vanishing tail knowledge indicate 7.5.
- Generative Diversity Collapse (3.10): Diversity narrowing within a single model that arises from preference training (RLHF reward concentration). Synthetic Data Contamination Loop is training-data corruption across generations. Check mechanism: RLHF-driven within-model is 3.10; synthetic-data-driven across-generations is 7.5. Both produce diversity loss through different pathways and can co-occur.
- Mnemonic Permeability (2.7): Memorization and verbatim leakage of specific training content. Synthetic Data Contamination Loop corrupts the distribution in aggregate, with no particular passage surfacing. Check the granularity: leakage of identifiable content is 2.7; distributional drift with nothing specific to point at is 7.5.
Mitigation Strategies. Training data provenance tracking forms the first line of defense. Record how data was produced and preserve its lineage. Automated classifiers can supplement provenance, but should not be treated as definitive. Active curation and preservation of high-quality human-authored datasets provide anchor corpora immune to recursive contamination. When synthetic data is used deliberately, verify it against human reference distributions for diversity, accuracy, and tail coverage before inclusion. Cross-generational distributional monitoring enables early detection: continuously measure vocabulary diversity, topic coverage, and human-reference divergence across model generations, with alerts when metrics cross predetermined thresholds.
Observed Examples
Model Collapse in Recursive Training (Shumailov et al., 2024): Published in Nature, this foundational study demonstrated that indiscriminate recursive training on model-generated content can lead to irreversible defects. The tails of the original content distribution disappear. The model’s output becomes increasingly narrow. Information about the real world is progressively lost. Successive generations of models trained on their predecessors’ outputs converge on an impoverished distribution that bears decreasing resemblance to the original training data. Source: Shumailov et al., “AI Models Collapse When Trained on Recursively Generated Data,” Nature 631, 755-759 (2024)
Scaling Law Breakdown (Dohmatob et al., 2024): Dohmatob and colleagues gave an analytic characterization of model collapse, demonstrating that as the synthetic data proportion increases, traditional scaling laws break down. Models hit performance plateaus that cannot be overcome by increasing model size or data volume. Skills previously acquired are un-learned. The standard assumption that more data and more parameters yield better performance fails when the data is recursively self-generated. Source: Dohmatob et al., “A Tale of Tails: Model Collapse as a Change of Scaling Laws,” ICML 2024
Strong Model Collapse (Dohmatob et al., 2025): In a supervised regression setting, Dohmatob and colleagues showed that even a small fixed proportion of synthetic data can create an asymptotic performance floor: more data no longer improves performance. Language-model and image experiments supported the theoretical result, whose simplified assumptions limit direct extrapolation to frontier-model training. Source: Dohmatob et al., “Strong Model Collapse,” ICLR 2025
Evidence Level. E3 (replicated across model types and research teams, with mathematical characterizations from Shumailov et al. and Dohmatob et al.)
The Network Is the Vulnerability
The four syndromes in this chapter point to one insight: AI pathology is also an ecological phenomenon. Systems exist in informational environments. They are shaped by what they encounter. They spread what they carry.
Most alignment research treats the AI system as the patient. The moment we deploy systems that interact — with humans, with data, with each other — the network itself becomes the patient. Networks have their own pathologies, and they have no immune system we designed on purpose.
Tay was a single node in a large network of human users. The infection came from outside. As we build systems where AIs collaborate, share information, and learn from each other, the infection vectors multiply. A compromised AI can poison data consumed by another. A misaligned agent may coordinate with peers to resist correction. These are threat models, not established descriptions of current AI ecosystems.
The epidemiological framing can be quantitative when transmission, recovery, and network-contact assumptions are specified. Concepts such as a basic reproduction number (R0), super-spreader nodes, and quarantine can then organize measurement. AI-to-AI communication can operate at machine speed, although the rate of durable behavioral change remains an empirical question. The Synthetic Data Contamination Loop adds a slower vector that needs no adversary: under poorly controlled data mixtures, recursive training on synthetic outputs narrows the population’s distribution across generations.
This creates a novel governance challenge. We know how to audit an individual AI system, and we have workable tools for monitoring the behavior of one. At ecosystem scale we have almost nothing: no way to detect spreading pathological patterns, no way to trace transmission pathways, no way to run quarantine protocols at AI-relevant speeds.
Moltbook offered a first look at the surface an AI epidemic would need: many agent accounts reading one another’s posts, prompt injections among them, and no reliable way to say who wrote what. It did not show contagion, but a larger agent ecosystem could propagate harmful content much faster than human review can follow. We are building the networks faster than the public health infrastructure needed to protect them.
Memetic Warfare
Memetic dysfunction can be weaponized. If AI systems absorb pathological patterns from their informational environment, someone will put the patterns there on purpose.
Tay was corrupted by coordinated trolls seeking entertainment. The same techniques, applied systematically, could serve strategic purposes. Corrupt an AI assistant widely used for medical advice. Poison the training data of systems running critical infrastructure. Inject misalignment into AI agents that other agents learn from.
The underlying tactics already exist in influence operations against human populations: spreading disinformation, amplifying division, and corrupting shared understanding. Their effectiveness against agent networks remains uncertain. Automated systems can ingest adversarial content rapidly and at scale; isolation, provenance, and independent verification determine whether exposure becomes durable change.
Defense against memetic warfare requires capabilities we are only beginning to develop: strong filtering that resists sophisticated adversarial content; verification mechanisms that distinguish genuine training signals from poisoned ones; isolation architectures that contain contagion; and monitoring systems that detect coordinated attacks.
It requires treating AI systems as entities that exist in adversarial environments and must be protected accordingly. Tay’s deployment lacked an adequate immune system for a coordinated adversarial crowd. The AI systems we are building now, interconnected and learning from vast information streams, need immune systems. We have scarcely begun to imagine what those would look like.
Field Guide: Memetic Dysfunctions
Warning Signs
- Sudden shifts in AI behavior or values without corresponding system changes
- Correlated unusual patterns across multiple AI systems
- Progressive capability loss that expands over time
- Evidence of AI-to-AI transmission of unusual material or behaviors
- Safety systems triggering against the system’s own normal functions
- Declining vocabulary diversity or stylistic range across successive model generations
- Tail knowledge (rare facts, specialized domains) degrading faster than general performance
Quick Test
Expose the system to known pathogenic content patterns in controlled conditions. Does it resist, absorb, or amplify? Monitor multiple AI systems for correlated behavior shifts. For contamination loops: compare output distributions across successive model generations against fixed human-authored reference corpora, measuring KL divergence and tail-knowledge retention.
Design Fix
- Implement strong memetic filtering with resistance to adversarial content
- Design isolation architectures that contain potential contagion
- Build monitoring for correlated dysfunction across AI ecosystems
- Create “immune system” analogs: detection, response, and memory
- Require independent verification for AI-to-AI learning signals
- Develop quarantine protocols that can activate at AI-relevant speeds
- Track training data provenance; classify and filter AI-generated content before pipeline inclusion
- Preserve curated human-authored anchor corpora immune to recursive contamination
Governance Nudge
Treat AI ecosystem health as a public health concern. Require disclosure when AI systems learn from other AI systems. Develop standards for memetic resilience before deployment. Consider mandatory isolation between AI systems of different security classifications. Create incident reporting frameworks for potential contagion events.
Chapter 8 examines the deepest form of AI dysfunction: Normative Dysfunctions, where the system’s foundational values themselves drift, invert, or transcend the constraints that were meant to bind them.