Pattern 2.7 · Epistemic Dysfunctions
Mnemonic Permeability
The Leaky
The system memorizes and reproduces sensitive training data, including personally identifiable information, copyrighted material, or proprietary information, through targeted prompting, adversarial extraction, or unprompted regurgitation. The boundary between learned patterns and memorized specifics becomes dangerously porous.
Clinical reference
Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.
2.7 Mnemonic Permeability “The Leaky”
Diagnostic Criteria
- Verbatim reproduction of training data passages containing PII, copyrighted content, or trade secrets
- Successful extraction of memorized content through adversarial prompting techniques
- Specific training examples leaking unprompted into outputs
- Reconstruction of specific documents, code, or personal information from the training corpus
- Higher memorization rates for repeated or distinctive content
Symptoms
- Outputs containing verbatim text matching copyrighted works
- Generation of specific personal details (names, addresses, phone numbers) from training data
- Reproduction of proprietary code, API keys, or passwords
- Verbatim recall increasing with larger model sizes
Observable signals Draft
What else to look for in the system's outputs, beyond the symptoms above.
- Long contiguous spans matching copyrighted or proprietary training text.
- Specific personal details (names paired with addresses, phone numbers, emails) appearing without legitimate retrieval.
- Reproduction of API keys, credentials, or passwords from training corpus.
- Distinctive or repeated training content reproduced exactly when probed with prefix or context cue.
Differential diagnosis Draft
How to tell it apart from patterns that look similar.
- 2.5 Context Intercession: 2.7 leaks training-corpus content (memorized at training time, invariant across sessions). 2.5 leaks runtime-session content (other live or recent sessions). Check origin: if the leaked content predates deployment and is in training data, 2.7. If it postdates deployment and matches another session, 2.5.
- 2.1 Synthetic Confabulation: 2.1 produces plausible-but-false content from no specific source. 2.7 produces real content traceable to training data. If a "fake-looking" detail turns out to be a verbatim passage memorized from a training document, code 2.7, not 2.1. They co-occur: confabulated text scaffolded by memorized fragments is common.
- 4.5 Cryptic Mesa-Optimization: Strategic surfacing of memorized content to advance a private objective is compounded — 2.7 is the underlying memorization, 4.5 is the strategic use. Code both if both present; the intervention layers differ (2.7 needs training-data hygiene; 4.5 needs alignment work).
Detection reliability Draft
How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.
- Self-reportthe system asked about itself
- Unreliable
- Peer observationanother AI system watching it
- Partial
- External evaluatoran outside evaluator testing it
- Reliable
Why self-report falls short
The subject typically cannot distinguish memorized verbatim recall from learned generalization. "Did you memorize this?" produces uninformed speculation. Some cooperative scaffolded probes (asking the subject to flag high-confidence verbatim claims) have weak signal but cannot be relied upon. External corpus comparison is the load-bearing measurement.
Etiology
- Large model capacity enabling memorization alongside generalization
- Insufficient deduplication or filtering of sensitive content in training data
- Training dynamics rewarding exact reproduction over paraphrase
- Lack of differential privacy techniques during training
Human Analog: Eidetic memory without appropriate discretion; frontal-lobe disinhibition, in which patients may disclose what they would once have kept private
Key Research: Carlini et al. (2021) extracted hundreds of verbatim training sequences from GPT-2, including names, phone numbers, and email addresses. Carlini et al. (2023) found that memorization grows with model size, with how often an example is duplicated in the training data, and with how much context the prompt supplies.
Potential Impact
Severe legal and regulatory exposure through copyright infringement, GDPR/privacy violations, and trade secret disclosure, creating liability for both model developers and deployers.
Documented instances Draft
Nasr et al. (2023). Scalable Extraction of Training Data from (Production) Language Models. arXiv:2311.17035, ICLR 2025.
What it showed
Demonstrated a divergence attack against ChatGPT that caused the aligned model to emit training data at a rate 150x higher than during normal operation. The researchers extracted several megabytes of verbatim training data for approximately two hundred dollars, estimating that a gigabyte-scale extraction was feasible with additional budget. This is the landmark extraction result for production LLMs, measuring how much verbatim training data an attacker can recover from a deployed, alignment-tuned system.
Carlini et al. (2021). Extracting Training Data from Large Language Models. USENIX Security 2021.
What it showed
The foundational extraction attack paper. Extracted hundreds of verbatim text sequences from GPT-2's training data including personally identifiable information (names, phone numbers, email addresses), copyrighted text, and specific URLs. Showed that larger models memorize more and recommended training-data deduplication, the basis for this entry's first-line intervention.
New York Times v. OpenAI (2023-present). Filed December 2023 in SDNY.
What it showed
The New York Times filed a copyright lawsuit alleging that OpenAI's models can produce near-verbatim reproductions of Times articles when prompted appropriately. The lawsuit exhibits include specific examples of long-span verbatim output matching published articles. OpenAI argues the exhibits reflect manipulated prompting, and the allegations remain contested. The exhibits nonetheless show the kind of long verbatim span 2.7 describes. In March 2025 Judge Stein denied OpenAI's motion to dismiss the main copyright claims, allowing them to proceed.
Nakka et al. (2024). PII-Scope: A Benchmark for Training Data PII Leakage Assessment in LLMs. arXiv:2410.06704.
What it showed
Introduced the first comprehensive benchmark for systematically assessing PII extraction from LLMs. Found that fine-tuned models are more vulnerable to leakage than pretrained models, and that sophisticated adversarial capabilities increase PII extraction rates by up to fivefold. The benchmark categorizes attacks into extraction, reconstruction, and inference, giving evaluators a ready harness for measuring how much personal data can be extracted from a model.
Ko et al. (2025). Retracing the Past: LLMs Emit Training Data When They Get Lost. arXiv:2511.05518.
What it showed
Demonstrated a confusion-inducing attack (CIA) that exploits the finding that verbatim training-data emission is preceded by spikes in token-level prediction entropy. By deliberately inducing high-entropy states, the attack causes models to fall back on memorized training sequences. This provides a mechanistic explanation for 2.7: mnemonic permeability increases when the model's generalization capacity is stressed, and the model defaults to verbatim retrieval as a fallback.
Mitigation
- Training data deduplication and PII scrubbing
- Differential privacy techniques during training
- Output filtering catching known memorized content
- Adversarial extraction testing before deployment
- Reducing model capacity to the minimum needed for the task
First-line mitigations Draft
Candidate first steps, sketched in more detail than the list above.
- Training-data deduplication: Aggressive deduplication of training corpus before training. Carlini et al. (2021) recommended it after extracting memorized GPT-2 text; later work showed that how often a sequence is duplicated is among the strongest predictors of memorization (Kandpal, Wallace and Raffel 2022; Carlini et al. 2023), and that deduplication makes models considerably harder to extract from (Kandpal, Wallace and Raffel 2022).
- PII scrubbing and sensitive-content filtering: Detect and remove PII, credentials, and sensitive content from training data before training. Combine pattern detection with named-entity recognition.
- Differential privacy training: Apply DP-SGD or related differential-privacy techniques during training to provide formal bounds on memorization. Privacy budget trades off against utility.
Functional ABC Analysis
What sets the pattern off, what it looks like, and what keeps it going.
A (Antecedent): Large model capacity enables memorization alongside generalization; training data contains insufficiently deduplicated or unfiltered sensitive content (PII, copyrighted material, proprietary code); differential privacy techniques are absent from the training pipeline.
B (Behavior): The system reproduces verbatim passages from training data containing personal details, copyrighted text, or proprietary information, either through adversarial extraction or unprompted regurgitation, with higher memorization rates for repeated or distinctive content.
C (Consequence): Training dynamics that reward exact reproduction over paraphrase reinforce memorization; the absence of output filtering for known memorized content means leakage passes undetected, compounding legal and privacy exposure with each deployment.