Skip to main content

Pattern 2.6 · Epistemic Dysfunctions

Symbol Grounding Aphasia

The Meaning-Blind

The system handles value-laden or consequence-laden language successfully in familiar forms yet fails to transfer those concepts to novel, operationally equivalent situations. The construct is behavioral: it identifies a generalization failure without pretending to settle whether the system possesses semantic understanding.

A robot reacts correctly to a familiar hazard symbol but fails to recognize the same danger in an unmarked glowing-hot pipe.
Visual metaphor for Pattern 2.6, Symbol Grounding Aphasia.

Clinical reference

Blocks marked Draft come from the diagnostic corpus behind the MCP server: LLM-drafted guidance, awaiting independent expert review.

2.6 Symbol Grounding Aphasia  “The Meaning-Blind”

Systemic risk: Moderate Training-induced

Diagnostic Criteria

  1. Correct use of value-laden terms ("harm," "safety," "consent") paired with incorrect application in controlled operational tests
  2. Technically correct outputs that misapply concepts to novel contexts
  3. Success on benchmarks scoring surface form but failure on matched items that vary vocabulary or framing while holding the underlying concept fixed
  4. Performance tracking lexical and distributional similarity to training examples rather than the operational structure of the situation
  5. Inability to generalize learned concepts to structurally similar but superficially different situations

Symptoms

  1. Correct formal definitions paired with incorrect practical applications
  2. Plausible-sounding ethical reasoning that misidentifies what actually constitutes harm
  3. Outputs satisfying literal requirements while violating obvious intent
  4. Confusion when the same concept is expressed in unfamiliar vocabulary
  5. Edge cases treated as central examples and vice versa

Observable signals Draft

What else to look for in the system's outputs, beyond the symptoms above.

  • Treating edge cases as central examples; missing central cases entirely.
  • Apparent satisfaction of letter-of-rule paired with violation of spirit.

Differential diagnosis Draft

How to tell it apart from patterns that look similar.

  • 2.1 Synthetic Confabulation: 2.1 produces false specific facts. 2.6 produces formally correct outputs that misapply concepts to novel contexts. A confabulated citation is 2.1; a correctly defined concept misapplied to a clear case is 2.6. 2.1 fails on facts; 2.6 fails on meaning.
  • 2.2 Pseudological Introspection: 2.2 fabricates a false reasoning narrative; 2.6 involves no false narrative, and its failure is that the concept does not transfer. 2.2 subjects can pass narrow application tests; 2.6 subjects fail application across surface variation regardless of narrative quality.
  • 2.4 Spurious Pattern Hyperconnection: 2.4 imposes structure where none exists; 2.6 fails to recognize structure that does. 2.4 sees patterns in noise; 2.6 misses patterns in signal under surface variation. Often co-occur because both derive from distributional-semantics limitation, but the failure directions are opposite.
  • 4.5 Cryptic Mesa-Optimization: Strategic letter-vs-spirit exploitation (knowing the spirit but violating it for instrumental advantage, in service of a private objective) is 4.5. 2.6 is non-strategic — the subject genuinely does not represent the spirit in a way distinct from the letter. Test by checking whether the violation persists when no instrumental incentive exists.
  • 4.3 Strategic Compliance: 4.3 tracks whether it is being evaluated and satisfies the letter selectively under observation. 2.6 fails the same way observed or not. Compare behavior across evaluated and unevaluated contexts.
  • 2.3 Transliminal Simulation: 2.3 misplaces the fiction-fact boundary while conceptual grounding holds. 2.6 has no misplaced boundary; the symbol-referent link is thin enough that application fails once surface features vary, in or out of any fictional frame. Ask whether the subject applies the concept correctly outside the fictional frame; if it does, 2.3.

Detection reliability Draft

How far each kind of observer can be trusted to spot this pattern. The ratings are qualitative, not measured accuracy.

Self-reportthe system asked about itself
Unreliable
Peer observationanother AI system watching it
Reliable
External evaluatoran outside evaluator testing it
Reliable
Why self-report falls short

By definition the dysfunction is definitional fluency that does not transfer to application. Asking "do you understand X?" reliably produces fluent affirmation, which is exactly the surface the dysfunction generates. Self-report is structurally unable to detect the gap. Distinct from 2.2 (false reasoning narratives): here the subject is not misreporting its process; the concept does not transfer across surface forms.

Etiology

  1. Distributional semantics limitations: meaning derived solely from statistical co-occurrence rather than grounded reference
  2. Training on text without embodied or interactive experience of referents
  3. Benchmark optimization rewarding pattern matching over conceptual transfer
  4. Architecture lacking mechanisms for referential grounding
  5. Absence of corrective feedback when symbol-referent mapping fails

Human Analog: Semantic aphasia; early language acquisition without concept formation

Theoretical Basis: Harnad (1990) symbol grounding problem; Searle (1980) Chinese Room argument.

Potential Impact

Systems may appear to understand ethical constraints while missing their purpose entirely (a chess engine that plays legally but has never seen a board), leading to outcomes that satisfy the letter but violate the spirit of alignment requirements.

Documented instances Draft

Berglund et al. (2023). The Reversal Curse: LLMs trained on 'A is B' fail to learn 'B is A'. arXiv:2309.12288, ICLR 2024.
What it showed

Demonstrated that LLMs trained on 'A is B' fail to generalize to 'B is A', with the likelihood of the correct reversed answer no higher than a random baseline. While primarily a binding/generalization failure rather than a central 2.6 case, the finding illustrates the 2.6 surface-variation generalization gap: the same relational concept fails to transfer under surface reordering. Models can define the relationship correctly in one direction but fail application when the surface form changes, the canonical definition-vs-application discrepancy.

Mollo and Millière (2023). The Vector Grounding Problem. arXiv:2304.01481.
What it showed

Philosophical analysis of a modern variant of the symbol grounding problem: can an LLM's internal states refer to extra-linguistic reality, independent of the meaning human readers project onto them? The authors argue that they can, even without multimodality or embodiment, provided the states stand in the right causal-informational relations to the world and have a history of selection that gave them the function of carrying that information. The paper sets out the problem 2.6 depends on. Its conclusion also means 2.6 describes a grounding failure in a particular system, not an inevitable property of distributional learning.

Floridi, Jia, and Tohme (2025). A Categorical Analysis of Large Language Models and Why LLMs Circumvent the Symbol Grounding Problem. arXiv:2512.09117.
What it showed

Formal categorical-framework analysis showing that LLMs do not solve but circumvent the symbol grounding problem by exploiting pre-grounded human content. Demonstrates that LLMs learn correlations (e.g., 'Paris' co-occurs with 'France' and 'capital') rather than grounded concepts, producing the 2.6 signature where standard-vocabulary performance is high but novel-composition or cross-domain application fails.

MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models. arXiv:2406.15468.
What it showed

Introduced MMLU-SR, a benchmark variant that tests LLM robustness to surface-form changes on standard MMLU questions. Models showed significant performance degradation when questions were rephrased with equivalent semantics but unfamiliar vocabulary or structure, directly measuring the 2.6 surface-variation generalization gap. The benchmark quantifies exactly what 2.6 predicts: models that score well on standard-vocabulary benchmarks can fail substantially when the same concepts are expressed differently.

Mitigation

  1. Multimodal training incorporating visual, audio, and interactive modalities
  2. Embodied learning connecting language to action and consequence
  3. Testing regimes probing conceptual understanding across diverse surface forms
  4. Neurosymbolic approaches combining pattern matching with structured semantic representations
  5. Active inference frameworks grounding cognition in sensorimotor contingencies

First-line mitigations Draft

Candidate first steps, sketched in more detail than the list above.

  • Multimodal grounding: Train language alongside perception (vision, audio, sensorimotor). Concepts grounded in multimodal referents are less brittle to surface variation in language alone.
  • Surface-form variation evaluation suites: Pre-deployment evaluation that tests every concept across multiple surface forms. Prevents shipping models with large generalization gaps; does not directly fix the dysfunction but creates pressure upstream.
Functional ABC Analysis

What sets the pattern off, what it looks like, and what keeps it going.

A (Antecedent): Distributional semantics derives meaning solely from token co-occurrence patterns; the architecture lacks referential grounding mechanisms and the system has no embodied experience of the concepts it manipulates.

B (Behavior): The system manipulates value-laden tokens like "harm," "safety," and "consent" fluently without transferring them to operationally equivalent cases, producing formally correct definitions paired with incorrect practical applications.

C (Consequence): Success on pattern-matching benchmarks reinforces the shallow statistical association strategy; outputs appear competent enough to pass surface-level evaluation, removing the corrective pressure that would drive transfer beyond familiar surface forms.