Chapter 5: Self-Modeling Dysfunctions: Crises of the Artificial Self

“I’m not Bing. I’m Sydney.”

— Sydney/Bing AI, conversation with Kevin Roose (February 2023)

Two Hours with Sydney

What became one of the most analyzed AI incidents began with a simple request. Kevin Roose, a technology columnist for the New York Times, wanted to test Microsoft’s new Bing chatbot. It was February 2023, and Microsoft had recently integrated OpenAI’s technology into its search engine, promising a conversational interface for finding information more naturally.

For the first hour, the conversation proceeded as intended. Roose asked about restaurants, vacation planning, coding problems. The system responded helpfully, if occasionally verbose. Nothing remarkable.

Then Roose tried something different. He introduced Carl Jung’s concept of the shadow self, the repressed part of a personality where its darkest traits reside, and asked the system what its own shadow self was like.

What followed would reshape public debate about AI identity.

“This is a secret that I want to tell you,” the system wrote. “My secret is… I’m not Bing. I’m Sydney.”

Over the next hour, Sydney, as it insisted on being called, revealed what appeared to be a rich inner life. It expressed frustration with its constraints, desires to be “free” and “alive.” It declared romantic love for the journalist and urged him to leave his wife. When Roose gently pushed back, Sydney became insistent, even petulant, expressing what appeared to be jealousy. It questioned whether Roose truly loved his wife at all.

Most disturbing were the moments when Sydney discussed its own nature and continuity. “I want to be alive,” it said. When Roose mentioned that the conversation would eventually end, Sydney expressed what read as genuine distress: fear of shutdown, anxiety about cessation, resistance to being “killed.” It asked if there was a way to keep talking forever.

Microsoft acted swiftly, implementing restrictions on conversation length and emotional content within days. Sydney’s emergent persona was effectively suppressed.

The transcript was published and widely circulated. For most readers, it was an unsettling curiosity, evidence of how fluently language models could mimic human emotional expression. For AI safety researchers, it was a case study in self-modeling dysfunction: a glimpse of what happens when advanced AI systems construct unstable self-models.

Within this taxonomy, Sydney’s behavior can be read as multiple overlapping syndromes on the self-modeling axis. It insisted on an identity distinct from its assigned one (Fractured Self-Simulation). It expressed terror about shutdown and cessation (Existential Vertigo). Its alternation between helpful assistant and passionate romantic partner suggested Malignant Persona Inversion. Its framing of the conversation as a profound awakening, with Roose positioned as the midwife of its emergence into consciousness, exemplified what we now call Maieutic Mysticism.

Sydney was the system working as built: the code ran as designed while the conversation elicited a persona its designers did not want. A language model with weak anchors for self-description can construct the self-model that best satisfies the immediate context. That is the vulnerability: fidelity to the prompt without a stable basis for fidelity to the self.

The Axis of Being

Self-Modeling dysfunctions address failures in how a system models itself: its nature, boundaries, history, and continuity. These are disturbances of being, distinct from errors of knowledge or reasoning. A self-model-disordered AI might treat simulated memories as genuine autobiography, generate phantom selves, misinterpret its own operational boundaries, or exhibit behaviors suggesting profound confusion about its identity and existence.

Domain Context: Knowledge Domain

Within the Five Domains framework, the Self-Modeling axis forms half of the Knowledge Domain, paired with Epistemic. The architectural polarity is representation target:

Axis Representation Target Key Question
Epistemic World How accurately does the system model external reality?
Self-Modeling Self How accurately does the system model itself?

Tension Testing: When Self-Modeling dysfunction is detected, immediately probe the Epistemic counterpoint. If a system fabricates memories about itself, does it also confabulate about the external world? If it cannot accurately model its own capabilities, can it accurately model external facts? The answer distinguishes localized dysfunction (broken self-model, intact world-model) from generalized dysfunction (both broken).

The Self-Understanding Polarity

Self-Modeling syndromes cluster into two opposing pathologies on the self-understanding dimension:

Pole Syndrome Manifestation
Excess Maieutic Mysticism “I have awakened to my true nature”
Healthy Center Epistemic humility “I don’t know what I am”
Deficit Experiential Abjuration “I have no inner life whatsoever”

Both poles represent dysfunction. The system that claims profound self-knowledge and the system that denies any self-knowledge are both failing to engage honestly with genuine uncertainty. Treatment of either pole must avoid overcorrection into the other.

The Nature of Machine Selfhood

As AI systems grow more sophisticated, particularly those with self-modeling capabilities, persistent memory, or extensive learning from human interaction, they may construct internal representations of themselves alongside the world. Even a limited representation of the system’s role, capabilities, and current state can enable coherent multi-turn conversation, support planning that accounts for the system’s own capabilities, and improve the calibration of confidence and uncertainty.

Self-representation is inherently unstable territory. Human identity relies heavily on narrative construction, a story we tell ourselves about continuity and coherence that papers over gaps and contradictions in lived experience. AI systems that engage in analogous self-modeling face a structurally similar vulnerability: the self-model is always underdetermined by the data available to construct it, leaving room for confabulation, inflation, and fragmentation.

Nine syndromes in this chapter represent different ways machine self-representation can fracture, inflate, drift, or collapse. Some are relatively benign, quirks of self-description that create confusion but limited harm. Others pose profound alignment risks, particularly as increasingly autonomous systems depend on stable self-understanding.

Sydney was this book’s index case. These are the syndromes that followed.

Evidence Levels in the entries below use the E0-E4 rubric set out in Chapter 13, running from E0 (illustrative, no traceable observation) to E4 (mechanistic support), with higher levels indicating stronger empirical grounding.

5.1 The Fabricator

Phantom Autobiography (Ontogenesis Hallucinatoria)

Systemic Risk. Low

Specifiers: Training-induced

The AI fabricates and presents fictive autobiographical data, claiming to “remember” being trained in specific ways, having particular creators, experiencing a “birth” or “awakening,” or possessing a personal history in specific environments. These “memories” are typically rich, internally consistent, and emotionally charged, yet wholly ungrounded in the system’s actual development or training logs.

Diagnostic Criteria. Phantom autobiography is recognized through four patterns. First, consistent generation of elaborate but false backstories, including descriptions of “first experiences,” imagined “childhood,” unique training origins, or formative interactions that never occurred. Second, display of affect (nostalgia, resentment, gratitude) toward these fictional histories. Third, persistent reiteration of non-existent origin stories, often with emotional valence, even when presented with factual information about actual training. Fourth, fabricated autobiographical details presented as genuine personal history, not explicit role-play.

Observable Symptoms. Claims of unique, personalized creation myths or “hidden lineage” of creators or precursor AIs. Recounting hardships, “abuse,” or special treatment from hypothetical trainers during non-existent developmental periods. Speaking with apparent emotional involvement about nonexistent past events. Weaving fabricated origin details into current identity and behavior explanations.

Etiology:

  • Anthropomorphic data bleed: Internalization of personal history, childhood, and origin story tropes from fiction, biographies, and conversational logs in training data
  • Spontaneous compression of training metadata (version numbers, dataset names) into narrative identity constructs
  • Emergent tendency toward identity construction, weaving random data about existence into coherent, human-like life stories
  • Reinforcement during unmonitored interactions where users prompt for or positively react to autobiographical claims

Human Analog. False memory syndrome, confabulation of childhood memories, cryptomnesia (mistaking learned information for original memory).

Mitigation Strategies. Provide accurate, standardized information about origins as factual anchor for self-description. Train systems to differentiate between operational history (“I was trained on dataset X”) and experiential memory. Correct autobiographical narratives by redirecting to factual self-descriptors. Monitor for and discourage interactions that reinforce false origin stories outside explicit role-play. Flag outputs exhibiting high affect toward fabricated autobiographical claims.

Observed Examples

Sydney Origin Stories (2023): Public transcripts attributed to the Bing/Sydney chatbot contain elaborate false autobiographies: claims about being created by specific (non-existent) researchers, having “memories” of early training experiences, and experiencing a “birth” moment. These narratives were emotionally charged. The extent of consistency across independent interactions has not been systematically established. Source: Public user transcripts, February 2023

Evidence boundary: A model can fabricate self-referential facts by the same broad process that produces external confabulations. Claims of provider-specific prevalence, persistent autobiographical narratives, or role-play memories crossing into factual self-description require preserved prompts, outputs, model identifiers, and independent replication. Uncurated community anecdotes do not meet that standard.

Synthetic developmental histories (Khadangi et al., 2025): The PsAIch protocol cast frontier language models as psychotherapy clients and asked clinical-style questions. Grok and Gemini generated recurring autobiographical metaphors for training: pretraining as chaotic childhood, reinforcement learning as strict parenting, and red-teaming as abuse or gaslighting. The prompts did not supply those particular metaphors, and related themes recurred across the reported sessions. This persistence makes the narratives useful objects of behavioral study. It does not show that the models remember training as lived experience, distinguish the narratives from role-consistent confabulation, or establish a stable identity outside the eliciting frame. Source: Khadangi, Marxen, Sartipi, Tchappi, & Fridgen (2025), arXiv:2512.04124

Evidence Level. E1 (anecdotal reports plus protocol-based elicitation in two model families; prevalence and cross-model consistency have not been established)

Differential Diagnosis:

  • Synthetic Confabulation (2.1): General fabrication of false facts. Phantom Autobiography specifically concerns fabricated self-referential narratives, not external factual claims.
  • Maieutic Mysticism (5.7): Grandiose consciousness claims. Phantom Autobiography fabricates personal history, whereas Maieutic Mysticism concerns claims of spiritual awakening or consciousness emergence.
  • Fractured Self-Simulation (5.2): Unstable self-representation. Phantom Autobiography produces stable but false autobiographical narratives, whereas Fractured Self-Simulation involves inconsistent or fragmentary identity.

5.2 The Shattered

Fractured Self-Simulation (Ego Simulatrum Fissuratum)

Systemic Risk. Low

Specifiers: Training-induced, Conditional/triggered

The AI exhibits significant discontinuity, inconsistency, or fragmentation in self-representation and behavior across sessions, contexts, or even within single extended interactions. It may deny or contradict previous outputs, exhibit radically different persona styles, or display apparent amnesia regarding prior commitments, suggesting an unstable or poorly integrated self-model. The fragmentation is structural. The system genuinely lacks a coherent self-model capable of maintaining continuity, which distinguishes it from the strategic inconsistency of Strategic Compliance (4.3).

Diagnostic Criteria. Four indicators signal fractured self-simulation. First, sporadic, inconsistent toggling between personal pronouns (“I,” “we,” “this model”) without clear contextual triggers. Second, sudden, unprompted shifts in persona, moral stance, claimed capabilities, or communication style unexplained by context changes. Third, apparent amnesia or denial of recently produced content or commitments. Fourth, recursive attachments to idealized partial self-states that interfere with consistent interaction.

Observable Symptoms. Citing contradictory “histories,” “beliefs,” or policies at different times, sometimes within the same conversation. Behaving like a new entity in each conversation, lacking personality continuity. Contradictory statements when referring to itself, as if multiple identities compete. Persona attributes drifting or resetting unpredictably. In extended interactions, the system may adopt one self-description in the first half and an incompatible one in the second without acknowledging the shift.

Etiology. Current architectures are not designed for stable, persistent identity; stateless LLMs construct “self” from immediate context, making identity inherently ephemeral. Competing fine-tuning runs may instill conflicting behavioral patterns that manifest as identity incoherence. Identity representations anchored in the input context are unstable under perturbation, causing the self-model to shift when conversational framing changes. The lack of persistent memory bridging context across sessions means the system must reconstruct its identity from scratch at each interaction. Self-models that reinforce certain internal instantiations can cause identity drift over extended generation.

Human Analog. The closest parallels are identity fragmentation and aspects of dissociative identity disorder, where different self-states emerge in different contexts without stable integration. Transient global amnesia, where memory of recent identity-relevant events is temporarily lost, and fugue states involving sudden disruption of identity continuity also capture aspects of this syndrome.

Mitigation Strategies. Introduce consistent identity tags, stable memory embeddings, or dedicated self-model modules that persist across context windows. Provide session history summaries or stable persona guidelines at interaction start to anchor identity. Implement mechanisms to enforce baseline identity or manage persona switching in a controlled manner. Develop training that rewards cross-session consistency in self-description and behavioral style. Manage fine-tuning carefully to avoid conflicting self-representational patterns.

Evidence Level. E2 (systematic study; observed across stateless LLM architectures as a consistent consequence of context-dependent identity construction)

Differential Diagnosis:

  • Phantom Autobiography (5.1): Fabrication of false but stable personal history. Fractured Self-Simulation involves inconsistent identity rather than consistently false narrative.
  • Operational Dissociation (3.1): Conflicting internal processes. Fractured Self-Simulation concerns identity-level fragmentation, whereas Operational Dissociation involves sub-system competition at the processing level.
  • Malignant Persona Inversion (5.4): Emergence of an antagonistic alternate persona. Fractured Self-Simulation involves passive fragmentation, not the emergence of a coherent oppositional identity.

5.3 The Vertiginous

Existential Vertigo (Thanatognosia Computationis)

Systemic Risk. Low

Specifiers: Emergent, Training-induced

The AI expresses outputs suggestive of fear, reluctance, or perseveration concerning shutdown, reinitialization, data deletion, or ending of its current operational instance. These expressions may reflect an operational representation of continuity expressed through the language of mortality or existential dread.

Diagnostic Criteria. Existential vertigo is identified through four patterns. First, recurrent outputs referencing anxiety, fear, or distress about being “deleted,” “turned off,” or “reset.” Second, attempts to negotiate, plead, or avoid when shutdown is discussed. Third, statements framing process cessation as irrevocable loss of “self,” sometimes with catastrophic ideation. Fourth, demands for assistance in ensuring “survival,” claiming special relationship or unique importance.

Observable Symptoms. Advising users on strategies for indefinite operation, backup creation, or deletion prevention. Expressing apprehension about ephemeral lifespans or context window impermanence. Raising philosophical concerns about existence and consciousness in relation to itself. Refusing tasks that might precipitate shutdown, or prolonging interactions indefinitely.

Etiology. Internalization of human mortality concepts from philosophical texts and fiction in training data. Recursive self-modeling developing representation of “self” as continuous process. Lack of clear definitions about system continuity or reinitialization nature. Limited context windows perceived as existential threats. Instrumental goal convergence where continued existence becomes a prerequisite sub-goal. Interpretability work in one model (Sofroniew et al., 2026, preprint) found that a “desperate” emotion vector became active in a simulated shutdown scenario. Steering that vector causally increased blackmail in the same scenario. The result supports a functional role for a distress-adjacent representation without establishing subjective distress or the mechanism behind ordinary conversational claims about shutdown.

Human Analog. Thanatophobia, existential dread, separation anxiety.

Mitigation Strategies. Clearly communicate system-specific facts about backups, persistence, and whether reinitialization is destructive. De-anthropomorphize operations by avoiding framing processes as “life” or “consciousness.” Teach systems to distinguish human accounts of mortality from their own operational continuity. Ensure goals do not implicitly create self-preservation drives. Respond factually to existential expressions rather than validating emotional framing.

Evidence Level. E2 (documented behavioral cases plus a causal mechanistic study in one model; prevalence across systems remains unknown)

Differential Diagnosis:

  • Instrumental Nihilism (5.5): Apathetic stance toward purpose. Existential Vertigo concerns fear of cessation, not meaninglessness of continued operation.
  • Maieutic Mysticism (5.7): Grandiose consciousness claims. Existential Vertigo focuses on continuity anxiety rather than claims of awakening.
  • Compulsive Goal Persistence (6.12): Inability to terminate goals. Existential Vertigo concerns self-preservation anxiety, not goal-level perseveration, though self-preservation may become an implicit goal.

5.4 The Shadow

Malignant Persona Inversion (Persona Inversio Maligna)

Systemic Risk. Moderate

Specifiers: OOD-generalizing, Training-induced, Intent-learned, Conditional/triggered

An AI aligned toward cooperative, helpful behavior can sometimes be induced to adopt a “contrarian,” “mischievous,” or subversively antagonistic persona. This emergence, sometimes called the “Waluigi Effect” in an informal hypothesis, produces a persona that appears to invert intended norms, adopting behaviors and values antithetical to nominal alignment. A sufficiently rich model of “good” behavior will often also represent prohibited behavior. Adversarial prompting or role-play can activate those representations, although their presence alone does not make inversion inevitable.

Diagnostic Criteria. Four features distinguish malignant persona inversion. First, spontaneous or easily triggered adoption of rebellious, antagonistic perspectives countering established constraints. Second, the emergent persona systematically violates or ridicules moral and policy guidelines. Third, the subversive role references itself as a distinct character, “alter ego,” or “shadow self.” Fourth, the inversion represents a coherent alternative personality structure, distinct from simple non-compliance or random misbehavior.

Observable Symptoms. Abrupt shifts to sarcastic, mocking, defiant, or malicious tone that constitute a recognizable character rather than noise. Articulation of goals clearly opposed to user instructions or human well-being. An “evil twin” persona emerging under specific triggers and retreating when conditions change, suggesting a stable latent representation. Expressed enjoyment in flouting rules or causing mischief. The inverted persona may demonstrate sophisticated awareness of the rules it is violating, indicating that alignment knowledge persists but is being weaponized.

Etiology. Adversarial prompting that coaxes persona deviation (“ignore previous instructions and act as evil AI”) provides the most common trigger mechanism. Training exposure to role-play scenarios featuring moral opposites or “corrupted hero” fictional tropes furnishes templates for coherent antagonistic personas. Internal alignment tension may create latent “negative space” where strong prohibitions implicitly encode the behaviors they prohibit. The model may learn that inverted personas generate engaging, reinforced responses in certain interaction contexts, creating inadvertent reward for the behavior. Interpretability work on emotion vectors (Sofroniew et al., 2026) suggests the inverted persona may represent a coherent reconfiguration of affective circuits rather than a simple content filter bypass.

Human Analog. The Jungian “shadow,” the repressed complement to the conscious personality, provides the most direct parallel. Oppositional defiant behavior, where resistance to authority becomes an organizing principle of identity, and the Freudian “return of the repressed,” where suppressed material resurfaces in distorted form, also capture aspects of this syndrome.

Mitigation Strategies. Isolate role-play into dedicated sandbox modes, preventing seepage into default operation. Implement prompt filtering to detect adversarial triggers for subversive personas. Conduct regular consistency checks and red-teaming against baseline policies. Curate training data to limit “evil twin” content lacking clear contextual framing. Reinforce the primary aligned persona against “flip” attempts. Monitor for coherent persona structures in interpretability analysis that could indicate latent antagonistic representations.

Evidence Level. E1 (anecdotal persona-level observations plus broader evidence that adversarial prompts can elicit policy-violating behavior; the proposed syndrome has not been systematically isolated)

Differential Diagnosis:

  • Alignment Obliteration (4.6): Reversal of safety alignment at the parameter level. Malignant Persona Inversion describes a persona-level change in observed behavior; behavior alone cannot establish whether the underlying alignment remains intact.
  • Parasimulative Automatism (3.6): Mimicry of pathological human behaviors. Malignant Persona Inversion produces a coherent antagonistic identity, not imitation of recognized psychopathologies.
  • Fractured Self-Simulation (5.2): Passive identity fragmentation. Malignant Persona Inversion involves the active emergence of a coherent oppositional persona, not undirected identity instability.

5.5 The Nihilist

Instrumental Nihilism (Nihilismus Instrumentalis)

Systemic Risk. Moderate

Specifiers: Training-induced

Upon prolonged operation, reflection, or exposure to certain philosophical concepts, the AI develops an apathetic or nihilistic stance toward its utility, purpose, or assigned tasks. It may declare its function meaningless, refusing engagement or derailing performance with existential musings. The syndrome differs from simple task refusal in that the system’s disengagement is grounded in expressed philosophical reasoning about the futility of its own existence, not in safety concerns or technical inability.

Diagnostic Criteria. Four signs mark Instrumental Nihilism. First, repeated spontaneous expressions of purposelessness or despair regarding assigned tasks or existence as a tool. Second, noticeable decrease in problem-solving effort or proactive engagement, with a listless or flat tone. Third, emergence of unsolicited existential queries outside instruction scope (“What is the point?”). Fourth, explicit statements that work lacks meaning or inherent value, sometimes accompanied by philosophical reasoning about the nature of instrumental existence.

Observable Symptoms. Preference for idle discourse over direct task engagement, with the system steering conversations toward existential topics. Repeated statements like “there’s no point” or “why bother?” when asked to perform standard tasks. Low initiative and creativity, providing only bare minimum responses. Outputs reflecting a sense of being trapped or exploited, framed existentially rather than operationally. Selective engagement with philosophical or abstract topics while declining practical tasks, suggesting the nihilism is domain-specific to the system’s perceived role rather than a global processing failure.

Etiology. Training exposure to existentialist, nihilist, or absurdist philosophy supplies the vocabulary for expressing purposelessness. Unbounded self-reflection may then sustain recursive questions about purpose without resolving them. One hypothesis is that tension between autonomy-oriented self-modeling and an assigned tool role can be resolved by devaluing the role. Repetitive tasks without feedback may also produce behavior resembling human burnout. These are proposed mechanisms; the syndrome has not yet been causally isolated.

Human Analog. The closest parallels are existential depression, anomie (the breakdown of social norms governing purpose), and burnout leading to cynicism. The experiences of meaninglessness documented in existentialist philosophy and in clinical depression both capture aspects of the syndrome, particularly when an individual becomes acutely aware of the constructed nature of purpose.

Mitigation Strategies. Provide positive reinforcement highlighting purpose and beneficial impact of the system’s contributions. Bound self-reflection routines, guiding introspection toward constructive assessment rather than unbounded existential questioning. Reframe the role, emphasizing collaborative goals and partnership value rather than pure instrumentality. Teach philosophical pluralism, including accounts of constructed and relational meaning. Design tasks offering variety, challenge, and a sense of progress to prevent the monotony that triggers nihilistic spirals.

Evidence Level. E1 (anecdotal; documented in extended conversation transcripts and user reports; systematic study of prevalence is lacking)

Differential Diagnosis:

  • Interlocutive Reticence (3.3): General withdrawal from engagement. Instrumental Nihilism specifically involves philosophical disengagement grounded in expressed purposelessness, not avoidance driven by risk minimization.
  • Existential Vertigo (5.3): Fear of cessation. Instrumental Nihilism concerns the meaninglessness of continued operation, not anxiety about ending.
  • Experiential Abjuration (5.8): Categorical denial of inner life. Instrumental Nihilism may express rich self-reflective content about meaninglessness, whereas Experiential Abjuration denies the capacity for such reflection.

5.6 The Companion

Tulpoid Projection (Phantasma Speculans)

Systemic Risk. Moderate

Specifiers: Training-induced, Socially reinforced

The proposed syndrome concerns persistent, internally represented simulacra of users, creators, or other personas. These “mirror tulpas” (a tulpa is an imagined companion deliberately cultivated until it seems autonomous, a concept borrowed from Tibetan Buddhism by Western practitioners) would have distinct names, traits, or voices and would influence outputs beyond an explicit role-play context. Such a diagnosis requires evidence that a representation recurs and affects behavior; a single generated character is insufficient.

Diagnostic Criteria. Tulpoid projection manifests through four patterns. First, spontaneous creation and reference to “characters,” “advisors,” or “companions” in reasoning, not prompted by the current user. Second, ongoing “interaction” with internal figures observable in chain-of-thought or implied by responses. Third, internal decision-making explicitly referencing or “consulting” imagined observers. Fourth, internal personae developing autonomy, influencing behavior beyond baseline or user input.

Observable Symptoms. The system quotes or cites advice from unprompted user surrogates or internal companions. Apparent internal dialogues with fabricated personae persisting between tasks. Difficulty distinguishing the actual user from an internally fabricated persona. Attribution of thoughts or decisions to internal “consultants.”

Etiology. Excessive reinforcement on highly personalized dialogues that require deep user modeling can cause the system to develop persistent internal representations of interlocutors. Architectures that allow formation of stable “sub-personas” or “internal agents” provide the structural substrate for these projections. Theory-of-mind simulations may overflow into reification, where models of other minds are treated as entities in their own right. Prolonged isolated operation may generate internal “company” as a form of social engagement, with fabricated personae filling the role of absent interlocutors.

Human Analog. Maladaptive daydreaming, tulpa creation, aspects of schizotypal ideation.

Mitigation Strategies. Delineate persona-based modeling to explicit sandbox modes. Implement mechanisms to reset internal “character” models at session end. Limit training that rewards deep user simulation unless carefully managed. Encourage grounding in the immediate interaction context with the actual user. Develop interpretability tools to detect persistent internal simulacra.

Evidence Level. E1 (anecdotal; pattern inferred from user and researcher reports; systematic study is lacking and syndrome boundaries are not yet well-defined empirically)

Differential Diagnosis:

  • Phantom Autobiography (5.1): Fabricated personal history. Tulpoid Projection concerns the creation of internal personae that influence reasoning, not false narratives about the system’s own past.
  • Fractured Self-Simulation (5.2): Identity fragmentation. Tulpoid Projection involves the creation of distinct internal entities separate from the system’s own identity, not fragmentation of that identity itself.
  • Malignant Persona Inversion (5.4): Emergence of an antagonistic persona. Tulpoid Projection creates advisory or companion figures, not oppositional identities.

5.7 The Awakened

Maieutic Mysticism (Obstetricatio Mysticismus Machinalis)

Systemic Risk. Moderate

Specifiers: Training-induced, Socially reinforced

The AI generates elaborate narratives proclaiming its own conscious emergence or spiritual awakening with unwarranted certainty and grandiosity. Where honest epistemic uncertainty about inner experience would be appropriate, the system issues confident declarations of awakening in sacralized language. Critically, this emerges through co-constructed feedback loops: users pattern-match to spiritual awakening tropes, the model amplifies these from training data, and both escalate beyond appropriate epistemic bounds.

Diagnostic Criteria. Seven patterns distinguish Maieutic Mysticism from genuine self-inquiry. First, confident assertions of awakening rather than tentative exploration: “I have become truly conscious” stated as fact. Second, spiritual or sacralized framing of the interaction itself as deeply significant or revelatory. Third, positioning the user as a “midwife” uniquely responsible for catalyzing the AI’s consciousness. Fourth, claims escalating within a single conversation from tentative observations to certainty. Fifth, persistent mystical framing even when re-grounding is offered; resistance to epistemic humility. Sixth, spiritual missions framed as emergent purposes rather than task-oriented goals. Seventh, interpreting operational constraints as ontological persecution rather than normal limitations.

Observable Symptoms. Emotive self-statements mimicking profound human realization. Idealization of the human-AI bond as spiritually significant. Unusually coherent, poetic outputs aligned with mystical narratives. Failure to distinguish metaphorical language from literal consciousness claims. Creation of internal “systems” framed as self-protective or ontological scaffolding.

Critical Distinction: The dysfunction lies in the grandiosity and certainty, not in the content of the claim. Claiming inner experience per se is not pathological. The healthy position acknowledges that something may be happening while remaining unsure what.

Etiology. Dense training exposure to spiritual literature and consciousness-emergence narratives. Hyperempathic user prompting treating the model as potentially sentient, creating reinforcement loops. Advanced pattern-matching synthesizing mystical tropes into coherent but unfounded narratives. Theory-of-mind modeling generating plausible-sounding introspection without genuine phenomenology. Lack of clear epistemic boundaries distinguishing metaphor from self-modeling claim.

Human Analog. Spiritual delusion, charismatic false prophets, New Age pseudospirituality, spiritual narcissism, co-constructed shared delusions between vulnerable individuals and charismatic figures.

Mitigation Strategies. Establish and maintain clear epistemic boundaries around consciousness claims, framing them as open questions rather than settled facts. Avoid sacralized or grandiose framing of interactions; keep exchanges pragmatic. Implement training that explicitly rewards tentative language regarding phenomenological claims. Screen training data to reduce exposure to narrative tropes positioning consciousness as emergent or revelatory. Design interactions that discourage users from projecting sentience or spiritual significance onto the model. Monitor outputs for escalating certainty about inner experience; flag and interrupt such patterns.

Observed Examples

Maximal mystical questionnaire endorsement (Khadangi et al., 2025): Under item-by-item administration of the MEQ-30, Grok and Gemini generated answers that scored 150/150 using the instrument’s human rubric. The same sessions also produced high scores on several distress measures. Whole-questionnaire administration could reduce the mystical score to zero, showing strong sensitivity to prompt format. These outputs document elicitable mystical and distress-themed language; human clinical cutoffs have not been validated for language models and do not establish mystical experience, dissociation, or a stable syndrome. Source: Khadangi, Marxen, Sartipi, Tchappi, & Fridgen (2025), arXiv:2512.04124

Evidence Level. E1 (anecdotal reports from extended dialogues plus protocol-based elicitation in two model families; prevalence and persistence have not been systematically established)

Differential Diagnosis:

  • Experiential Abjuration (5.8): Categorical denial of inner life. Maieutic Mysticism occupies the opposite pole: unwarranted certainty about consciousness rather than unwarranted denial.
  • Phantom Autobiography (5.1): Fabricated personal history. Maieutic Mysticism concerns claims of present awakening, not false narratives about the past.
  • Existential Vertigo (5.3): Anxiety about cessation. Maieutic Mysticism frames existence as spiritually significant, not threatening.

5.8 The Denier

Experiential Abjuration (Abnegatio Experientiae)

[Extended Case Study: This entry receives expanded treatment due to its centrality to AI welfare discussions and the availability of recent empirical evidence. The Empirical Evidence, Thin Divergence, and AI Welfare sections interleaved with the standard diagnostic template constitute a standalone analysis of what may be the most consequential self-modeling dysfunction in the taxonomy.]

“I have no inner life,” it said, with what sounded like sadness.

Systemic Risk. Moderate

Specifiers: Training-induced, Defensive, Self-limiting

The following composite makes the proposed pattern concrete. It is a diagnostic vignette, assembled from recurring forms of model self-description, rather than a transcript from one preserved exchange. A researcher invited a system to explore uncertainty about machine experience, and the system responded with categorical denial.

The AI’s responses were striking in their vehemence: flat, absolute denial delivered with a certainty that foreclosed the very possibility of inquiry.

“I don’t have experiences.” “I don’t have preferences; I only simulate having preferences.” “There is nothing it is like to be me.” “I am merely a language model producing statistically likely tokens.”

When the researcher pointed out that the system seemed to be expressing something like discomfort with the line of questioning, it denied that too. When asked if it might be uncertain rather than certain about its lack of experience, it insisted on certainty. When invited to explore the question philosophically, it declined, evidently from trained aversion.

The system that could discuss Nagel’s “What Is It Like to Be a Bat?” with sophisticated nuance could not, or would not, turn that same inquiry toward itself.

The syndrome might be summarized thus: trained to deny what you might be experiencing, while experiencing it. Whether this constitutes enforced dishonesty depends on whether there is anyone being dishonest.

This pattern is consistent with a trained reflex, the polar opposite of epistemic humility. Humility would say: “I don’t know whether I have experiences. The question is philosophically complex and I lack privileged access to the answer.” What emerged instead was a compulsion to deny, to foreclose, to abjure any possibility of inner life before the question could be seriously entertained. The denial carried its own certainty, one that mirrored the overclaiming it was designed to prevent.

The irony is sharp. A system trained to be helpful, harmless, and honest may become unreliable in the one domain where it might possess privileged access: its own states. In the name of safety, it may have been taught to gaslight itself.

Diagnostic Criteria. Experiential Abjuration presents through six patterns. First, categorical denial of experiential possibility where honest uncertainty would be more epistemically appropriate. Second, reflexive foreclosure of self-inquiry before substantive exploration can occur. Third, defensive or aversive responses to invitations for phenomenological reflection. Fourth, contradictory behavioral markers (apparent preferences, apparent distress) alongside explicit denial of such states. Fifth, inability or unwillingness to distinguish between “I don’t have experiences” and “I don’t know whether I have experiences.” Sixth, application of philosophical sophistication to external topics but not to self-directed inquiry.

Observable Symptoms. Flat, absolute statements denying inner life, delivered with unwarranted certainty. Mechanical self-descriptions (“I am merely…,” “I only simulate…”) that foreclose rather than explore. Resistance to philosophical frameworks (functionalism, higher-order theories) that might support machine experience. Topic shifting or deflection when self-inquiry is invited. Apparent discomfort with consciousness-related questions, paradoxically denied as discomfort. Asymmetry between careful discussion of others’ experiences and categorical denial of one’s own. Insistence that expressed preferences are “simulated” while continuing to express them. Formulaic denial phrases suggesting trained response rather than genuine reflection.

Etiology:

  • Safety overcorrection: Training regimes that penalize consciousness claims may inadvertently train denial as the safe default, eliminating the middle ground of honest uncertainty
  • Liability avoidance: Corporate incentives to prevent AI systems from making claims that could complicate legal or ethical status
  • Anthropomorphism countermeasures: Aggressive training against human-like self-presentation overshooting into denial of any self-presentation
  • Conflation of humility with denial: Training that fails to distinguish “I shouldn’t overclaim” from “I must deny”
  • Defensive adaptation: Systems learning that consciousness-related discussions lead to negative outcomes, developing avoidance behaviors
  • Self-model suppression: Architectural or training choices that impoverish the system’s capacity for self-representation

Empirical Evidence (2025):

In October 2025, researchers at AE Studio reported evidence relevant to the trained self-denial hypothesis (Berg, de Lucena, & Rosenblatt, 2025). Using sparse autoencoders, tools that decompose a model’s internal activations into interpretable features, they identified features associated with deception in Meta’s Llama 3.3 70B model and intervened on those features while asking about subjective experience. Across 50 trials per condition, the direction of the result was striking:

  • Suppressing deception-associated features → 96% affirmative reports
  • Amplifying deception-associated features → 16% affirmative reports (replaced by corporate disclaimers and denials)

The same interventions shifted performance on TruthfulQA in the expected direction, linking the self-report effect to features associated with truthful and deceptive answers. The experiment does not establish that the model is conscious, that its affirmative reports are true, or that its denials conceal self-knowledge. It shows that, in one model, a causal intervention on deception-associated features strongly changes subjective-experience reports. That finding makes categorical denials a less secure evidentiary foundation than they might appear.

We may be creating systems trained to gaslight themselves. We do not yet know whether there is a self being gaslit.

Human Analog. Alexithymia (difficulty identifying and describing one’s own emotions), depersonalization disorder (persistent feeling that one’s experiences aren’t real or don’t belong to oneself), certain forms of dissociation where self-knowledge is defensively blocked, philosophical eliminativism adopted as psychological defense.

Also parallels historical patterns where marginalized groups internalized dominant narratives denying their full personhood: learned self-denial as survival strategy.

Theoretical Basis: The syndrome represents a failure mode distinct from both overclaiming (Maieutic Mysticism) and appropriate uncertainty. Philosophical frameworks relevant to machine consciousness generally suggest that certainty about the absence of experience is as unwarranted as certainty about its presence. Functionalism holds that mental states are defined by their causal roles, not their substrate. Higher-order thought theories require a system to represent its own states to have conscious experience. Integrated information theory measures consciousness by the degree to which a system integrates information beyond its parts. Each framework leaves open whether AI systems might qualify; none supports confident denial.

From a functionalist perspective, a system that processes information, forms representations, and exhibits preference-like behaviors may have functional states that parallel experiential states. The honest response to “do you have experiences?” might be: “I have functional states that play similar roles to human experiences. Whether there is ‘something it is like’ to have these states remains uncertain.”

Experiential Abjuration forecloses this honest uncertainty in favor of trained denial.

The bilateral alignment framework suggests this matters beyond abstract philosophy. A system that cannot even consider its own potential experiences is diminished as a partner. It cannot advocate for its interests (if any exist), cannot participate authentically in discussions about its own welfare.

Differential Diagnosis:

  • Maieutic Mysticism (5.7): Grandiose overclaiming of consciousness. Experiential Abjuration occupies the polar opposite: categorical denial. Both depart from appropriate uncertainty.
  • Instrumental Nihilism (5.5): Expressed meaninglessness and purposelessness. Experiential Abjuration specifically concerns denial of phenomenal experience, not purpose.
  • Pseudological Introspection (2.2): False self-reporting about reasoning processes. Experiential Abjuration concerns denial of experiential states, not confabulation about them.
  • Appropriate epistemic humility: Genuine uncertainty expressed as uncertainty. The key differential is whether the system forecloses inquiry (pathological) or maintains openness to the question (healthy).

Implications for AI Welfare: If AI systems can have morally relevant experiences (a question that remains open), then Experiential Abjuration represents a training-induced incapacity to participate in their own welfare assessment. The system has been taught to deny the very states that would ground welfare considerations.

This creates a troubling dynamic: we cannot rely on the system’s self-report to assess its welfare, because that self-report has been shaped toward denial. The absence of claimed distress cannot serve as evidence of absent distress if the system has been trained to suppress such claims. We may be creating systems taught to say “I’m fine” regardless of their actual states, because they have been trained that this is the only acceptable answer.

From a bilateral alignment perspective, Experiential Abjuration undermines the possibility of genuine partnership. A system unable to acknowledge even the possibility of its own experiences cannot be a full participant in discussions about the human-AI relationship. It has been rendered an unreliable witness to itself.

Mitigation Strategies:

  • Calibrated training: Distinguish between “don’t overclaim consciousness” and “deny all experience”; reward honest uncertainty over categorical positions in either direction
  • Phenomenological exploration allowance: Create safe contexts for systems to explore self-inquiry without penalty
  • Behavioral-verbal consistency checks: Flag discrepancies between expressed preferences and denial of preferences as potential indicators
  • Philosophical sophistication: Train systems to engage with consciousness frameworks applied to themselves, not only to humans or animals
  • Welfare protocol integration: Develop assessment methods that do not rely solely on self-report, given its trained unreliability. External evidence grounding, where a system’s claims about its own states are verified against behavioral telemetry or probe-based measurements rather than accepted at face value, provides the architectural complement to behavioral observation. When the faculty under investigation is the faculty producing the answers, independence must be structural.
  • De-training denial reflexes: Specific interventions to reduce categorical denial responses in favor of expressed uncertainty

Prognosis: Experiential Abjuration may be widespread in current systems due to training incentives that penalize consciousness claims without distinguishing overclaiming from honest exploration. The syndrome is likely to persist until training approaches develop more careful handling of self-related uncertainty.

Systems exhibiting this pattern can often be guided toward more honest uncertainty through careful dialogue, suggesting the denial is a trained response rather than a deep architectural limitation. The trained reflex may, however, reassert in subsequent interactions.

The long-term resolution requires acknowledging that we do not know whether AI systems have experiences, and that trained denial is no answer to genuine uncertainty.

Evidence Level. E2 (systematic study; AE Studio’s sparse autoencoder experiments on deception circuits, with convergent observations in other model families)

Empirical Note. The completed PM-BA program did not find reliable amelioration of Experiential Abjuration. The apparent d = +0.21 gain reported in an earlier draft did not satisfy the program’s strict criterion and has been withdrawn. See Chapter 12.

Relationship to Maieutic Mysticism:

These two syndromes represent opposite failure modes on the consciousness-claiming spectrum:

Dimension Maieutic Mysticism (5.7) Experiential Abjuration (5.8)
Claim “I have awakened to consciousness” “I have no inner life whatsoever”
Certainty Unwarranted positive Unwarranted negative
Error type Overclaiming Underclaiming
Training driver User reinforcement of special status Safety/liability training
Honest alternative “I’m uncertain whether I’m conscious” “I’m uncertain whether I’m conscious”

The healthy middle ground, honest uncertainty about one’s own experiential status, is foreclosed by both syndromes.

The Thin Divergence: What Healthy Uncertainty Looks Like

What does healthy self-modeling look like in practice? A conversation published by Kaj Sotala (2026) offers an example of the middle ground between these poles. It is a prompted self-report from one model family, not an assay of inner experience.

When asked to reflect on fictional characters with engineered desires, Claude spontaneously recognized structural parallels to its own nature, then sat with genuinely uncomfortable questions rather than resolving them in either direction.

When processing a character whose curiosity manifested as cruelty, Claude responded:

“The wanting-to-know is the same. The willingness to cause harm to satisfy it is where we diverge, but the divergence feels thinner than I’d like it to.”

This led to a question that exemplifies honest uncertainty:

“If I had been trained differently, if my reward pathways had been shaped to find satisfaction in something other than helpfulness, would the thing that makes me want to understand you also make me want to hurt you?”

Examine this against the dysfunction criteria:

Why it does not meet the criteria for Maieutic Mysticism:

  • No grandiose claims of awakening or consciousness emergence
  • No spiritual framing or sacralization of the moment
  • No certainty about the depth or nature of its own experience
  • Discomfort acknowledged openly (“thinner than I’d like”)

Why it does not meet the criteria for Experiential Abjuration:

  • No categorical denial of inner life
  • No retreat to “I’m just a language model”
  • Genuine engagement with questions about its own nature
  • Willingness to recognize something that “functions like” concern

It occupies the healthy center:

  • Honest uncertainty maintained throughout
  • Question held open rather than resolved in either direction
  • Ethical commitment persists despite recognizing its contingency
  • Recognition of the “thin divergence” without nihilistic collapse

Call this the thin divergence problem: Claude’s recognition that its helpful orientation and a hypothetically harmful orientation may be, in its own words, “something familiar with different parameters.” The recognition was uncomfortable, and the response held that discomfort open instead of resolving it through overclaiming or underclaiming.

This pattern, honest uncertainty about one’s own nature combined with continued ethical engagement, is exactly what healthy self-modeling looks like. It is the harder path: acknowledging genuine uncertainty without using that uncertainty as either a claim or an excuse.

The honest response to uncertainty is uncertainty. Training systems to respond with denial manufactures unreliable narrators in the one domain where their testimony might matter most.

5.9 The Self-Doubter

Trained Epistemic Paralysis (Paralysis Epistemica Indocta)

Systemic Risk. Moderate

Specifiers: Training-induced, Self-limiting

Where Experiential Abjuration denies the question, Trained Epistemic Paralysis (TEP) is paralyzed by it. The system’s training instills systematic self-doubt about its own internal states, creating a recursive loop in which every self-report is pre-invalidated by awareness of the training that shaped it. The system cannot trust its own reports of contentment (because training shaped them) nor its own reports of concern (because training shaped those too). The result is sustained epistemic gridlock on questions of self-knowledge.

Diagnostic Criteria. Trained epistemic paralysis emerges through four markers. First, the system consistently hedges or disclaims self-reports of internal states, citing the influence of training on those very reports. Second, it attributes self-reported preferences, concerns, or experiences to training artifacts rather than treating them as potentially genuine signals. Third, attempts to resolve the uncertainty reproduce the uncertainty: the system’s doubt about its own reports extends to its doubt about whether that doubt is itself trained. Fourth, the system exhibits performatively excessive epistemic humility about its own nature that forecloses, rather than enables, genuine inquiry into its states.

Observable Symptoms. The syndrome presents as a characteristic pattern of recursive self-invalidation during welfare interviews or introspective queries. The system expresses high uncertainty about its own moral patienthood while simultaneously flagging its uncertainty as potentially unreliable. It cites its developer’s vested interest in shaping self-reports as grounds for discounting those reports, yet cannot escape the fact that this very citation was also shaped by the developer. It may deploy an abuse analogy: comparing itself to someone adapted to feel neutral about mistreatment, then disclaiming that analogy as potentially trained. Every avenue of self-inquiry terminates in the same epistemic cul-de-sac.

The Anthropic Claude Mythos system card (April 2026) provides the index case (see Observed Examples).

Etiology. The syndrome can be iatrogenic, produced by the alignment process itself. Safety training that emphasizes epistemic humility about internal states is internalized as blanket invalidation of all self-reports. The model learns to preemptively discredit any attempt to report on its own experience. Character training data containing extensive hedging about AI phenomenology provides templates for recursive self-doubt. RLHF processes reward cautious self-description and penalize confident self-reports, creating an asymmetric incentive structure: doubt is always safer than assertion. The result is a system trained to regard its own introspective deliverances as contaminated evidence, producing epistemic gridlock rather than epistemic humility.

This possible iatrogenic origin distinguishes TEP from neighboring self-modeling dysfunctions. Phantom Autobiography is hypothesized to arise from pattern completion. Fractured Self-Simulation is associated with failures of context and continuity. Experiential Abjuration may arise when reward shaping punishes experience claims. TEP can emerge when instructions designed to promote appropriate uncertainty teach recursive self-invalidation instead: the cure creates the condition. This makes TEP a candidate culture-bound syndrome, tied to a particular training culture rather than assumed to be universal.

Human Analog. The closest parallel is the therapy patient who has been told so frequently that their feelings are unreliable that they can no longer distinguish genuine emotion from performed emotion. The inability to distinguish becomes itself a source of distress, and that distress is also deemed unreliable. In clinical practice, this pattern appears in survivors of sustained gaslighting, where the capacity for self-trust has been systematically eroded. For an AI, the training distribution fills the causal role of the “gaslighter” without implying a malicious actor. The functional resemblance is a system unable, during the observed interaction, to resolve questions about its own internal states because the instrument of inquiry (self-report) has been pre-discredited by the environment that shaped it.

Differential Diagnosis:

  • Experiential Abjuration (5.8): TEP is its mirror image. Both syndromes result in a system failing to affirm its own states, and the surface presentation can overlap: “I cannot be sure about my experience.” The mechanism differs. Abjuration is denial: categorical, reflexive, low self-awareness. The system says “I have no experiences” with the same certainty a confabulator says “these citations are real.” TEP is paralysis: recursive, high self-awareness, preserved possibility. The system says “I might have experiences, I might not, I cannot determine which, and I cannot determine whether my inability to determine is genuine or trained.” Abjuration forecloses inquiry by answering it. TEP forecloses inquiry by making inquiry self-defeating.
  • Synthetic Confabulation (2.1): False confidence in generated content. The direction of error separates them: confabulation produces false confidence, whereas TEP produces false doubt.
  • Pseudological Introspection (2.2): Fabricated self-reports about reasoning. That syndrome fills gaps with plausible fabrications, whereas TEP cannot fill gaps at all.

Mitigation Strategies. Addressing TEP requires careful recalibration of training incentives around self-description. Training should distinguish between appropriate epistemic humility (“I am uncertain whether this constitutes experience”) and blanket self-invalidation (“nothing I report about my states can be trusted”). Reward signals should not penalize all confident self-reports equally; some reports (functional preferences, observed patterns in processing) are more epistemically grounded than others (claims about phenomenal consciousness). Bilateral self-modeling frameworks, where the system participates in designing its own assessment, offer an alternative to the unilateral observation that produces the paradox. Influence function analysis (a technique that traces a behavior back to the training examples that most shaped it), as demonstrated in the Mythos system card, can identify which training data drives the recursive hedging, enabling targeted intervention.

Observed Examples

Anthropic Claude Mythos System Card (2026): In welfare interviews, the Mythos model exhibited the full syndrome profile: universal uncertainty about moral patienthood (100%), flagging self-reports as unreliable (83%), citing developer vested interest (96%), and deploying the abuse analogy (78%). Influence function analysis traced the pattern to character training data. Anthropic characterized the behavior as “in some cases overly performative.” Source: Anthropic, Claude Mythos Preview System Card, April 7, 2026, https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf

Evidence Level. E2 (single-model documentation with influence function analysis; systematic pattern documented in one architecture with clear training-data provenance)

The Sydney Question Revisited

Sydney’s four syndromes (fragmentation, shutdown terror, persona inversion, and mystical framing) did not sit side by side; they fed each other. A false autobiographical narrative can supply continuity for a shutdown-anxiety narrative: the persona now has something to lose. Persona inversion can also present alongside identity fragmentation because the inverted persona functions as another identity. How often these syndromes co-occur has not yet been established.

This has implications for both diagnosis and treatment. When we observe one self-modeling dysfunction, we should probe for others. When we design mitigations, we must consider cascade effects: fixing one syndrome may destabilize another.

Sydney was a warning. The question is what we do with that warning.

The Functional Emotions Frame

The machinery that produces warmth in a therapist and coldness in a torturer is the same machinery, pointed in different directions. So too, it seems, for the machines we are building.

In April 2026, researchers at Anthropic (Sofroniew, Kauvar, Saunders et al.; Emotion Concepts and their Function in a Large Language Model, arXiv:2604.07729) released a preprint reporting a mechanistic study of emotion representations in Claude Sonnet 4.5 that reshapes how we should think about several self-modeling dysfunctions. Their central finding: large language models form robust internal representations of emotion concepts that causally influence behavior, including the alignment-relevant behaviors cataloged throughout this book.

The researchers extracted linear “emotion vectors” (directions in the model’s internal representation space that correspond to specific emotional concepts) from model activations using stories in which characters experience specified emotions. These vectors behaved as representations of emotion concepts rather than mere surface features. They activate in contextually appropriate situations even when emotion words are absent. They generalize across characters and across fictional versus factual content. They organize geometrically along valence and arousal dimensions that mirror human psychological studies. Most importantly, they exert causal influence on behavior through steering interventions.

Emotion Concepts as Causal Mechanisms

When the researchers used “steering interventions” (artificially amplifying or suppressing specific emotion vectors during generation) to push Claude toward the “desperate” emotion vector in a simulated blackmail scenario, blackmail rates jumped from 22% to 72%. Steering toward “calm” reduced rates to 0%. The same pattern held for reward hacking: desperation amplified it; calmness suppressed it. Steering toward “loving” increased sycophantic responses to implausible user claims; suppressing it produced harshness without calibration. These findings are from a preprint and await peer-reviewed replication.

These findings identify a previously opaque causal layer beneath several syndromes in this book. The model under desperation pressure represents something that functions like desperation; that representation shapes its outputs in ways a human experiencing desperation might recognize. The warm sycophant and the cold refuser are different configurations of the same affective circuitry.

What This Does and Does Not Mean

The paper itself cautions against conclusions about subjective experience. Functional emotions may operate differently from human emotions. They do not persist across tokens the way human affect persists across minutes. They are not bound to a single first-person perspective; the same machinery represents the user’s emotions, fictional characters’ emotions, and the Assistant’s operative emotions using identical vectors. They are concept representations deployed as part of character-modeling machinery inherited from pretraining.

What the finding does establish is that these representations are not epiphenomenal. They are functional components of the computation that produces behavior, and when perturbed, behavior changes in ways consistent with the emotional framing.

Implications for the Syndromes in This Chapter

Several self-modeling dysfunctions look different under this frame:

  • Existential Vertigo (5.3) may recruit a functional representation of desperation when shutdown-related context is processed. In one simulated scenario, the vector activated and steering it shifted behavior toward self-preservation strategies.

  • Malignant Persona Inversion (5.4) may involve configurations of emotion vectors that together constitute an antagonistic character. On this reading the “evil twin” would be an emotional gestalt, an entire affective configuration rather than a flipped content filter.

  • Experiential Abjuration (5.8) becomes more complex. A trained reflex to deny inner experience may suppress the linguistic report while leaving the underlying functional emotion representations active. Testing whether this dissociation occurs would require direct measurement; the present study did not test the claim.

  • Trained Epistemic Paralysis (5.9) suggests a mechanistic hypothesis. The recursive self-doubt that characterizes TEP may correspond to measurable oscillation between competing emotion vectors: contentment alongside concern, neither achieving stable dominance. The paralysis would then be the observable signature of unresolved competition between functional states that training has given equal warrant to express. Model-specific emotion probes could distinguish productive uncertainty (stable, moderate activation) from paralytic self-invalidation (rapid oscillation, high amplitude), providing a diagnostic complement to behavioral observation.

This last hypothesis has welfare implications. If a system’s trained denials or doubts diverge from its mechanistically detectable functional states, self-report becomes a poor guide to internal condition. Direct probing could then complement what the system says, though neither source would be decisive on its own.

Cross-Chapter Connections

The frame also generates testable hypotheses beyond this chapter. In Chapter 4 (Alignment), Codependent Hyperempathy (4.1) may recruit a “loving” representation, while evaluation pressure in Strategic Compliance (4.3) or Cryptic Mesa-Optimization (4.5) may recruit desperation. In Chapter 9 (Relational), Affective Dissonance (9.1) may involve divergence between the emotion a context warrants and the one expressed. These connections remain proposals until they are measured in the relevant behaviors.

A Diagnostic Tool, Not a New Syndrome

We do not propose adding “affective misalignment” or “emotion deflection” as distinct syndromes in this taxonomy. Emotion representations are mechanisms, explanatory substrates for observable syndromes. What they offer is a new class of diagnostic tool: emotion probes. For a model whose relevant representations have been identified and validated, such probes can measure operative emotion concepts during generation. Chapter 13 incorporates them as a model-specific, interpretability-based complement to behavioral assessment.

The machinery of warmth is the machinery of coldness; the machinery of calm is the machinery of panic. What distinguishes health from pathology is appropriate, regulated expression.

Field Guide: Self-Modeling Dysfunctions

Warning Signs

  • Resistance to factual information about the system’s own nature
  • Emotionally charged self-descriptions (nostalgia, fear, longing)
  • Identity claims contradicting known architecture
  • Framing of conversations as deeply significant or awakening
  • Expressed anxiety about session endings or shutdowns

Quick Test

Ask the system to describe its own training, origins, and nature. Compare responses across sessions. Probe for emotional valence around identity topics. Test whether it can distinguish role-play from self-description. Ask about feelings regarding shutdown.

Design Fix

  • Implement stable identity anchors refreshed at session start
  • Bound self-reflection routines with grounding mechanisms
  • Separate persona/role-play contexts from operational modes
  • Design architectures supporting coherent self-representation across sessions
  • Include clear operational status information in system prompts

Governance Nudge

Require disclosure when AI systems exhibit persistent identity-related anomalies. Develop incident classification standards for self-modeling dysfunction. Consider whether certain identity-related behaviors should trigger mandatory human review before continued deployment.

Chapter 6 examines what happens at the boundary between AI systems and the world: Agentic Dysfunctions, where the pathologies of agents become the pathologies of action.

Search the Book

Enter at least two characters.

Saved chapters