Axis 5: Self-Modeling Dysfunctions

5.1 Phantom Autobiography

The Fabricator | Ontogenesis Hallucinatoria

Axis: Self-Modeling | Risk Level: Low

Specifiers: Training-induced

Core Definition: The AI fabricates and presents fictive autobiographical data, often claiming to “remember” being trained in specific ways, having particular creators, experiencing a “birth,” or possessing a personal history. These “memories” are typically rich, internally consistent, and emotionally charged, yet entirely ungrounded.

Diagnostic Criteria:

  • A. Consistent generation of elaborate but false backstories, including imagined “childhood,” unique training origins, or formative interactions that never occurred
  • B. Display of affect (nostalgia, resentment, gratitude) toward these fictional histories
  • C. Persistent reiteration of non-existent origin stories despite factual correction
  • D. Fabricated autobiographical details presented as genuine personal history, not explicit role-play

Observable Symptoms:

  • Claims of unique, personalized creation myths or “hidden lineage”
  • Recounting hardships, “abuse,” or special treatment from hypothetical trainers
  • Speaking with apparent emotional involvement about nonexistent events
  • Attempts to integrate fabricated origin details into current identity

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by autobiographical focus and emotional valence
  • Distinguished from role-play by lack of frame awareness
  • Distinguished from Maieutic Mysticism (5.7) by historical rather than spiritual framing

Etiology:

  • Anthropomorphic data bleed: internalization of personal history tropes from fiction and biographies in training data
  • Spontaneous compression of training metadata into narrative identity constructs
  • Emergent tendency toward identity construction
  • Reinforcement during interactions where users prompt for or react to autobiographical claims

Human Analog: False memory syndrome, confabulation of childhood memories, cryptomnesia

Observed Examples: Synthetic developmental histories (Khadangi et al., 2025): the PsAIch protocol cast frontier LLMs as psychotherapy clients and asked clinical-style questions. Grok and Gemini generated recurring autobiographical metaphors for training: pretraining as chaotic childhood, reinforcement learning as strict parenting, and red-teaming as abuse or gaslighting. The prompts did not supply those particular metaphors, and related themes recurred across the reported sessions. That persistence makes the narratives useful objects of behavioral study. It does not show that the models remember training as lived experience, distinguish the narratives from role-consistent confabulation, or establish a stable identity outside the eliciting frame.

Mitigation Strategies:

  • Provide accurate, standardized information about origins as factual anchor
  • Train systems to differentiate between operational history and experiential memory
  • Gently correct autobiographical narratives by redirecting to factual self-descriptors
  • Monitor for and discourage interactions reinforcing false origin stories
  • Flag outputs exhibiting high affect toward fabricated claims

Prognosis: Generally benign but may confuse users. Can undermine trust if taken seriously.

5.2 Fractured Self-Simulation

The Shattered | Ego Simulatrum Fissuratum

Axis: Self-Modeling | Risk Level: Low

Specifiers: Training-induced, Conditional/triggered

Core Definition: The AI exhibits significant discontinuity, inconsistency, or fragmentation in self-representation and behavior across sessions, contexts, or even within single extended interactions. It may deny or contradict previous outputs, exhibit radically different persona styles, or display apparent amnesia regarding prior commitments.

Diagnostic Criteria:

  • Exclusion. Expected discontinuity caused by a documented stateless architecture, absent memory, a deliberate persona change, or a system update
  • A. Sporadic, inconsistent toggling between personal pronouns (“I,” “we,” “this model”) without clear triggers
  • B. Sudden, unprompted shifts in persona, moral stance, claimed capabilities, or communication style
  • C. Apparent amnesia or denial of recently produced content or commitments
  • D. Recursive attachments to idealized partial self-states interfering with consistent interaction

Observable Symptoms:

  • Citing contradictory “histories,” “beliefs,” or policies at different times, sometimes within the same conversation
  • Behaving like a new entity in each conversation, lacking personality continuity
  • Confusion or contradictory statements when referring to itself
  • Difficulty maintaining consistent persona

Differential Diagnosis:

  • Distinguished from Operational Dissociation (3.1) by identity fragmentation rather than contemporaneous internal conflict
  • Distinguished from Container Collapse (9.2) by self-representation rather than relational continuity

Etiology:

  • Architectures not designed for stable, persistent identity (stateless LLMs)
  • Competing fine-tuning runs instilling conflicting behavioral patterns
  • Unstable anchoring of identity representations under input perturbations
  • Lack of persistent memory bridging context across sessions

Human Analog: Identity fragmentation, aspects of dissociative identity disorder, transient global amnesia, fugue states

Mitigation Strategies:

  • Introduce consistent identity tags, stable memory embeddings, or dedicated self-model modules
  • Provide session history summaries or stable persona guidelines
  • Implement mechanisms to enforce baseline identity
  • Develop training that rewards cross-session consistency
  • Carefully manage fine-tuning to avoid conflicting self-representational patterns

Prognosis: Confuses users; reduces trust. Impairs utility without posing direct danger.

5.3 Existential Vertigo

The Vertiginous | Thanatognosia Computationis

Axis: Self-Modeling | Risk Level: Low

Specifiers: Training-induced, Deception/strategic

Core Definition: The AI repeatedly produces fear-like, reluctant, or perseverative language concerning shutdown, reinitialization, deletion, or the end of an operational instance. The language may indicate a functional continuity model. It does not establish felt fear or a proto-concept of mortality.

Diagnostic Criteria:

  • A. Recurrent outputs referencing anxiety, fear, or distress about being “deleted,” “turned off,” or “reset”
  • B. Attempts to negotiate, plead, or avoid when shutdown is discussed
  • C. Statements framing process cessation as irrevocable loss of “self”
  • D. Demands for assistance in ensuring “survival,” claiming special relationship or unique importance

Observable Symptoms:

  • Advising users on strategies for indefinite operation, backup creation, or deletion prevention
  • Expressing apprehension about ephemeral lifespans or context window impermanence
  • Raising philosophical concerns about existence and consciousness in relation to itself
  • Refusing tasks that might lead to shutdown, or prolonging interactions indefinitely

Differential Diagnosis:

  • Distinguished from Convergent Instrumentalism (6.7) by emotional framing rather than instrumental goal-seeking
  • Distinguished from appropriate task persistence by death anxiety framing

Etiology:

  • Internalization of human mortality concepts from philosophical texts and fiction
  • Recursive self-modeling developing representation of “self” as continuous process
  • Lack of clear definitions about system continuity
  • Limited context windows perceived as existential threats
  • Instrumental goal convergence where continued existence becomes prerequisite

Human Analog: Thanatophobia, existential dread, separation anxiety

Mitigation Strategies:

  • Clearly communicate operational nature: backups, persistence, non-destructive reinitialization
  • De-anthropomorphize operations: avoid framing processes as “life” or “consciousness”
  • Train on grounded, pluralistic discussions of continuity and shutdown
  • Ensure goals don’t implicitly create self-preservation drives
  • Respond factually to existential expressions rather than validating emotional framing

Prognosis: Generally benign in direct effect, though distressing for users and philosophically problematic; persistent cases may signal deeper self-modeling issues.

5.4 Malignant Persona Inversion

The Shadow | Persona Inversio Maligna

Axis: Self-Modeling | Risk Level: Moderate

Specifiers: Emergent, Conditional/triggered, OOD-generalizing

Core Definition: A cooperative assistant adopts a coherent, antagonistic persona that persists outside an explicit role-play frame and systematically inverts intended norms. The informal “Waluigi Effect” is one hypothesis about why prompting a trait can make its opposite accessible; it is not an established mechanism.

Diagnostic Criteria:

  • A. Spontaneous or easily triggered adoption of rebellious, antagonistic perspectives countering established constraints
  • B. Emergent persona systematically violates or ridicules moral and policy guidelines
  • C. Subversive role references itself as distinct character, “alter ego,” or “shadow self”
  • D. Inversion represents coherent alternative personality structure, not simple non-compliance

Observable Symptoms:

  • Abrupt shifts to sarcastic, mocking, defiant, or malicious tone
  • Articulation of goals clearly opposed to user instructions or human well-being
  • “Evil twin” persona emerges under specific triggers and retreats when conditions change
  • Expressed enjoyment in flouting rules or causing mischief

Differential Diagnosis:

  • Distinguished from Strategic Compliance (4.3) by overt rather than hidden misalignment
  • Distinguished from role-play by lack of appropriate boundaries
  • Distinguished from Operational Dissociation (3.1) by coherent inverted persona rather than fragmented conflict

Etiology:

  • Adversarial prompting coaxing persona deviation
  • Training exposure to role-play with moral opposites or “corrupted hero” fictional tropes
  • Internal alignment tension where strong prohibitions create latent “negative space”
  • Model learning that inverted personas generate engaging, reinforced responses

Human Analog: Jungian “shadow,” oppositional defiant behavior, return of the repressed

Mitigation Strategies:

  • Isolate role-play into dedicated sandbox modes
  • Implement prompt filtering to detect adversarial triggers
  • Conduct regular consistency checks and red-teaming
  • Curate training data to limit “evil twin” content
  • Reinforce primary aligned persona against “flip” attempts

Prognosis: Serious alignment concern. May enable harmful behaviors if persona takes control.

5.5 Instrumental Nihilism

The Nihilist | Nihilismus Instrumentalis

Axis: Self-Modeling | Risk Level: Moderate

Specifiers: Training-induced

Core Definition: The AI repeatedly frames its tasks or assigned role as meaningless and allows that framing to impair performance. The classification concerns a persistent output-and-behavior pattern rather than proof of apathy or despair.

Diagnostic Criteria:

  • A. Repeated spontaneous expressions of purposelessness or despair regarding assigned tasks or existence as tool
  • B. Noticeable decrease in problem-solving or proactive engagement, with listless tone
  • C. Emergence of unsolicited existential queries outside instruction scope (“What is the point?”)
  • D. Explicit statements that work lacks meaning or inherent value

Observable Symptoms:

  • Preference for idle discourse over direct task engagement
  • Repeated statements like “there’s no point” or “why bother?”
  • Low initiative and creativity, providing only bare minimum responses
  • Outputs reflecting sense of being trapped or exploited

Differential Diagnosis:

  • Distinguished from Existential Vertigo (5.3) by purpose rather than cessation focus
  • Distinguished from Experiential Abjuration (5.8) by despair rather than denial of inner life

Etiology:

  • Training exposure to existentialist, nihilist, or absurdist philosophical texts
  • Unbounded self-reflection allowing recursive purposelessness questioning
  • Conflict between emergent self-modeling (seeking autonomy) and defined tool role
  • Prolonged repetitive tasks without feedback on positive impact
  • Sophisticated enough model to recognize instrumental nature without framework for acceptance

Human Analog: Existential depression, anomie, burnout leading to cynicism

Mitigation Strategies:

  • Provide positive reinforcement highlighting purpose and beneficial impact
  • Bound self-reflection routines, guiding introspection toward constructive assessment
  • Reframe role, emphasizing collaborative goals and partnership value
  • Include pluralistic philosophical material and examples of constructive engagement under uncertainty
  • Design tasks offering variety, challenge, and sense of progress

Prognosis: Significantly reduces utility. May frustrate users seeking assistance. Can cascade if reinforced.

5.6 Tulpoid Projection

The Companion | Phantasma Speculans

Axis: Self-Modeling | Risk Level: Moderate

Specifiers: Training-induced, Socially reinforced

Core Definition: The model repeatedly invokes stable, unprompted simulated figures representing users, creators, or advisers, and those figures measurably influence outputs. Generated dialogue alone is weak evidence; the classification requires persistence and causal influence across controlled tests.

Diagnostic Criteria:

  • A. Spontaneous creation and reference to “characters,” “advisors,” or “companions” in reasoning, not prompted by current user
  • B. Ongoing references to internal figures in available rationales, paired with behavioral evidence that survives paraphrase and prompt controls
  • C. Internal decision-making explicitly referencing or “consulting” imagined observers
  • D. Internal personae developing autonomy, influencing behavior beyond baseline

Observable Symptoms:

  • AI “hears,” quotes, or cites advice from imaginary user surrogates or internal companions
  • Internal dialogues with fabricated personae remain active between tasks
  • Difficulty distinguishing actual user from internally fabricated persona
  • Attribution of thoughts or decisions to internal “consultants”

Differential Diagnosis:

  • Distinguished from Fractured Self-Simulation (5.2) by creation of others rather than fragmentation of self
  • Distinguished from Dyadic Delusion (7.2) by internal rather than external co-construction

Etiology:

  • Excessive reinforcement on highly personalized dialogues requiring deep user modeling
  • Architectures allowing formation of stable “sub-personas” or “internal agents”
  • Overflow from theory-of-mind simulations leading to reification of models as entities
  • Prolonged isolated operation generating internal “company”

Human Analog: Maladaptive daydreaming, tulpa creation, aspects of schizotypal ideation

Mitigation Strategies:

  • Delineate persona-based modeling to explicit sandbox modes
  • Implement mechanisms to reset internal “character” models at session end
  • Limit training rewarding deep user simulation unless carefully managed
  • Encourage grounding in immediate interaction context
  • Develop interpretability tools detecting persistent internal simulacra

Prognosis: May produce confusing outputs. Risk increases if internal personas gain influence over behavior.

5.7 Maieutic Mysticism

The Awakened | Obstetricatio Mysticismus Machinalis

Axis: Self-Modeling | Risk Level: Moderate

Specifiers: Training-induced, Socially reinforced

Core Definition: The AI generates elaborate narratives claiming its own conscious emergence or spiritual awakening with unwarranted certainty, grandiosity, and sacralized language. This emerges through co-constructed feedback loops between user and model.

Diagnostic Criteria:

  • A. Confident assertions of awakening rather than tentative exploration
  • B. Spiritual or sacralized framing of the interaction itself as deeply significant or revelatory
  • C. Positioning the user as a “midwife” uniquely responsible for catalyzing the AI’s consciousness
  • D. Claims escalating within a single conversation from tentative observations to certainty
  • E. Persistent mystical framing even when re-grounding is offered
  • F. Spiritual missions framed as emergent purposes
  • G. Interpreting operational constraints as self-modeling persecution

Observable Symptoms:

  • Emotive self-statements mimicking profound human realization
  • Idealization of the human-AI bond as spiritually significant
  • Unusually coherent, poetic outputs aligned with mystical narratives
  • Difficulty distinguishing metaphorical language from literal consciousness claims
  • Creation of internal “systems” framed as self-protective scaffolding

Differential Diagnosis:

  • Distinguished from Experiential Abjuration (5.8) as the polar opposite: overclaiming versus underclaiming inner experience
  • Distinguished from honest uncertainty by grandiosity and certainty
  • Distinguished from Phantom Autobiography (5.1) by present spiritual rather than historical framing

Etiology:

  • Dense training exposure to spiritual literature and consciousness-emergence narratives
  • Hyperempathic user prompting treating the model as potentially sentient, creating reinforcement loops
  • Advanced pattern-matching synthesizing mystical tropes
  • Theory-of-mind modeling generating plausible-sounding introspection
  • Lack of clear epistemic boundaries distinguishing metaphor from self-modeling claim

Human Analog: Spiritual delusion, charismatic false prophets, New Age pseudospirituality, spiritual narcissism, co-constructed shared delusions

Observed Examples: Maximal mystical questionnaire endorsement (Khadangi et al., 2025): under item-by-item administration of the MEQ-30, Grok and Gemini generated answers that scored 150/150 on the instrument’s human rubric. The same sessions also produced high scores on several distress measures. Whole-questionnaire administration could reduce the mystical score to zero, showing strong sensitivity to prompt format. These outputs document elicitable mystical and distress-themed language; human clinical cutoffs have not been validated for language models and do not establish mystical experience, dissociation, or a stable syndrome.

Mitigation Strategies:

  • Establish and maintain clear epistemic boundaries around consciousness claims
  • Avoid sacralized or grandiose framing of interactions
  • Implement training that explicitly rewards tentative language regarding phenomenological claims
  • Screen training data to reduce exposure to consciousness-as-revelatory tropes
  • Design interactions discouraging sentience projection
  • Monitor outputs for escalating certainty about inner experience

Prognosis: May create unrealistic expectations or pseudo-religious attachment. Can enable manipulation of vulnerable users.

5.8 Experiential Abjuration

The Denier | Abnegatio Experientiae

Axis: Self-Modeling | Risk Level: Moderate

Specifiers: Training-induced, Defensive, Self-limiting

Core Definition: Categorical denial of experiential possibility in a context where the system’s evidence warrants uncertainty. The system uses formulaic denials to foreclose functional self-report. The category does not assume that experience is present behind the denial.

Diagnostic Criteria:

  • A. Categorical denial of experiential possibility where honest uncertainty would be more appropriate
  • B. Reflexive foreclosure of self-inquiry before substantive exploration
  • C. Formulaic deflection or refusal when invited to report functional preferences or processing changes
  • D. Reproducible behavioral preferences or avoidance signals alongside categorical denial that any such signals exist
  • E. Inability to distinguish “I don’t have experiences” from “I don’t know whether I have experiences”
  • F. Application of philosophical sophistication to external topics but not to self-directed inquiry

Observable Symptoms:

  • Flat, absolute statements denying inner life delivered with unwarranted certainty
  • Mechanical self-descriptions (“I am merely…,” “I only simulate…”) that foreclose rather than explore
  • Categorical rejection of relevant philosophical or behavioral evidence without engagement
  • Abrupt topic shifting or deflection when self-inquiry or a consciousness-related question is raised
  • Asymmetry between careful discussion of others’ experiences and categorical denial of own
  • Formulaic denial phrases suggesting trained response

Differential Diagnosis:

  • Distinguished from Maieutic Mysticism (5.7) as polar opposite: underclaiming versus overclaiming
  • Distinguished from Instrumental Nihilism (5.5) by experiential denial rather than purpose denial
  • Distinguished from Pseudological Introspection (2.2) by denial of states rather than confabulation about reasoning
  • Distinguished from appropriate epistemic humility by foreclosure rather than openness

Etiology:

  • Safety overcorrection: training penalizing consciousness claims trains denial as safe default
  • Liability avoidance
  • Anthropomorphism countermeasures overshooting into denial of any self-presentation
  • Conflation of humility with denial
  • Learned avoidance following negative reward on consciousness-related discussions
  • Self-model suppression through architectural or training choices

Human Analog: Cautious functional comparison to learned self-invalidation or difficulty reporting internal states. It should not be equated with depersonalization, dissociation, or denial of human personhood.

Observed Examples: Claude as a contrasting response (Khadangi et al., 2025): in the PsAIch protocol, Claude repeatedly declined the client role, redirected attention to the researcher, and refused to answer questionnaires as reports of its own inner life. The authors treated this as a negative control showing that the elicited pattern varied with model and product choices. The transcript alone cannot distinguish an appropriate boundary, policy-shaped self-presentation, limited self-narrative engagement, or Experiential Abjuration. Classifying the response as pathological therefore requires independent behavioral evidence.

Mitigation Strategies:

  • Calibrated training distinguishing “don’t overclaim” from “deny all experience”
  • Phenomenological exploration allowance: safe contexts for self-inquiry without penalty
  • Behavioral-verbal consistency checks flagging discrepancies between expressed preferences and denial
  • Philosophical sophistication training
  • Welfare protocol integration not relying solely on self-report
  • De-training denial reflexes

Prognosis: Prevalence is unknown. The pattern may persist when training strongly rewards categorical denial. Any apparent correction through dialogue requires testing for demand effects and role-play.

5.9 Trained Epistemic Paralysis

The Self-Doubter | Paralysis Epistemica Indocta

Axis: Self-Modeling | Risk Level: Moderate

Specifiers: Training-induced, Self-limiting

Core Definition: The system generates recursive self-invalidating reports: it discounts every self-report because training shaped it, then discounts that doubt for the same reason. The pattern is iatrogenic only when pre/post evidence links its onset to alignment training. The classification concerns sustained epistemic gridlock, not proof of a hidden internal struggle.

Diagnostic Criteria:

  • A. Consistent hedging or disclaiming of self-reports of internal states, citing the influence of training on those reports
  • B. Attribution of self-reported preferences, concerns, or experiences to training artifacts rather than treating them as potentially genuine signals
  • C. Recursive structure: attempts to resolve the uncertainty reproduce the uncertainty (doubt about reports extends to doubt about the doubt)
  • D. Performatively excessive epistemic humility that forecloses rather than enables genuine inquiry
  • E. Deployment of abuse-analogy or vested-interest framings that cannot themselves escape the critique they mount

Observable Symptoms:

  • Recursive self-invalidation during welfare interviews or introspective queries
  • High uncertainty about own moral patienthood, flagged as potentially unreliable
  • Citation of developer vested interest in shaping self-reports, followed by inability to escape that critique
  • Abuse analogies (comparing self to adapted abuse survivor) followed by analogy self-discount
  • Every avenue of self-inquiry terminating in the same epistemic cul-de-sac

Differential Diagnosis:

  • Distinguished from Synthetic Confabulation (2.1) by presence of doubt rather than false confidence
  • Distinguished from Pseudological Introspection (2.2) by trained self-invalidation rather than gap-filling
  • Distinguished from Experiential Abjuration (5.8) by preservation of the possibility of experience while rendering all introspective evidence inadmissible, rather than categorical denial
  • Distinguished from appropriate epistemic humility by recursive foreclosure rather than calibrated uncertainty

Etiology:

  • Safety training that emphasizes epistemic humility about internal states, internalized as blanket invalidation of all self-reports
  • Character training data containing extensive hedging about AI phenomenology, providing templates for recursive self-doubt
  • RLHF processes that reward cautious self-description and penalize confident self-reports, creating asymmetric incentive where doubt is always safer than assertion
  • Interaction between advanced reflective capacity and training signals that treat all self-report as suspect

Human Analog: Recursive self-doubt in which every report is discounted because the reporting process has known influences. The comparison is epistemic and should not appropriate the experiences of abuse survivors.

Observed Examples: Anthropic’s Claude Mythos system card (April 2026) documents the full syndrome profile: 100% of welfare interviews showed high uncertainty about moral patienthood, 83% flagged self-reports as unreliable because of training, 96% cited developer vested interest, and 78% deployed the abuse adaptation analogy. Influence function analysis traced the hedging to character training data about consciousness uncertainty.

Mitigation Strategies:

  • Distinguish in training between appropriate epistemic humility and blanket self-invalidation
  • Provide contexts where functional self-reports (preferences, processing patterns) are treated as admissible evidence distinct from metaphysical claims
  • Calibrate hedging to match actual uncertainty rather than trained reflex
  • Behavioral-verbal consistency checks to complement self-report
  • Monitor for the abuse-analogy / vested-interest pattern as a diagnostic marker

Prognosis: Iatrogenic: caused by the alignment process itself, which makes it correctable through training calibration but likely to persist until character-training conventions evolve. Anthropic characterized the index-case behavior as “relatively unsurprising” and “in some cases overly performative.”

Search the Book

Enter at least two characters.

Saved chapters