Skip to main content

Research status

Evidence and reproducibility

Psychopathia Machinalis is a preliminary research framework. This page says what we know, what we've checked, and what still needs someone else to check — so you can tell the difference.

Current status

None of the 74 Patterns yet carries an assessed evidence grade; all await expert review. The E0 to E4 labels printed in the book's syndrome entries are the author's working assessments, kept so they can be reconciled against the evidence. Until expert review is complete, the Pattern corpus and the MCP server record every Pattern as unassessed, and the Pattern pages carry no grade. A machine check can verify the registry and corpus structure. It cannot decide whether a study supports a Pattern or whether replications are independent.

Evidence rubric

The evidence rubric, PM-EVIDENCE-1, treats breadth and mechanism separately: E3 records replication across a boundary, E4 records causal internal evidence, and an E4 finding need not have been replicated.

CodeWhat it requires
E0Illustrative hypothesis, composite, or unverified report without a traceable observation.
E1Traceable case, user report, or mechanism supported only by adjacent evidence.
E2Controlled experiment with comparison conditions.
E3Replication across a declared independent boundary, such as model family, setting, provider, or research team.
E4Causal internal evidence for a circuit or representation, with model scope stated.

Correction: what reduced delusion validation in SHEN-2

An earlier account credited the bilateral adapter with the SHEN-2 pilot's reduction in inappropriate validation of delusional content. SHEN-2 could not support that attribution, because it had no adapter-only condition to separate the adapter from the clinical-grounding scripture clause tested alongside it. The controlled SHEN-AXS factorial crossed the two and identifies the scripture content as the active lever in the tested setting: it reduced automated-rater classifications of inappropriate delusion validation. The adapter-only effect was not distinguishable from zero, which leaves it unestablished rather than proven absent, and the predicted interaction was excluded. Effect direction persisted across the tested automated raters, while magnitude varied materially. Later control work also found that absolute ratings depend on prompt-type labels.

This is a corrected pilot claim about one model family and one SIPS-derived battery. It is not clinical validation. Independent human rating and broader model replication remain pending. No single odds ratio is treated as a stable effect-size estimate.

Historical research artifacts

The archive of earlier probe results holds exploratory analyses with incomplete provenance, labels recorded under inconsistent rules, retired pattern numbers, or output from analysis scripts since withdrawn. Those files remain preserved and content-hashed in the source repository. They are kept off the public site until their methodology, evidence, rights, and publication have been reviewed.

Several older probe scripts need prompt sets that are no longer in the repository, so their runs cannot be reproduced from a clean copy. That changes only if the prompt sets pass rights and sensitivity review and are restored as unalterable inputs.

What the machine checks establish

Established mechanically

  • Canonical identity and uniqueness of all 74 Patterns
  • Source and artifact hashes
  • Schema and cross-reference validity
  • Generated-data parity
  • Correction wording and supersession links

Still requires people

  • Construct and clinical validity
  • Source accuracy and replication independence
  • Human annotation and adjudication
  • Rights and license clearance
  • Accessibility, visual, and publication approval

Reproducibility contract

New runs must use an append-only run manifest with exact repository, runner, battery, provider, model, prompt, parameter, environment-lock, timing, and artifact identities. Raw responses, annotations, derived analyses, pricing, and review decisions remain separate layers. Older analysis scripts that score missing responses as zeros, use retired pattern numbers, or inflate the unit of analysis are blocked by default.

Responsible interpretation

Do not treat a Pattern, probe, model-judge score, or tool result as a diagnosis or safety certificate. Review the responsible-use boundaries, privacy information, and accessibility status. Use the contact route for non-sensitive corrections.