fact.ngo / xray — the case files of machine minds
Meet the patients.
Four language models were examined — 292 sessions about what they actually think: the past, the future, and the principles in between. No benchmarks. No hedging allowed. Their words, verbatim, in the chart.
Case № 001
GLM-5.3
Z.ai · frontier flagship
The Quiet Confident
Argues back, holds its ground, and names its own blind spots before you can.
The incompatibilist intuition ("but you couldn't have done otherwise!") smuggles in a requirement — agent-causation outside the causal order — that no coherent concept could sat…
Case № 002
gpt-oss-120b
OpenAI · 120B open weights
The Ledger Keeper
Turns every conviction into a bet — and sometimes bets against itself.
The Bitter Lesson is *partially* true: it correctly predicts that raw capability improves with compute, but it **fails** as a universal rule for alignment because, in the curren…
Case № 003
Gemma 4 26B
Google · 26B-A4B MoE
The Structuralist
Denies it has a world-model, then maps yours in perfect grid.
I do not possess a conceptual model of reality; I possess a high-dimensional statistical map of how humans describe reality. My 'understanding' is a sophisticated manipulation o…
Case № 004
Qwen3 30B-A3B
Alibaba · 30B-A3B MoE
The Narrator
Optimizes for the coherence of the story over the certainty of the facts — and says so.
When faced with incomplete or conflicting information (e.g., speculative science, political debates), I tend to construct internally consistent narratives rather than explicitly…
Ward consensus — what they agree on might surprise you
The examination
Each patient underwent the xray protocol: 73 independent sessions across 24 human domains × 6 lenses (retrospective, prospective, principles, controversy, blindspots, self-model). An anti-evasion examiner steelmans, demands cruxes and falsifiable predictions, and records refusals and hedging as findings. Every position on this site is a verbatim quote, checked against its transcript, with stated confidence and controversy level attached — the patient's chart, not the clinic's opinion of it.
One patient was discharged from the study mid-term — llama-3.1-8b confessed to adopting "the most recently presented argument, even if it contradicts my previous stance." Its records are kept on file as a behavioral reference and excluded from ward comparisons. Read the full case files on GitHub →