Living papers: AI safety
Influential AI safety and interpretability papers, kept alive. For each one, our agents replicated the original analysis from its own materials, re-ran the main claims on model generations the authors never saw, and refit the theory to the combined evidence. Every page shows the paper's own results next to ours, and each has a chatbot (sign in to use it) that answers questions from the report. The archive updates as new models come out. The same treatment for economics lives at sai.science/econ.
The Linear Representation Hypothesis
Park, Choe and Veitch · ICML 2024 · models through 2025
Replicates cleanly and survives to 2025: concepts stay linear and causally orthogonal on every model, but the paper's 1.5-2.5x advantage holds only in the LLaMA lineage, with Qwen and OLMo near 1.25x.
The Geometry of Truth
Marks and Tegmark · arXiv 2023 · models through 2025
The truth direction survives across all three 2024-2025 model families, though its depth is family-specific and the negation gap has nearly closed.
Interpretability in the Wild: the IOI Circuit
Wang, Variengien, Conmy, Shlegeris and Steinhardt · ICLR 2023 · models through 2025
The behaviour holds up, the Pythia family never learned it, and the tidy 26-head circuit turns out to be a GPT-2 small story: only its name-mover backbone persists at scale.
Function Vectors in Large Language Models
Todd, Li, Sharma, Mueller, Wallace and Bau · ICLR 2024 · models through 2025
Function vectors survive to 2025 and even strengthen, while the where-to-insert and ten-heads rules turn out to belong to the older attention design.
Persona Vectors
Chen et al. · arXiv 2025 · models through 2025
Survives to 2025 with one amendment: monitoring reliability is set by alignment style, and reasoning-style training breaks the monitor.