Research
A selection of research conducted by pivotal fellows.
A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
Shikhar Shiromani
Mentor: Noah Y. Siegel (Google DeepMind)
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
Arnau Marin-Llobet
Mentor: Stefan Heimersheim (FAR.AI)
ICML 2026: Mechanistic Interpretability Workshop
Evidence for Feature-Specific Error Correction in LLMs
Francisco Ferreira da Silva
Mentor: Stefan Heimersheim (FAR.AI)
Training on Documents About Monitoring Leads to CoT Obfuscation
Reilly Haskins
Mentors: Joshua Engels (Google DeepMind) & Bilal Chughtai (Google DeepMind)
Causal Foundations of Collective Agency
Frederik Hytting Jørgensen
Mentor: Lewis Hammond (Cooperative AI Foundation)
Proceedings of CLeaR 2026 (PMLR 323)
Bayesian Influence Functions for Hessian-Free Data Attribution
Philipp Alexander Kreer
Mentor: Jesse Hoogland (Timaeus)
ICLR 2026 Poster
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
Guillaume Corlouer & Avi Semler
Mentors: Alexander Strang (UC Berkeley) & Alexander Gietelink Oldenziel (Timaeus)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
Joachim Schaeffer & Arjun Khandelwal
Mentor: Tyler Tracy (Redwood Research)
ICLR 2026: AIWILD Workshop
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
Aaron Sandoval
Mentor: Cody Rushing (Redwood Research)
Beyond Vibe Decision Theory: Asymmetric Manipulation Vulnerabilities in LLM Multi-Agent Coordination
Sukanya Krishna
Mentor: Tobin South (Stanford HAI)
NeurIPS 2025 Workshop on Algorithmic Collective Action; AAMAS 2026
The Multi-Agent Off-Switch Game
Soroush Ebadian
Mentor: Lewis Hammond (Cooperative AI Foundation)
Trustworthy Agentic AI Workshop@AAAI26
Decomposition of Small Transformer Models
Casper L. Christensen
Mentor: Logan Riggs Smith
NeurIPS 2025: Mechanistic Interpretability Workshop
Influence Dynamics and Stagewise Data Attribution
Jin Hwa Lee & Matt Smith
Mentor: Jesse Hoogland (Timaeus)
ICLR 2026
Activation Probes Are Reliable With a Handful of Positive Examples
Riya Tyagi
Mentor: Stefan Heimersheim (Apollo Research)
NeurIPS 2025: Mechanistic Interpretability Workshop
Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data
Nicola Croce
Mentor: Tobin South (MIT)
Technical AI Governance Forum 2025
Benchmarking Deception Probes via Black-to-White Performance Boosts
Avi Parrack
Mentor: Stefan Heimersheim (Apollo Research)
ICML 2025: Actionable Interpretability Workshop
