Research

A selection of research conducted by Pivotal fellows.

1 Aug 2026

A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

Researcher: Shikhar Shiromani

Mentor: Noah Y. Siegel

Technical AI Safety

6 Jul 2026

Compressed Computation under L⁴ Loss is likely Computation in Superposition

Researcher: Francisco Ferreira da Silva

Mentor: Stefan Heimersheim

Technical AI Safety

3 Jul 2026

Individual Parameters in Weight-Sparse Transformers Appear Interpretable

Researcher: Arnau Marin-Llobet

Mentor: Stefan Heimersheim

Technical AI Safety

24 Jun 2026

Door's Locked, Try the Window: Measuring Circumvention Propensity in Coding Agents

Researcher: Prakrat Agrawal

Mentor: Jérémy Scheurer

Technical AI Safety

23 Jun 2026

Evidence for Feature-Specific Error Correction in LLMs

Researcher: Francisco Ferreira da Silva

Mentor: Stefan Heimersheim

Technical AI Safety

14 May 2026

Training on Documents About Monitoring Leads to CoT Obfuscation

Researcher: Reilly Haskins

Mentors: Joshua Engels & Bilal Chughtai

Technical AI Safety

30 Apr 2026

Causal Foundations of Collective Agency

Researcher: Frederik Hytting Jørgensen

Mentor: Lewis Hammond

Technical AI Safety

23 Apr 2026

Bayesian Influence Functions for Hessian-Free Data Attribution

Researcher: Philipp Alexander Kreer

Mentor: Jesse Hoogland

Technical AI Safety

7 Apr 2026

Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks

Researchers: Guillaume Corlouer & Avi Semler

Mentors: Alexander Strang & Alexander Gietelink Oldenziel

Technical AI Safety

4 Feb 2026

Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring

Researchers: Joachim Schaeffer & Arjun Khandelwal

Mentor: Tyler Tracy

Technical AI Safety

1 Dec 2025

Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI

Researcher: Aaron Sandoval

Mentor: Cody Rushing

Technical AI Safety

1 Dec 2025

Beyond Vibe Decision Theory: Asymmetric Manipulation Vulnerabilities in LLM Multi-Agent Coordination

Researcher: Sukanya Krishna

Mentor: Tobin South

Technical AI Safety

20 Nov 2025

The Multi-Agent Off-Switch Game

Researcher: Soroush Ebadian

Mentor: Lewis Hammond

Technical AI Safety

12 Nov 2025

Decomposition of Small Transformer Models

Researcher: Casper L. Christensen

Mentor: Logan Riggs Smith

Technical AI Safety

14 Oct 2025

Influence Dynamics and Stagewise Data Attribution

Researchers: Jin Hwa Lee & Matt Smith

Mentor: Jesse Hoogland

Technical AI Safety

30 Sep 2025

Activation Probes Are Reliable With a Handful of Positive Examples

Researcher: Riya Tyagi

Mentor: Stefan Heimersheim

Technical AI Safety

26 Jul 2025

Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data

Researcher: Nicola Croce

Mentor: Tobin South

Technical AI Safety

16 Jul 2025

Benchmarking Deception Probes via Black-to-White Performance Boosts

Researcher: Avi Parrack

Mentor: Stefan Heimersheim

Technical AI Safety

26 May 2025

Understanding the learned look-ahead behavior of chess neural networks

Researcher: Diogo Cruz

Technical AI Safety

Get involved

Join our team

Work with us directly to advance AI safety research. No roles are open at the moment. If you think you may be a good fit for the team, introduce yourself and we will reach out when something opens up.

Mentor top emerging researchers

Guide a fellow on a research project, share your expertise, and help them grow as a researcher. We handle recruitment, logistics, and day-to-day research support.

Participate in a fellowship

Join our in-person research fellowship in London and work with experienced mentors on AI safety, AI governance, AIxBio, or biodefense.