Projects

Explore the projects our fellows work on with their mentors, cohort by cohort.

Q3 2026

AI Safety Fellowship

Explore the projects, fellows, and mentors of Pivotal's 2026 Q3 AI Safety Research Fellowship.

  • Ongoing
  • AI Governance

A benchmark for AI agent compliance with US sanctions law

A benchmark evaluating whether AI agents comply with US sanctions law.

Fellow:
Arjun Sharma
Mentor:
Kevin Wei
  • Ongoing
  • Technical AI Safety

A benchmark for “deal-making” with potentially scheming models

Building a benchmark that probes whether and how models engage in "deal-making" — offering a potentially-scheming model something in exchange for revealing misalignment — with a focus on qualitatively analysing the model's reasoning about credibility, cheating and being evaluated.

Fellow:
Mark Keavney
  • Ongoing
  • AI Governance

A disclosure framework for AI-lab insiders

Developing a framework to help AI-lab insiders judge whether and when to disclose high-impact information.

  • Ongoing
  • Technical AI Safety

A simpler “cross-domain solution” for protecting model weights

Designing and prototyping a simpler, smaller-attack-surface "cross-domain solution" for securely moving data — the kind of security-critical infrastructure needed to protect model weights against sophisticated state-level attackers under the SL5 standard.

Mentor:
Guy
  • Ongoing
  • Technical AI Safety

A taxonomy for compositional interpretability and hierarchical concepts

Working toward a taxonomy for compositional interpretability and hierarchical concepts — top-down interpretability techniques aimed at pragmatic alignment goals.

Fellow:
Ethan Nguyen
  • Ongoing
  • AI Governance

A threat model for Autonomous Replication and Adaptation (ARA)

Producing a rigorous threat model and capability-threshold analysis for Autonomous Replication and Adaptation (ARA) — working out where the weakest links are in the causal chain, and whether current lab safeguards adequately cover the threat.

Mentor:
Jide Alaga
  • Ongoing
  • Biodefense

An agent that refreshes biology-experiment evals from recent publications

Recent evals have begun to go beyond human-expert capability by asking models to predict real-world biology experiment outcomes — but that needs a constantly refreshed pool of experiments outside the training corpus to avoid measuring memorisation. Kamal is building an agent that automatically refreshes that pool by finding and formatting experiments from recent publications and preprints as eval tasks.

Fellow:
Kamal Maher
  • Ongoing
  • Biodefense

Applying Item Response Theory to item-wise eval data

Applying Item Response Theory — the method behind the Epoch Capabilities Index — to item-wise eval data at SecureBio, to find the most information-rich items.

Fellow:
Ben Pomeranz
  • Ongoing
  • Technical AI Safety

Auditing Bloom as a measurement instrument for evaluating language models

Auditing the effectiveness of Bloom as a measurement instrument for evaluating language models.

Fellow:
Ben Slater
Mentor:
Kevin Wei
  • Ongoing
  • Technical AI Safety

Can LLM agents model themselves and other agents?

Testing whether LLM agents can model themselves and other agents in multi-agent settings — and whether, when a model fails to predict another agent, its guess gets pulled toward what it itself would have done. Relevant to any setting where one model monitors another.

Fellow:
Luis Montoya
  • Ongoing
  • Technical AI Safety

Challenging LLM self-introspection techniques from an adversarial standpoint

Improving and challenging current LLM self-introspection techniques from an adversarial standpoint.

Fellow:
Guy Galun
Mentor:
Peter Hase
  • Ongoing
  • Technical AI Safety

Character traits, reward hacking, and chain-of-thought monitorability

Studying how character traits shift reward-hacking rates and chain-of-thought monitorability, and whether a teacher model's "corrupted" self-account transmits to a student distilled on its reasoning.

Fellow:
Ionuț Stan
  • Ongoing
  • Technical AI Safety

Deception detectors as reward signals for honesty training

Using deception detectors as reward signals for honesty training, developing methods to reduce obfuscation.

Fellow:
Lily Shi
Mentor:
Peter Hase
  • Ongoing
  • Technical AI Safety

Extending AI control protocols with interrogation

Building a control protocol that uses interrogation: a trusted model asks an untrusted model pointed questions about each proposed action, and a trusted judge uses the answers to catch dangerous actions while keeping useful ones flowing. The idea is that convincing justifications are easier to give when there’s a genuine reason, which should make the monitor more effective.

Mentor:
Adam Kaufman
  • Ongoing
  • Technical AI Safety

How agents influence one another over long-horizon interactions

Investigating how agents influence one another over long-horizon, multi-turn interactions and tasks.

Fellow:
Ben Maltbie
  • Ongoing
  • Technical AI Safety

How framing persona evidence as memory shapes model behaviour

Testing whether the same persona evidence shapes a model more strongly when framed as its own prior behaviour or memory than when framed as another assistant's output or an external document — with implications for how memory, summaries and handoffs should be designed.

  • Ongoing
  • Technical AI Safety

Interpretability for monitoring and steering LLMs

Working in Peter Hase's stream on interpretability for monitoring and steering LLMs — scoping a concrete project that turns interpretability into something practical for oversight.

Fellow:
Thea Xu
Mentor:
Peter Hase
  • Ongoing
  • AI Governance

Mapping extreme power concentration across frontier labs and governments

Mapping extreme power concentration across frontier labs and governments: where the power actually lies, what kind of power it is, and how reversible it is.

Fellow:
Haimi Tefera
  • Ongoing
  • AI Governance

Mapping the U.S. response function for a middle-power AI coalition

Mapping the full U.S. response function — from active partnership and export exemptions through to restrictive instruments such as export controls, sanctions and supply-chain pressure — to identify the design choices that most affect the viability of a middle-power frontier-AI coalition, drawing on historical examples.

  • Ongoing
  • Technical AI Safety

Mapping the “attractor” states LLMs drift into over long interactions

Investigating the "attractor" states LLMs drift into over long interactions (Claude's spiritual-bliss basin, Gemma's frustration spiral). He'll run many fast experiments to find basins systematically and see how character training, memory and compaction reshape them.

Fellow:
Tim Farrelly
  • Ongoing
  • AI Governance

Middle-power interests and prospective AI coalitions

Mapping the interests of middle-power nations and the kinds of coalitions they'd be interested in forming.

  • AI Governance

Physical side-channels as verification layers for international AI agreements

Developing verification layers for international AI agreements: researching physical side-channels (electromagnetic and power emissions) as fingerprints for LLM workload verification, letting a verifier classify what a cluster is actually running without trusting the operator.

Mentor:
Gabriel Kulp
  • Ongoing
  • Technical AI Safety

Scientific reliability when alignment research is automated

In Kozzy Voudouris's stream, modelling what happens to scientific reliability when alignment research itself is automated.

  • Ongoing
  • AI Governance

Should more people be working on AI-enabled power concentration?

Should more people be working on AI-enabled power concentration?

Fellow:
Hugo Bos
  • Ongoing
  • Technical AI Safety

Tensor-network mechanistic interpretability for biological foundation models

Tensor-network mechanistic interpretability for biological foundation models.

  • Ongoing
  • AI Governance

The scheming threat landscape in national-security AI deployment

Mapping the scheming threat landscape within national-security and military AI deployment — which pathways are most plausible, most dangerous and least addressed, and whether scheming there is more likely to come from external interference or internal misalignment.

Fellow:
Liha Leung
Mentor:
Jide Alaga
  • Ongoing
  • Technical AI Safety

What causes activation plateaus in models?

Investigating what causes “activation plateaus” in models — testing whether they arise because activations live on a non-linear manifold, and whether perturbing along the manifold versus across it predicts where the plateaus and their sensitive directions appear. This could give a data-independent way to discover a model’s features and steer it.

  • Ongoing
  • Technical AI Safety

Why models become evaluation-aware during post-training

Understanding why models become evaluation-aware during post-training, and whether there are robust interventions for mitigating the rise in evaluation awareness and metagaming. Ryan is also collaborating with Ben Slater on the Bloom auditing project.

Mentor:
Kevin Wei

Q1 2026

AI Safety Fellowship

Explore the projects, fellows, and mentors of Pivotal's 2026 Q1 Research Fellowship.

  • AI Governance

Analysing the Shanghai Cooperation Organisation’s (SCO) evolution from “information security” to “AI governance” frameworks

Fellow:
Otto Barrow
  • Technical AI Safety

Can LLMs develop goal misgeneralization without reward misspecification?

Fellow:
Jinzhou Wu
  • Technical AI Safety

Improving human-AI coordination in collective intelligence and decision-making

2024

AI Safety Fellowship

Developing frameworks for red teaming DNA synthesis screening

Fellow:
Abel Ashby

Guarding Democracies with Evaluations of Political Persuasion

Modeling pre-superhuman persuasion risks from agentic systems

Predicting models capabilities through observational scaling laws

Fellow:
Jeanne Salle

Runaway Military AI: Governance Initiatives for AI in the Military Domain

The Multi-agent Dynamics of Elliott Thornley's TD-agents

Fellow:
Joss Oliver

Third-party auditing of frontier models: case studies from non-AI industries

Transferability of adversarial attacks between modalities of vision-language models

Fellow:
Jakub Krys

Understanding apocalyptic bioterrorism and its effects on rational deterrence theories

Fellow:
Felix Porée

Which confidence-building measures are most politically viable for US-China Track II AI dialogues?

2023

AI Safety Fellowship

Beyond the United States: Potential Disruptive Actors in International AI Governance

China's Compute Governance Potential: Advancing Cooperative AI Governance

Fellow:
Yueyang Cao

Current AI Policy Structures in China

Fellow:
Bill Chen

Developing Interpretability Techniques by Reverse Engineering Toy Models

Direct-Ascent Anti-Satellite Missiles: Moratorium, Resolution, and Lessons for the Future

Large Language Models as Active Inference models

Fellow:
Aude Maier

Red Teaming Mainstream AI Safety Regulation Proposals

Uncorrelated Failure Modes: The Swiss Cheese Approach to AI Safety

2022

AI Safety Fellowship

Cooperative AI Leveraging Group Identity and Team Reasoning to Frame Cooperative Equilibria

Data for IRL: What do we need to learn human values?

Fellow:
Jan Wehner

How New Biotechnologies Can Strengthen the BWC

Fellow:
Theo Knopfer

Improving Institutional Decision-Making: Which Improvements?

Incentive Compatibility Dilemma in Multi-Agent Inverse Reinforcement Learning

On the normative intuitions of Xi Jinping's political thought and what it implies for the future of humanity.

Optimal GCR mitigation through different decision frameworks.

Proliferation and Catastrophic Risks from Generation IV Nuclear Reactors / SMRs

Socio-psychological Variables that Impact Decision-Making in Food-scarce Post-disaster Scenarios

Fellow:
Isla Gibson

The Central Challenge to Global Catastrophic Biological Risks: A Reconsideration of Epidemic Sovereignty

Training Language Models with Natural Language Feedback

What are some ideal structures of AGI governance?

2021

AI Safety Fellowship

AI Surveillance: Comparing Chinese & EU Regulations

How can economists best contribute to pandemic prevention and preparedness? / Catastrophic rectangles - visualising catastrophic risks

Integrating the Interests and Needs of Future Generations into the Swiss Political System

Moral Circle Expansion: A Way to Mitigate Global Catastrophic Risks?

What Lawyers can contribute to Longtermism and Existential Risk Research

Get involved

Join our team

Work with us directly to advance AI safety research. No roles are open at the moment. If you think you may be a good fit for the team, introduce yourself and we will reach out when something opens up.

Mentor top emerging researchers

Guide a fellow on a research project, share your expertise, and help them grow as a researcher. We handle recruitment, logistics, and day-to-day research support.

Participate in a fellowship

Join our in-person research fellowship in London and work with experienced mentors on AI safety, AI governance, AIxBio, or biodefense.