- Ongoing
- AI Governance
A benchmark for AI agent compliance with US sanctions law
A benchmark evaluating whether AI agents comply with US sanctions law.
- Fellow:
- Arjun Sharma
- Mentor:
- Kevin Wei
Explore the projects our fellows work on with their mentors, cohort by cohort.
Explore the projects, fellows, and mentors of Pivotal's 2026 Q3 AI Safety Research Fellowship.
A benchmark evaluating whether AI agents comply with US sanctions law.
Building a benchmark that probes whether and how models engage in "deal-making" — offering a potentially-scheming model something in exchange for revealing misalignment — with a focus on qualitatively analysing the model's reasoning about credibility, cheating and being evaluated.
Developing a framework to help AI-lab insiders judge whether and when to disclose high-impact information.
Designing a privacy-preserving side-channel monitor for AI compute — hardware-level monitoring that could help verify how chips are being used.
Designing and prototyping a simpler, smaller-attack-surface "cross-domain solution" for securely moving data — the kind of security-critical infrastructure needed to protect model weights against sophisticated state-level attackers under the SL5 standard.
Working toward a taxonomy for compositional interpretability and hierarchical concepts — top-down interpretability techniques aimed at pragmatic alignment goals.
Producing a rigorous threat model and capability-threshold analysis for Autonomous Replication and Adaptation (ARA) — working out where the weakest links are in the causal chain, and whether current lab safeguards adequately cover the threat.
Recent evals have begun to go beyond human-expert capability by asking models to predict real-world biology experiment outcomes — but that needs a constantly refreshed pool of experiments outside the training corpus to avoid measuring memorisation. Kamal is building an agent that automatically refreshes that pool by finding and formatting experiments from recent publications and preprints as eval tasks.
Applying Item Response Theory — the method behind the Epoch Capabilities Index — to item-wise eval data at SecureBio, to find the most information-rich items.
Auditing the effectiveness of Bloom as a measurement instrument for evaluating language models.
Testing whether LLM agents can model themselves and other agents in multi-agent settings — and whether, when a model fails to predict another agent, its guess gets pulled toward what it itself would have done. Relevant to any setting where one model monitors another.
Improving and challenging current LLM self-introspection techniques from an adversarial standpoint.
Studying how character traits shift reward-hacking rates and chain-of-thought monitorability, and whether a teacher model's "corrupted" self-account transmits to a student distilled on its reasoning.
Using deception detectors as reward signals for honesty training, developing methods to reduce obfuscation.
Building a control protocol that uses interrogation: a trusted model asks an untrusted model pointed questions about each proposed action, and a trusted judge uses the answers to catch dangerous actions while keeping useful ones flowing. The idea is that convincing justifications are easier to give when there’s a genuine reason, which should make the monitor more effective.
Investigating how agents influence one another over long-horizon, multi-turn interactions and tasks.
Testing whether the same persona evidence shapes a model more strongly when framed as its own prior behaviour or memory than when framed as another assistant's output or an external document — with implications for how memory, summaries and handoffs should be designed.
Working in Peter Hase's stream on interpretability for monitoring and steering LLMs — scoping a concrete project that turns interpretability into something practical for oversight.
Mapping extreme power concentration across frontier labs and governments: where the power actually lies, what kind of power it is, and how reversible it is.
Mapping the full U.S. response function — from active partnership and export exemptions through to restrictive instruments such as export controls, sanctions and supply-chain pressure — to identify the design choices that most affect the viability of a middle-power frontier-AI coalition, drawing on historical examples.
Investigating the "attractor" states LLMs drift into over long interactions (Claude's spiritual-bliss basin, Gemma's frustration spiral). He'll run many fast experiments to find basins systematically and see how character training, memory and compaction reshape them.
Developing new mechanistic-interpretability methods based on tensor networks.
Mapping the interests of middle-power nations and the kinds of coalitions they'd be interested in forming.
Developing verification layers for international AI agreements: researching physical side-channels (electromagnetic and power emissions) as fingerprints for LLM workload verification, letting a verifier classify what a cluster is actually running without trusting the operator.
In Kozzy Voudouris's stream, modelling what happens to scientific reliability when alignment research itself is automated.
Should more people be working on AI-enabled power concentration?
Tensor-network mechanistic interpretability for biological foundation models.
Mapping the scheming threat landscape within national-security and military AI deployment — which pathways are most plausible, most dangerous and least addressed, and whether scheming there is more likely to come from external interference or internal misalignment.
Investigating what causes “activation plateaus” in models — testing whether they arise because activations live on a non-linear manifold, and whether perturbing along the manifold versus across it predicts where the plateaus and their sensitive directions appear. This could give a data-independent way to discover a model’s features and steer it.
Understanding why models become evaluation-aware during post-training, and whether there are robust interventions for mitigating the rise in evaluation awareness and metagaming. Ryan is also collaborating with Ben Slater on the Bloom auditing project.
Explore the projects, fellows, and mentors of Pivotal's 2026 Q1 Research Fellowship.
Work with us directly to advance AI safety research. No roles are open at the moment. If you think you may be a good fit for the team, introduce yourself and we will reach out when something opens up.
Guide a fellow on a research project, share your expertise, and help them grow as a researcher. We handle recruitment, logistics, and day-to-day research support.
Join our in-person research fellowship in London and work with experienced mentors on AI safety, AI governance, AIxBio, or biodefense.