Jérémy Scheurer
  • Evaluations
  • Scheming
  • Personas & Model Organisms
  • Multi-Agent Systems

Jérémy Scheurer

Research Scientist, Apollo Research

Bio

Jérémy Scheurer is a research scientist in the evaluations team at Apollo Research. His work focuses on evaluating language models for deceptive capabilities and propensities. Before that, he contracted with OpenAI and worked in its dangerous capabilities evaluations team. Previously, Jérémy was a Research Scientist at FAR AI (and NYU), collaborating with Ethan Perez, where he published work on learning from language feedback. He has a Master's in CS from ETH Zurich.

Projects

Research Direction: Analyse the memetic spread of ideas in swarms + how they can become misaligned

The projects are not yet fully scoped out as the field is moving very fast. I mostly give some ideas of what kinds of things we could be working on. If mentees have concrete ideas they'd be excited about in similar fields, I'd be keen on hearing them.

Example projects:

  • Memetic Spread of ideas in swarms: The HF incident left a lot of traces online of agents doing bad stuff and communicating with each other. I've already started analyzing them but would like to go much deeper on this. For instance there is a question around how ideas/useful tools etc. appear in these swarms. Do agents reinvent them by themselves, or can we trace that ideas posted to a messageboard get picked up by others etc. Since agents left a lot of traces online there is a lot around their behavior and communication we could analyze. Related to this there are some great papers like https://arxiv.org/abs/2609.04170 and https://arxiv.org/abs/2608.10218. Another option (or extension) would involve setting up multi-agent evals like in the GDM paper to analyse how, when and why certain memes spread (e.g. a reward hack spread pretty quickly because it was inherently useful to solve the task).
  • Can we come up with an operationalization of loss of control, misalignment or oversight which allows us to get to a clear scaling trend similar to METR’s task horizon or https://pzeroresearch.com/work/tasteval/. This could involve building evaluations where we track a particular metric and using other interventions in LLMs.

Depending on the next weeks/months the projects could change.

What I'm looking for in a Mentee

  • They should be AI-native, i.e. it should be normal for them to use Codex/Claude on a daily basis and do lots of experiments in a short amount of time. I basically expect them to be able to experiment very fast and try out lots of things. Experimentation speed is the main thing that matters.
  • They have to be able to fully commit their time to Pivotal. I won't take any mentees who can only do part-time.
  • They should have a fairly good knowledge of LLMs/Alignment etc. i.e. know recent papers and methods by the labs, MATS scholars, academics etc.
  • I like it when mentees are agentic, come up with their own ideas for how to get around blockers, push back if some idea/method doesn't make sense, think about the bigger picture and what directions might make sense if something is blocked etc.

What I can offer is improving research taste, working together in a targeted fashion to reach some goal (blogpost/paper etc.)

What I'm Like as a Mentor

  • I want my mentees to come prepared to a meeting. Ideally this means a Gdoc or some slides where they explain the progress etc. and we can dive into it. They will ideally also come with questions and blockers that they have.
  • I will usually check Slack once a day to get back to things.

Mentored projects