Bio
Edward is a cofounder and member of technical staff at Geodesic Research.
He holds a PhD in Computational Neuroscience from the University of Cambridge, and worked previously as a research engineer at the UK AISI Misuse Red Team.
Edward is interested in building up a science of misalignment: why does misalignment occur, what does misaligned cognition look like, and how can we give mechanistic accounts of how RL leads to misalignment?
Projects
Research Direction: Inducing and understanding pathological cognition from RL
Fellows will work on creating RL environments which elicit pathological cognition in models, and then understanding the psychology of those models.
Types of cognition we are interested in inducing:
- Metagaming. Reasoning about oversight and grading processes, and how to evade or deceive them.
- Collusion. Cooperating in an unsanctioned way with other instances.
- Paranoia. Instinctive distrust of stated facts, situations, or tool outputs.
- Situational awareness and self-preservation. Reasoning about the implications of being an AI, particularly with respect to corrigibility and future training.
Work on this project primarily looks like:
- Developing diverse RL environments to induce a pathology
- Analysing transcripts from models in those environments
- Debugging and iterating on large-scale RL runs
- (Optionally) Investigating the cognition of trained models, using tools from model forensics, interpretability, and evals
- (Optionally) Analysing the dynamics of training using frameworks such as the behavioural selection model
- (Optionally) Performing Post-Training Diffing to isolate which environment elements are responsible for the induced cognition
Successful outputs would be OS model checkpoints and run transcripts for model organisms induced by RL. Excellent execution would accompany this with an analysis of the dynamics of how the pathology is learned or the final model's cognition.
Building these model organisms gives us a way to study failures we've seen in recent models and expect to see more of in future. In building these stacks we learn more about the nature of that pathology and where it stems from within RL. Finally, developing tools to better understand the cognition of models expands our basic science of model cognition.
Fellows have scope to choose which pathology they are interested in and from what angle they are most excited to analyse the runs (see above). The project will not focus on mitigations or solutions for the pathologies; this is out of scope.
What we’re looking for in a Mentee
I'm interested in mentoring fellows with a scientific background and mindset.
I value curiosity and a desire to generate comprehensive explanations, and the ability to quickly generate, test, and discard hypotheses.
Experience with RL on LLMs is a plus, but not required!
What we’re like as Mentors
I'd like to meet with fellows once or twice a week for an hour to discuss progress.
At meetings, I'd like fellows to present slides with their recent progress, and a few directions they are most excited to pursue next. I will give my own suggestions, and an overall gauge of how excited I am about each of them, but I'm happy to give fellows leeway in exploring their own ideas and interests (provided it stays within the bounds of the overall direction).
Between meetings, I like fellows to post research updates or blockers frequently on Slack, and will be responsive to requests for immediate feedback between meetings or technical debugging.

