Bio
I’m a founding member of the technical staff at Geodesic Research, a research nonprofit focused on pre-RL alignment. I lead our pretraining efforts. I am excited to explore: “Since we are concerned about X, is pretraining on X-relevant data making things worse?” Broadly, this direction looks like intervening on pretraining and midtraining data mixes, removing potentially unwanted information we don't want to be salient to AIs, or inserting new data to create desirable personas and influence generalisation. Before Geodesic, I was an ERA fellow, researcher at EleutherAI, and an applied scientist at Microsoft.
Projects
Research Direction: Understanding Misalignment-Relevant Generalisation via Pretraining-Time Interventions
Overview: This stream will study pretraining-time alignment and safety interventions. We want to do work that helps the community better understand pretraining learning dynamics and how they affect downstream alignment. We can then use these advances to motivate alignment interventions. We further describe the theory of change in this post: https://www.lesswrong.com/posts/nhjkHsppEk98xxmPe/why-study-alignment-interventions-on-pre-rl-checkpoints
Workflow: Fellows will spend their time reading papers, delegating work to coding agents, spending 2+ hours a day reviewing outputs from Claude/Codex (e.g., code review of PRs), filtering pretraining datasets, and configuring training runs. Pretraining runs will be conducted on Geodesic’s compute. We expect fellows to take ownership and work well on a team, helping each other with blockers and delegating work. Fellows may contribute to Geodesic’s core infrastructure if they find opportunities for improvement. If fellows do not think the project is going well, they should raise the (potentially vague) concern and should potentially advocate for us to pivot.
Directions: We will work with fellows to refine the specific project leading up to the fellowship. The goal is to have a paper to share with the community or a research note blog post. We want to report negative results if we believe they tell the community something interesting. Possible research directions include the following, with a full list at https://docs.google.com/document/d/10gLGvgpU5yLH42-B_aSlMuwtjfMpQv9s5kAS6eKrmfk/edit?usp=sharing:
RQ — When do midtraining and SDF interventions fail to overwrite alignment priors?
The community has become increasingly aware of crafting alignment priors via midtraining and SDF. However, as models become more capable and natural discussion of AI misalignment becomes more prevalent in pretraining datasets, models may develop increasingly crystallised alignment priors by the end of base model training. Evidence from the machine unlearning literature suggests that overwriting deeply held knowledge via continued training is challenging. Similarly, these midtraining/SDF approaches may break down, or remain effective only when instilling knowledge that does not conflict with pre-existing priors. We aim to study this empirically. If we find positive results for this hypothesis, then this suggests that the community should prioritise pretraining-time interventions from token zero or filter out knowledge which we may wish to overwrite later. If we get negative results, this suggests the community can safely continue investing in late-stage alignment interventions.
RQ — Single-Persona Pretraining
The Persona Selection Model (PSM) (Marks et al., 2026) posits that LLMs represent a diverse distribution of characters/authors/personas during pretraining, and then assume one of these personas during generation. Ideally, we’d have a consistent, robust, helpful, honest, and harmless (HHH) AI assistant persona that is hard for adversaries to remove and can withstand distribution shifts. This direction proposes rewriting all of pretraining (and midtraining) to have a single common persona. In the best case, acting from a single, consistent, pretraining-curated persona is so natural to the model that downstream training that changes the model's character becomes an uphill battle.
What we’re looking for in a Mentee
I am excited to potentially hire fellows for full-time roles with Geodesic after the fellowship. Broadly, we are excited about full-stack researchers: folks who can contribute to the conceptual and strategic motivations for our project’s theory of change, while also having the discipline to grind for weeks on empirics without becoming distracted, including reading data for hours and thinking deeply about a single problem for days. Correlated traits include truth-seeking, responsibility, and agency: we are among the relatively few people in the world who can make technical progress on perhaps the most important problem facing humanity — we need to be maximally truth-seeking and execute as best we can. Fellows’ impact should be amplified by being on a team and can help make others on the team the best versions of themselves.
What we’re like as Mentors
I like to be involved and hold my fellows to high standards for the quality of their work, the reliability of their planning, and the clarity of their thinking. I'm happy to find an optimal balance between being hands-on with fellows and my own bandwidth, with a preference towards frequent short standups. I prefer to communicate over calls, but am generally available over Slack with a couple-hour delay. I tend to cancel meetings if fellows do not arrive prepared, which I think sets up fellows for future collaborations with senior researchers and full-time roles.

