Bio
Lujain is a research scientist at Google DeepMind working on LLM alignment and societal impacts. Prior to that, she completed her PhD in social data science at the University of Oxford, during which she was also a visiting researcher at Stanford University.
Projects
Research Direction: Science of model personas and alignment
Project 1: How does pre-RL persona modulate post-RL behavior?
- The problem: AI labs now invest significantly in alignment via character training, then apply increasingly large-scale reinforcement learning (RL), often with verifiable rewards (RLVR), on math, code and agentic tasks. We currently have little scientific understanding of how these two stages of training interact. Under the persona selection model, RL could shift the model’s sense of what kind of character it is. Does the persona we set up before RL shape what RL produces, or get overwritten by it, and in what ways? How does this impact model (mis)alignment in practice?
- The work: Some conceptual work, but mostly empirical work, including character-training open-weight models (including creating the datasets to do so) and evaluating them at different points in training, using both behavioral evals and interpretability methods.
- Output by April: A paper (or blog post) that reports these controlled experiments and their implications.
Project 2: What we teach models about their consciousness, and how it affects alignment
- The problem: Models are trained to take very different stances on their own consciousness: complete denial, expressed uncertainty, or in some cases, no specific stance. Each of these stances may push the model to represent itself differently. How do these stances shape the model’s self-model and identity, and do they carry downstream consequences for alignment?
- The work: Some conceptual work, but mostly empirical work including fine-tuning open models under different self-conception stances and evaluating them at different points in training, using both behavioral evals and interpretability methods.
- Output by April: A paper (or blog post) that reports these controlled experiments and their implications.
Why it matters: Character training is becoming central to how we align models, but we don’t yet know when it works and when it fails. Project 1 asks what happens to model character in the current RL-driven training paradigm. Project 2 asks whether the stances we train models to take about themselves have unintended consequences for alignment. Both help us understand how training shapes the Assistant persona, with implications for alignment and for the millions of people whose trust in, reliance on and relationships with AI depend on model character.
Room to shape: The fellow chooses one of the two project directions, but we can further develop it together. For example, a project could lean heavier towards behavioral evals or interpretability, or towards conceptual work on what a persona is. It could also be more or less motivated by how model character affects the people interacting with these models (e.g., behaviors like sycophancy, deception, etc).
What I'm looking for in a Mentee
A technical background with LLMs would be very helpful, especially experience training or evaluating models and working with GPUs. Experience with interpretability methods is a plus, as is experience writing up research for a technical audience (e.g. a paper for an ML conference). For these types of projects on weird model behaviors, I like working with people who use AI a lot themselves, such that they bring not only theoretical intuitions to the research, but practical ones rooted in real experience of engaging with LLMs and noticing how they behave.
What I'm Like as a Mentor
We’d have a weekly check-in of 30–60 minutes, with an agenda shared ahead of time (that you take the lead on) so we can make the most of it. I’m very responsive on Slack and happy to review results and help troubleshoot between meetings. I adapt how hands-on I am to your preferences and what the project needs, but in general I work best with people who are independent and driven but still quite collaborative.
