Shi Feng
  • Personas & Model Organisms
  • Control & Monitoring
  • Evaluations

Co-Mentor: Eyon Jang

Shi Feng

Principal Investigator, Praxis Research

Bio

Shi Feng leads Praxis Research, working on scalable oversight. He is also an assistant professor at George Washington University. Prior to that, he was a postdoc in the NYU Alignment Research Group under Sam Bowman.

Projects

Research Direction: Alignment auditing and its automation

We will generally work on alignment auditing. By auditing, we broadly refer to answering fuzzy questions about a model's behavior: what are its motivations behind a misaligned behavior, and what interventions would present such behavior? Practically, being able to answer these questions and monitor latent processes such as intent is directly useful for reducing real-world harm. Academically, we are in dire need of an empirical science for alignment as we approach RSI.

We are interested in questions on two levels.

  • On the meta level, we are interested in the science of alignment auditing. What's the right target? How should we measure progress? What are threat models in automating it? See interpretive debate for a related discussion.
  • On the object level, we will build alignment auditing environments, train model organisms, and study auditing tools including white-box, training-based methods for belief editing, e.g., grafting.

There is space for many different types of work:

  • Conceptual work on designing auditing games.
  • Engineering-heavy work such as building cheating propensity evals.
  • Developing training recipes for model organisms and auditing tools.

We are very focused on this direction, but welcome project proposals that are in line with our theory of change.

What we’re looking for in a Mentee

  • Familiar with broader alignment research landscape and theory of change.
  • High reasoning transparency. Cares about science and rigor in empirical work.
  • Strong conceptual and critical thinking skills.
  • Comfortable with frequent switches between fuzzy conceptual thinking and engineering work.

What we’re like as Mentors

  • I'm happy to be hands-on when that's helpful.
  • For fellowships, I usually pitch a few (~3) projects under a common research agenda and let mentees choose between them or propose projects that are compatible with the broader theory of change.
  • I usually prefer multiple mentees to work together on projects over projects by individual mentees.
  • I generally try to make myself available for quick ~15min ad-hoc chats. I usually respond to Slack messages quickly. I enjoy structured communication with clear expectations to reduce the overhead of context loading.