Bio
Jan Kirchner is an AI alignment researcher at Anthropic in San Francisco, previously at OpenAI. He holds a PhD in computational neuroscience and writes the Universal Prior Substack on decision theory, economics, and cognitive science. Originally from Germany, he now calls California home.
Projects
Research Direction: Using frontier AI to accelerate alignment research
AI systems are now capable enough to do real research. Can we point them at alignment itself? Here are some ideas that I would be excited to pursue with the right fellow.
- Mapping alignment research, four years on. In 2022 I co-authored a dataset and unsupervised analysis of the alignment literature. The field has grown and fragmented enormously since. Rebuild the dataset, use language models to cluster and summarize it, and track how subfields and research communities have shifted. Output: an updated public dataset and paper.
- Automated alignment researchers. Recent Anthropic work showed Claude agents autonomously finding methods that close most of the gap on weak-to-strong supervision and on benchmarks for ten alignment failures. Open questions remain: do these methods hold up on unseen tasks and larger models, can agents make progress on fuzzier problems without a clean score, and how do we reliably catch them gaming the setup? Output: experiments extending the open-source harness, written up as a paper.
- Agent foundations with AI mathematicians. Use frontier models' rapidly improving math abilities to attack open problems in agent foundations, such as decision theory, logical uncertainty, or embedded agency. Output: new formal results, or a careful account of where AI-assisted proof helps and where it doesn't.
Day to day, you'll design experiments, run language model agents, write code, and scrutinize results. By the end of April, a strong result is a public preprint or Alignment Forum post with code or data released.
Why it matters: if AI can do alignment research, safety work can keep pace with capabilities. If it can't, we need to know where it fails before we rely on it.
What I'm looking for in a Mentee
I enjoy working with someone driven and self-directed who already has real empirical research experience: you can design experiments, write the code, and interpret messy results without hand-holding. Career stage matters less to me than ambition and a track record of getting things done.
What I'm Like as a Mentor
I match the energy the fellow brings: if you're pushing hard, I'll meet weekly, dig into your results, and share my takes and help unblock you in any way I can. If momentum stalls, I'll step back and let the project run its course. I communicate mostly asynchronously over Slack and appreciate concise, regular written updates.
