Bio
Cecilia Tilli leads Allelo, a new AI safety research team studying which propensities and character traits make AI agents safe to deploy in multi-agent settings. Previously she spent three years at the Cooperative AI Foundation, running its grantmaking program, leading workshops, and building evaluations of agent-agent influence capabilities. Before that she founded a vegan cheese company, worked in strategy consulting for research-heavy startups, led a foundation working to prevent antibiotic resistance, and did postgraduate research in scientific computing at Uppsala University.
Projects
Research Direction: AI agents' susceptibility to influence and related safety propensities
The problem. AI agents will increasingly influence, and be influenced by, other agents and humans. Developers will have to decide how persuadable to make their agents, but it's unclear what level of susceptibility to influence is desirable: an easily persuaded agent can be manipulated, while training an agent to resist influence could have side effects, such as making it less corrigible. At scale, getting this wrong could enable manipulation and failures that spread across many agents. I want to understand how susceptibility to influence relates to other safety-relevant propensities, and how training changes it.
Project options (we'll choose the direction together based on your interests and fit):
- Model organisms. Develop open-weight models with different levels of susceptibility to influence, comparing approaches such as different fine-tuning methods and activation steering. Strong result: a set of model organisms with targeted, well-understood changes, and a comparison of which methods work best.
- Propensity evaluations. Build evals that complement our basic susceptibility-to-influence evals, covering other aspects of influence or other safety-relevant propensities. Strong result: open-source evals others can use.
- Influence in agent populations. Study how agents with different levels of susceptibility behave in groups, possibly using real-world deployment data. More observational. Strong result: an analysis of how influence spreads and which agents drive or resist it.
By January, I expect to have minimal model organisms and basic susceptibility evals ready for you to build on. I'm also open to mechanistic interpretability angles for fellows with that background.
What I'm looking for in a Mentee
I'm looking for fellows with hands-on experience working with LLMs, such as running evals or fine-tuning models; the specific tools matter less. You should be able to work independently and write clearly. I'd especially like to work with people interested in research careers outside academia.
What I'm Like as a Mentor
I meet fellows 1:1 at least weekly, include them in Allelo's team meetings, and review documents async in between. I expect fellows to drive their project, and I focus on giving substantive feedback on direction, design and writeups. I prefer Slack, give direct feedback, and want the same back. I value clear plans and kept commitments, and I'm quick to drop ideas that aren't working.
