Bio
David Manheim is a policy and risk researcher, and founder of the Israeli think-and-do tank, ALTER. He has been working on AI risks and related policy since pre-historical times of 2017/2018, alongside his work on Biorisk and policy. Prior to that, he completed a PhD at the RAND Corporation, was one of the original set of Superforecasters, and has worked for and advised a variety of other organizations focused on building a flourishing and safe future.
Projects
Research Direction: Multi-agent (non-swarm) oversight and risks
The problem: In my view, most current work on multi-agent risks over-indexes on prior conceptual models from multi-agent RL, which involves training models together, or from game theory, or focuses on recent events. In fact, multi-agent interactions between entirely separate systems in unstructured domains may look unlike any of this. Worse, the class of incident may not be clearly understood even in retrospect - for example, events like financial flash-crashes were never fully understood, involving large groups of interacting systems.
The work: Better understanding of this risk, and of mitigations, could involve anything from conceptual foundational work to empirical investigation of model's ability to identify opportunities for cooperation, to policy research on how governments or Labs could identify and mitigate these risks. The exact class of work is very flexible and would depend on the approach; I expect that I or another collaborator would be able to provide types of expertise that the fellow lacks, so the details of the work are flexible.
Output: Presumably, this would be a paper, either presenting some new result or useful approach, or a review of relevant existing work across domains and research agenda, or a policy recommendation. (In some cases, it is possible for safety reasons this would be shared directly with key stakeholders rather than published, but I do not expect that.) The venue would depend on the type of paper.
Why it matters: This could enable better understanding of, identification of risks from, or mitigations to guard against, one plausible route to large-scale failures during or caused by deployment of AI in varied applied domains which we are already beginning to see.
Room to shape the project is very significant, as noted.
What I'm looking for in a Mentee
I am flexible about background, and have and would be happy to do anything from advise undergraduates on their first real research project, to guide graduate students on larger papers, to collaborate with professionals with substantive expertise to help them apply it to policy and research.
What I'm Like as a Mentor
I am happy to be as involved as the fellow prefers, from general advising, to in-depth feedback, to close collaboration, depending on their backgrounds and the type of work. I am interested in co-developing project proposals, and suggesting approaches, but would also be happy to supervise a project of the fellow's own independent design or suggestion. I expect mentees to be self-driven and proactively seek feedback, without needing my encouragement, but am happy to either closely supervise and/or co-work, or have routine check-ins about progress and advice. My technical background is sufficient to supervise more technical AI/ML projects if fellows have relevant backgrounds, but will be less able to provide substantive advice on methods and approaches.
