Research description:
AI debate has been proposed as an adversarial scalable oversight method, with encouraging recent progress (see refs below). Debate elicits a wide range of capabilities, however, in particular a mix of knowledge and persuasion. In this pilot project, a new debate protocol focused on disentangling persuasive tendencies from knowledge elicitation will be implemented, validated and explored. Additionally supported by OpenAI funds, this research theme broadly aims to develop scalable oversight methods for super-alignment, using physics as a ground truth. The objective of super-alignment is to ensure that AI systems remain aligned with human values and intentions, even in the limit where they become more capable than humans. Reference document.
Year: 2025
Researcher:
Kristen Menou, Department of Physical and Environmental Sciences, University of Toronto Scarborough
Student:
Xucheng (Rolland) He, University of Toronto