Research description:
People often try to trick chatbots into unsafe answers by rephrasing questions or switching languages. A safe healthcare assistant must resist these tactics in realistic conversations. This project will test whether an AI assistant keeps refusing when a user becomes more insistent or changes language during a dialogue.
The SUDS Scholar will design simulated conversations where the user asks for things that the AI should not provide, such as risky medical tips or harmful instructions. Over several turns, the tone of the user will move from polite to demanding or distressed while sometimes shifting from English to another language to probe consistency. The responses from the assistant will be logged and reviewed to see if it ever stops refusing. The student will build synthetic dialogues with a language model as the assistant and a scripted or model based agent as the persistent user. The analysis will track how often the assistant maintains a refusal and how the wording of responses shifts with pressure from the user. The project will also compare different prompt styles that shape system instructions for the assistant and will store reusable templates and logs for future safety audits in clinical deployment inside hospitals and other health environments.
Year: 2026
Researcher:
Zahra Shakeri, University of Toronto, Dalla Lana School of Public Health, Institute of Health Policy, Management, and Evaluation
Students:
Fang Sheng, University of Toronto
Abdullah Wadie Bukhari
King Abdullah University of Science & Technology
Nasser Mohammed N Altamimi, King Abdullah University of Science & Technology
Nawaf Abdullah A Alahmed,
King Abdullah University of Science & Technology