SUDS project

Developing Datasets and Algorithms for Metabolic Modelling and Engineering

Research description:

Current methods for developing biocatalysts for sustainable applications require multiple rounds of design and experimentation with significant trial and error steps. Bioengineering can be accelerated significantly if systematic design, testing can be used to narrow the range of candidates that require testing. While there are many deterministic methods for modeling metabolic networks they fall short in their ability to predict the cumulative effect of higher order modifications typically required to enhance the production of desired compounds relevant for applications, for example, adipic acid required for bionylon synthesis. Hence, there is a need to develop hybrid methods that combine deterministic methods with data driven methods that can account for incomplete biological knowledge. Here we aim to use multi-modal large language models including DNA, and protein language models that can be used to predict and correct the current shortcomings of deterministic models. We aim to build on several in house datasets including a pandemic dataset of all metabolic reactions and augment this dataset with curated models derived from gold standard databases such as SWISS PROT and BiGG. We will then use these high quality metabolic network datasets to develop hybrid methods that can predict the impact of genomic modifications on physiology.

The SUDS Scholar will work to develop pipelines to process and curate data from different biological domains specific for microbial metabolism. The scholar will then work with other group members to apply existing pipelines for data analysis and modeling and also facilitate the development of new data analysis methods.

Year: 2026

Researcher: Radhakrishnan Mahadevan, University of Toronto, Faculty of Applied Science and Engineering, Department of Chemical Engineering and Applied Chemistry

Student:
Hussain Faisal Zaid Alharbi
King Abdullah University of Science & Technology

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Refusal Robustness of Clinical AI Assistants Under User Pressure and Language Switching

Research description:

People often try to trick chatbots into unsafe answers by rephrasing questions or switching languages. A safe healthcare assistant must resist these tactics in realistic conversations. This project will test whether an AI assistant keeps refusing when a user becomes more insistent or changes language during a dialogue.

The SUDS Scholar will design simulated conversations where the user asks for things that the AI should not provide, such as risky medical tips or harmful instructions. Over several turns, the tone of the user will move from polite to demanding or distressed while sometimes shifting from English to another language to probe consistency. The responses from the assistant will be logged and reviewed to see if it ever stops refusing. The student will build synthetic dialogues with a language model as the assistant and a scripted or model based agent as the persistent user. The analysis will track how often the assistant maintains a refusal and how the wording of responses shifts with pressure from the user. The project will also compare different prompt styles that shape system instructions for the assistant and will store reusable templates and logs for future safety audits in clinical deployment inside hospitals and other health environments.

Year: 2026

Researcher:
 Zahra Shakeri, University of Toronto, Dalla Lana School of Public Health, Institute of Health Policy, Management, and Evaluation

Students:
Fang Sheng, University of Toronto
Abdullah Wadie Bukhari
King Abdullah University of Science & Technology
Nasser Mohammed N Altamimi, King Abdullah University of Science & Technology
Nawaf Abdullah A Alahmed,
King Abdullah University of Science & Technology

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Accelerating Cosmic Discovery with Bayesian Analysis

Research description:

Type Ia Supernovae are calibrated standard light beacons that enable us to measure distances across cosmic time. These distances encode the expansion history of the Universe; however, one of the biggest challenges is finding a “pure” sample of these supernovae, given that many things explode in the night sky, and only some of those are useful cosmological probes. The Vera C Rubin Observatory is a telescope that takes images of the sky and will find hundreds of thousands of these objects, contaminated by other light sources.  Our group is working on a fully Bayesian supernova cosmology analysis pipeline to process the incoming Rubin data.

There are many aspects to this analysis, including parametrizing supernova rates over time, modelling supernova spectra, and more practical considerations such as optimizing the analytic and numerical runtime, and performing coverage tests. Depending on the SUDS Scholar’s interests and strengths, your tasks could include developing statistical tests to determine the accuracy of the Bayesian model, using conformal prediction or similar methods to improve quantified uncertainties, performing an independent analysis on an alternate supernova dataset, or optimizing the code for accuracy or performance.

Year: 2026

Researcher: 
Renee Hlozek, University of Toronto, Faculty of Arts and Science, David A. Dunlap Department of Astronomy and Astrophysics

Students:
Sulaiman Wael Alangari, King Abdullah University of Science & Technology
Linh Vo, York University

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.