Research description:
Current methods for developing biocatalysts for sustainable applications require multiple rounds of design and experimentation with significant trial and error steps. Bioengineering can be accelerated significantly if systematic design, testing can be used to narrow the range of candidates that require testing. While there are many deterministic methods for modeling metabolic networks they fall short in their ability to predict the cumulative effect of higher order modifications typically required to enhance the production of desired compounds relevant for applications, for example, adipic acid required for bionylon synthesis. Hence, there is a need to develop hybrid methods that combine deterministic methods with data driven methods that can account for incomplete biological knowledge. Here we aim to use multi-modal large language models including DNA, and protein language models that can be used to predict and correct the current shortcomings of deterministic models. We aim to build on several in house datasets including a pandemic dataset of all metabolic reactions and augment this dataset with curated models derived from gold standard databases such as SWISS PROT and BiGG. We will then use these high quality metabolic network datasets to develop hybrid methods that can predict the impact of genomic modifications on physiology.
The SUDS Scholar will work to develop pipelines to process and curate data from different biological domains specific for microbial metabolism. The scholar will then work with other group members to apply existing pipelines for data analysis and modeling and also facilitate the development of new data analysis methods.
Year: 2026
Researcher: Radhakrishnan Mahadevan, University of Toronto, Faculty of Applied Science and Engineering, Department of Chemical Engineering and Applied Chemistry
Student:
Hussain Faisal Zaid Alharbi
King Abdullah University of Science & Technology