SUDS project

Learning the DNA characteristics of mutational processes in cancer

Research description:

Cancer is a genetic disease caused by small mutations in DNA that occur in individual’s cells over time. Most mutations are harmless “passenger” mutations while a small minority of mutations termed “driver” mutations unlock the features of cells that lead to cancer. Passenger mutations tell us about the history of the cancer and how mutations arise due to age, carcinogens, or deficient DNA repair processes in cells. Thousands of cancer genomes with millions of mutations are now available. These datasets show that mutations do not occur randomly but instead have nucleotide characteristics (such as C>T mutations correlated with patient age vs. C>A mutations associated with tobacco smoking). However, these “mutational signatures” are based on very limited DNA context, usually just the two nucleotides around the mutated position. The objective of this research project is to develop sequence-base machine learning models that classify or generate cancer mutations based on their mutational process that caused the mutations, or the cancer type they occur in. In addition to developing accurate models, we aim to enhance model interpretation and decipher the sequence features contributing most to model performance, allowing us to better understand how mutations contribute to cancer development and molecular complexity. The student is expected to develop and test ML models using R or python coding, interpret data from computational and biological angles, visualize data, prepare documentation, and present at lab meetings. We will finetune the project based on the computational and/or biological or disease research interests of the student.

Year: 2024

Researcher:
Judi Reimand, Ontario Institute for Cancer Research

Students: 
Keren Zhang, University of Toronto
Yahya Abdullah Alhabboub, King Abdullah University of Science & Technology

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Interactive approaches to automatic source code summarization using deep learning

Research description:

Automatic source code summarization is the task of generating a readable summary that describes the functionality of the code in natural language. In recent years, the use of deep learning-based approaches has led to significant improvement in the performance of automatic code summarization, e.g., using Transformers and Graph Neural Networks. However, the performance is still far from optimal and developers that are unsatisfied with a given summary are not able to provide feedback or additional information that can be used to refine the output.

In this research project, the goal is to investigate ways in which additional input from the developer can further improve the performance of automatic code summarization. Specifically, the main tasks in the project are:

  1. Investigating existing failures of state-of-the-art source code summarization solutions
  2. Developing new computational approaches and interactive schemes for incorporating developer input and feedback in order to improve the performance of deep learning-based approaches for source code summarization
  3. Evaluating the impact of the new approaches on existing large code summarization datasets.

The responsibilities of the SUDS student will be:

  1. Read about, implement, and empirically evaluate state-of-the-art models for automatic source code summarization.
  2. Investigate existing failures of state-of-the-art source code summarization solutions and develop interactive schemes for incorporating developer input and feedback in order to improve their performance.

Evaluating the impact of the new approaches on existing large code summarization datasets.

Year: 2024

Researcher:
Eldan Cohen, Department of Mechanical and Industrial Engineering, Faculty of Applied Science & Engineering, University of Toronto

Student: 
Yifan Liu, University of Toronto

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Improving a flexible search system for high-accuracy identification of biological entities and molecules using AI

Research description:

Year: 2024

Researcher:
Gary Bader, Terrence Donnely Centre for Cellular & Biomedical Research, University of Toronto

Student: 
Fatemah Alsolaiman, King Abdullah University of Science & Technology

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Dating stars with contrastive learning

Research description:

Because we can only observed the Milky Way at the present time, obtaining ages for large numbers of stars is crucial to unraveling our Galaxy’s history. However, ages are notoriously difficult to obtain using traditional astronomical techniques. The most robust method for determining ages uses time series of red giant stars; a star’s age is directly reflected in the random oscillations that such stars undergo and that we can observe using detailed time series observations. However, these observations are expensive and difficult to model.

Obtaining high-resolution spectra using a diffraction grating is much easier and such samples now consist of about a million stars. But while we believe these spectra contain age information, we have no robust theory to extract it. This is where machine learning comes in! In this project, we will use contrastive learning to extract the age information from stellar spectra using similar techniques as used to, for example, provide captions for images (see, e.g., OpenAI’s clip). We will use this to obtain ages for large numbers of stars in the APOGEE and SDSS-V surveys and determine the age distribution of stars across the Milky Way’s disk. The student will be responsible for implementing the contrastive learning process in Pytorch using data that we will provide and for evaluating the model’s performance using a test set and by comparing to the results from other, previous techniques.

Year: 2024

Researcher:
Jo Bovy, David A. Dunlap Department of Astronomy and Astrophysics, Faculty of Arts & Science, University of Toronto

Student: 
Yiwei Jiang, University of Toronto

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Developing an intelligent health equity dashboard

Research description:

Disparities in health outcomes represent one of the most challenging issues our healthcare system is currently facing. The development an equity dashboard in hospitals has been proposed as a solution to facilitate the identification of variations in outcomes, encourage accountability, and support ongoing monitoring. However, limited data, lack of data (including demographic attributes), and insensitive measures can render the development of equity dashboards challenging. In order to gain insights on potential variations in care, multiple sources of quantitative and qualitative data – including EMR documentation, incident reports, patient feedback, and various outcomes – need to be linked and leveraged to create a broader understanding of clinical systems inequities. Using maternal care as a case study, this project will utilize incident report data, patient experience data, and outcome data to develop an equity dashboard that can be used to inform decision-making.

The responsibilities of the student will be as follows:

  • Complete TCPS 2.0 Research Ethics Training
  • Review literature on maternal mortality and disparities
  • Conduct statistical analysis on disaggregated data to identify differences in outcomes and narrow down outcomes of interests – SMM indicators, adverse events, process measures
  • Assist with developing and evaluating predictive models based on patient characteristics and social vulnerability indices
  • Conduct data preprocessing
  • Train and evaluate different modelsAssess fairness
  • Utilize explainable artificial intelligence techniques
  • Design and test an interactive dashboard of the outcomes (in Python, Tableau, or Power BI)
  • Develop visualizations that meaningfully convey potential disparities in care and outcomes
  • Incorporate XAI explanations as necessary to support transparency
  • Conduct usability testing to iterate design
  • Compose abstract of the findings
  • Present research at UnERD conference

Year: 2024

Researcher:
Myrtede Alfred, Department of Mechanical and Industrial Engineering, Faculty of Applied Science and Engineering, University of Toronto

Students: 
Ning Bao, University of Toronto
Abdulaziz Essam AlTelmissani, King Abdullah University of Science & Technology

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.