Research description:
Cancer is a genetic disease caused by small mutations in DNA that occur in individual’s cells over time. Most mutations are harmless “passenger” mutations while a small minority of mutations termed “driver” mutations unlock the features of cells that lead to cancer. Passenger mutations tell us about the history of the cancer and how mutations arise due to age, carcinogens, or deficient DNA repair processes in cells. Thousands of cancer genomes with millions of mutations are now available. These datasets show that mutations do not occur randomly but instead have nucleotide characteristics (such as C>T mutations correlated with patient age vs. C>A mutations associated with tobacco smoking). However, these “mutational signatures” are based on very limited DNA context, usually just the two nucleotides around the mutated position. The objective of this research project is to develop sequence-base machine learning models that classify or generate cancer mutations based on their mutational process that caused the mutations, or the cancer type they occur in. In addition to developing accurate models, we aim to enhance model interpretation and decipher the sequence features contributing most to model performance, allowing us to better understand how mutations contribute to cancer development and molecular complexity. The student is expected to develop and test ML models using R or python coding, interpret data from computational and biological angles, visualize data, prepare documentation, and present at lab meetings. We will finetune the project based on the computational and/or biological or disease research interests of the student.
Year: 2024
Researcher:
Judi Reimand, Ontario Institute for Cancer Research
Students:
Keren Zhang, University of Toronto
Yahya Abdullah Alhabboub, King Abdullah University of Science & Technology