Joshua Speagle

Galactic Paleontology: Uncovering the Assembly History of Dwarf Galaxies using their Surviving Stars

Research description:

The growth and assembly of galaxies involves many complex processes, which culminate in the diverse collection of galaxies we observe today. One of the best ways of understanding these processes is through “Galactic Paleontology”, which tries to reconstruct the assembly history of nearby galaxies through their surviving “fossils” (which are their present-day surviving stars!). Using this data, we simulate the birth, evolution, and death of many thousands/millions of stars, compare the end result with the stars we observe today, and repeat this process many times for many different evolutionary pathways to see which ones match the observed data better.

For the past few years, astronomers have largely relied on simulation studies and more “ad hoc” approaches to try to compare simulated data with real data, often involving “binning” the data into larger groups. In this project, co-supervised with Prof. Ting Li, we will develop a new, more principled approach based on Inhomogeneous Poisson Point Processes (IPPP) that will allow us to utilize all of the available data. If time/interest permits, we will also try to compare these results with traditional approaches and potentially explore new probabilistic machine learning-driven methods. The main responsibilities of the student will be to review relevant literature, lead coding and data analysis efforts (using simulation studies and/or real data), and meet regularly with me and various collaborators to discuss progress on the project.

Year: 2024

Researcher:
Joshua Speagle, Department of Statistical Sciences, Faculty of Arts & Science, University of Toronto

Student:
Luke Weizhi, University of Toronto

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Combining Theory and Data in Machine Learning Applications in Astronomy

Research description:

Many machine learning (ML) applications rely on having high-quality, labeled training data that is representative of the type of data the ML model will eventually be applied to. However, in many astronomical applications, we have the exact opposite, with observed training data that is substantially biased (often to the brightest, closest, best-measured objects) relative to the underlying populations of interest (the fainter, faraway, noisier objects). To account for these domain mismatch issues, astronomers often resort to various data augmentation strategies that include making the training data “noisier” and supplementing observed data with simulated data from theoretical models. While these broadly address the fundamental problems, they also tend to degrade the performance of the initial ML model. This project will explore new approaches to improve on these data augmentation strategies using state-of-the-art data from the DESI and SDSS-V astronomical surveys, with the goal of having a model that does strictly better on both observed (real) and simulated (theoretical) data under almost all circumstances. The main responsibilities of the student will be to review relevant literature, lead coding and data analysis efforts (using simulation studies and/or real data), and meet regularly with me and various collaborators to discuss progress on the project.

Year: 2024

Researcher:
Joshua Speagle, Department of Statistical Sciences, Faculty of Arts & Science, University of Toronto

Student: 
Zack Steine, University of Toronto

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.

Hamiltonian Optimization and Sampling for Deep Learning

Research description:

In the last decade, deep learning models have become an integral part of our lives, from image recognition software to large language models such as ChatGPT. These models are often overparameterized, with many (many!) more parameters than training examples. While this naively implies that these models should just memorize their training data, instead, we find that they generalize extremely well, finding solutions that are often even better than more traditional underparameterized models. This almost-magical ability to generalize rather than memorize turns out to be (in part) the result of how we optimize and sample from these overwhelmingly large models’ parameters. Motivated by this behaviour, this project will investigate new variants of optimization and/or sampling methods based on ideas from Hamiltonian optimization, dynamics, and optics, to see how well they perform across a wide class of problems (potentially including large language models). This will involve a combination of theoretical work as well as empirical studies of how these methods perform on various methods on benchmarks, their stability and dynamics under various conditions, and the implicit and/or explicit regularization that they provide.This project will be co-supervised with Prof. Ricardo Baptista, Department of Statistical Sciences, Faculty of Arts & Science, University of Toronto. 
 
The SUDS Scholar will be actively involved in pursuing a combination of both (1) theoretical work and literature review, as well as (2) conducting empirical studies of how these methods perform (including their stability, dynamics, and regularization properties) across various benchmarks.

Year: 2026

Researcher:
Joshua Speagle, University of Toronto, Faculty of Arts and Science, Department of Statistical Sciences

Students:
Chuxuan Ai, University of Toronto
Yasir Abdullah M Alsugair, King Abdullah University of Science & Technology

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.