Research description:
Speech and natural language processing have become major forces in AI, with increasingly perceptible impacts on our daily lives. But very little is really understood about the way that modern speech representations trained with machine learning really work (e.g., the popular wav2vec 2.0 or Whisper features that have revolutionized speech processing in recent years). They are “black boxes”. At the same time, phonetics, the scientific study of speech sounds and perception, has not taken advantage of the massive potential of recent speech representations trained with machine learning, which promise to help us better understand how human speech works. This is in part because they are hard to use, and in part because they are opaque and would require in-depth analysis to be able to interpret what they are doing.
This project has two aspects. The student will be engaged in adding features to Speech Features Online, an existing online platform aimed at making speech representations accessible to non-experts. Second, the student will develop approaches for analyzing these representations that help bridge the gap between modern language and speech sciences, and modern machine learning approaches to speech processing. This project is a step towards changing the way we understand human speech.
Year: 2024
Researcher:
Ewan Dunbar, Department of French, Faculty of Arts & Science, University of Torotno
Student:
Robin Huo, University of Toronto