SUDS Project:

Weight and activation quantization using mathematical optimization for efficient training and inference in large language models

Research description:

This project focuses on developing new quantization methods for representing the weights and activations of large language models as numbers with lower precisions to achieve faster training and inference for large language models while minimizing the reconstruction error. In 2024, several new methods including EasyQuant and SqueezeLLM are proposed for quantizing LLMs to reduce training and inference time under acceptably low reconstruction errors. While the existing methods provide remarkable performance, it is expected that a quantization algorithm that relies on mathematical optimization can exceed the performance of existing methods. In this project, the SUDS scholar will be supervised by a faculty member from the MIE department to complete a series of weekly assignments. These tasks will encompass activities such as data analysis, computational experiments, and the implementation and testing of new algorithm enhancements in a git environment. This project leverages cutting-edge techniques in mathematical optimization to advance the quantization of LLMs by reducing reconstruction error. The results of this summer research initiative contribute to the development of a new algorithm for weight and activation quantization of large language models, thereby enhancing a widely used AI technology in using data science.

Year: 2025

Researcher:
Samin Aref, Department of Mechanical and Industrial Engineering, Faculty of Applied Science and Engineering, University of Toronto

Student: 
Yixin (Amanda) Yin, University of Toronto

Through SUDS, undergraduate students engage in hands-on research focused on data sciences and AI methodology applications.