Research area

Distributed Machine Learning

How can distributed learning systems become more efficient and scalable?

At MSRG, we tackle core challenges in distributed deep learning systems. Our research focuses on optimizing training and inference across distributed infrastructure, developing novel approaches for resource allocation and scheduling of training jobs, and creating system-level improvements for deep learning and graph learning frameworks. We are advancing federated learning for resource-constrained and embedded devices, optimizing performance and communication efficiency in edge computing environments. Through our work, we aim to make distributed deep learning more efficient and accessible, enabling faster training and better resource utilization.

7 researchers · 48 publications · 0 data sets

MSRG studies performance optimization, privacy-preserving and federated learning, and graph processing as connected challenges in distributed machine learning.

Distributed Deep Learning

Large training and inference workloads need substantial computing resources and distributed infrastructure. We investigate how resource allocation, job scheduling, and optimizations within learning systems can improve performance and reduce costs.

Our federated learning research also addresses embedded devices with limited resources. We work across hardware-level optimizations and distributed learning algorithms to improve efficiency in edge deployments.

Graph Processing & Learning

Graph workloads combine irregular memory access, interdependent computations, and large datasets. We develop partitioning and sampling methods alongside system optimizations to make these workloads faster.

Our distributed graph systems target graphs containing billions of vertices and edges. We also optimize the training and inference stages of graph neural networks.

Current directions

  • Resource allocation and scheduling for distributed training and inference
  • Privacy-preserving and federated learning on embedded and edge devices
  • Graph partitioning, sampling, and scalable graph neural network execution

Professor

Hans-Arno Jacobsen

Hans-Arno Jacobsen

Professor

Toronto, Canada

Post-Docs

Yiran Li

Yiran Li

Post-Doc

Toronto, Canada

Graph AnalyticsGraph LearningQuantum Data ManagementAI Agent Reliability

Graduate Students PhD

Jana Vatter

Jana Vatter

PhD Student

Munich, Germany

Graph Neural NetworksDistributed Deep LearningGraph Processing
Yifang Tian

Yifang Tian

PhD Student

Toronto, Canada

Machine Learning SystemsDistributed Systems
Dawn (Mengli) Duan

Dawn (Mengli) Duan

PhD Student

Toronto, Canada

Machine Learning Systems

Research Trainees Undergraduate

Ben Cheng

Ben Cheng

Undergraduate Student

Toronto, Canada

Machine Learning SystemsData Management
CZ

Chenkai Zhang

Undergraduate Student

Toronto, Canada

AI Agent ReliabilityFault Diagnosis