Electrical Engineering and Computer Science Technical Seminar Series
"Improving Scheduling Performance for HPC Applications in Distributed Systems"
Mina Naghshnejad
Department of Engineering and Computer Science, UC Merced
Faculty Host: Mukesh Singhal
Abstract
The main characterization of HPC workloads is high computation loads and large inputs and outputs. With the increasing usage of HPC Clusters, effective scheduling and resource management for these workloads are crucial. Cluster managers try to deliver the desired performance to the customers while keeping their resource usage efficient. This is not possible without careful scheduling of HPC workloads into available resources. However, most scheduling algorithms require accurate estimation application runtime to perform the effective scheduling of applications. In this talk, I first present a hybrid scheduling platform that improves the existing scheduling algorithms by using estimations of runtime prediction accuracy. After that, I talk about using deep mixture density networks to predict application runtime and why it is an appropriate supervised approach to predict application runtimes. Mixture Density Networks estimate application runtimes as a probability density function and train a mixture density network to predict the runtime of a newly submitted application.
Biography
Mina Naghshnejad is a Ph.D. candidate at the Electrical Engineering and Computer Science department of the University of California Merced. Her research interest includes scheduling for distributed systems and applied machine learning. Before that, she got her Master’s in Computer Science at Shahid Beheshti University and her Bachelor of Computer Science from Sharif University of Technology.


