Machine Learning Systems

LLMs place heavy demands on the systems that serve them: a single forward pass moves gigabytes of weights and cache data through the memory hierarchy. The deployment costs depend heavily on the throughput and latency one can extract from the GPUs. This course studies how modern LLMs are made to run efficiently, working from the hardware upward. We begin with an analytical framework for reasoning about performance: arithmetic intensity, the roofline model, and careful accounting of FLOPs and bytes a transformer layer consumes. We then treat the GPU as a machine, discussing its memory hierarchy and execution model. We also study the CUDA programming model used to write high-performance kernels, with matrix multiplication as the running example. The second half of the course focuses on computational characteristics of LLM inference (prefill, decode, KV-cache) and the system techniques that make serving practical at scale.

Learning Objectives

After completing the course, students should be able to:

  • Reason quantitatively about ML systems performance
  • Write and optimize high-performance GPU kernels
  • Characterize the anatomy of LLM inference
  • Apply system techniques for efficient LLM serving
  • Kirk, David B., and W. Hwu Wen-Mei. Programming massively parallel processors: a hands-on approach. Morgan kaufmann, 2016.
  • Austin et al., "How to Scale Your Model", Google DeepMind, online, 2025.