Papers in Order

Machine Learning Systems

From mixed-precision arithmetic to trillion-parameter training runs and paged KV caches — the frameworks, parallelism strategies, and serving engines that turn models into products.

16 papers 5 levels Included papers 2016 – 2023 MVRP 5 papers
0 of 16 papers read. Progress stays in this browser.

Minimum viable reading path

The 5 papers that give you most of the field's mental model, in reading order.

  1. Mixed Precision Training
  2. PyTorch: An Imperative Style, High-Performance Deep Learning Library
  3. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
  4. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
  5. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM)
Level 0 Foundations numerics & data parallelism
01 Mixed Precision Training MVRP
TL;DR

Trains in FP16 with an FP32 master copy and loss scaling, halving memory with no accuracy loss.

Why read this

The numerics recipe under every modern training run — and the reason tensor cores exist.

Prerequisites

Floating-point representation, backprop

Key takeaway

Most of training tolerates low precision if you protect the accumulations.

Read the paper
02 Horovod: Fast and Easy Distributed Deep Learning in TensorFlow
TL;DR

Ring-allreduce data parallelism as a drop-in library, replacing parameter servers.

Why read this

Made multi-GPU training accessible and popularised the allreduce pattern now inside every framework.

Prerequisites

SGD, basic networking

Key takeaway

Bandwidth-optimal collectives, not central servers, are the right shape for gradient exchange.

Read the paper
Level 1 Frameworks & compilers
03 TensorFlow: A System for Large-Scale Machine Learning
TL;DR

A dataflow-graph execution engine spanning mobile to datacenter, with automatic differentiation.

Why read this

The first industrial-scale ML framework paper; defines the static-graph design point.

Prerequisites

Backprop, dataflow concepts

Key takeaway

Representing computation as a graph enables distribution, optimisation, and heterogeneity.

Read the paper
04 PyTorch: An Imperative Style, High-Performance Deep Learning Library MVRP
TL;DR

Eager, Pythonic execution with dynamic autograd, engineered to stay fast via async dispatch.

Why read this

The framework that won research; the design argument for usability as a systems property.

Prerequisites

TensorFlow, Python

Key takeaway

Developer experience is a legitimate systems-design objective, not a trade-off.

Read the paper
05 TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
TL;DR

Compiles models to diverse hardware using learned cost models to search schedule spaces.

Why read this

Defines the ML-compiler layer between frameworks and silicon; ancestor of today's kernel autotuners.

Prerequisites

Compilers basics, GPU programming

Key takeaway

Kernel optimisation is a search problem — and ML can drive the search.

Read the paper
Level 2 Distributed training
06 GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
TL;DR

Splits a model into stages across accelerators and pipelines micro-batches through them.

Why read this

The canonical pipeline-parallelism paper — one of the three axes of 3D parallelism.

Prerequisites

Data parallelism, backprop memory costs

Key takeaway

Micro-batching keeps pipeline stages busy and makes re-materialisation affordable.

Read the paper
07 Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism MVRP
TL;DR

Splits attention and MLP matrix multiplies across GPUs with just two communication points per layer.

Why read this

Tensor parallelism as used in practically every large training run since.

Prerequisites

Transformer, collective communication

Key takeaway

Partition the matmuls, not the layers, and communication stays cheap.

Read the paper
08 ZeRO: Memory Optimizations Toward Training Trillion Parameter Models MVRP
TL;DR

Shards optimizer states, gradients, and parameters across data-parallel workers instead of replicating them.

Why read this

The memory analysis every practitioner needs; the engine inside DeepSpeed and FSDP.

Prerequisites

Adam, data parallelism, mixed precision

Key takeaway

Optimizer state, not parameters, dominates training memory — so shard it.

Read the paper
09 GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
TL;DR

Annotation-driven compiler sharding plus mixture-of-experts, demonstrated on a 600B model.

Why read this

Where sparse scaling met compiler-managed parallelism; read alongside Switch Transformer.

Prerequisites

Megatron-LM, MoE basics

Key takeaway

Let the compiler place the computation; the programmer only annotates intent.

Read the paper
Level 3 Datacenter-scale orchestration
10 Ray: A Distributed Framework for Emerging AI Applications
TL;DR

A unified task-and-actor engine with a distributed object store for heterogeneous ML workloads.

Why read this

The substrate under much of modern RL, hyperparameter search, and LLM serving (vLLM runs on it).

Prerequisites

Distributed systems basics, futures/actors

Key takeaway

ML workloads need dynamic task graphs, not just static dataflow.

Read the paper
11 Pathways: Asynchronous Distributed Dataflow for ML
TL;DR

Single-controller orchestration of thousands of accelerators with gang-scheduled asynchronous dataflow.

Why read this

The infrastructure paper behind PaLM; the design point for multi-pod training.

Prerequisites

TensorFlow, GPipe

Key takeaway

A single logical controller can scale if dispatch is asynchronous and sharded.

Read the paper
12 In-Datacenter Performance Analysis of a Tensor Processing Unit
TL;DR

Retrospective on the TPUv1's systolic array delivering 15–30× perf/watt over contemporary GPUs.

Why read this

The hardware end of the ML-systems stack; read with Roofline in mind.

Prerequisites

Roofline model, matrix multiply

Key takeaway

Domain-specific silicon wins when the workload is one operator.

Read the paper
Level 4 Inference & serving
13 FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
TL;DR

Reworks FlashAttention's loop structure and warp partitioning to reach ~70% of peak FLOPs.

Why read this

The production attention kernel; a masterclass in mapping algorithms to GPU execution hierarchy.

Prerequisites

FlashAttention, GPU architecture

Key takeaway

Getting from 30% to 70% of peak is about work partitioning, not new math.

Read the paper
14 Efficiently Scaling Transformer Inference
TL;DR

An analytical cost model for partitioning LLM inference, separating prefill and decode regimes.

Why read this

The clearest thinking on why inference economics differ from training economics.

Prerequisites

Megatron-LM, Roofline model

Key takeaway

Decode is memory-bandwidth-bound; batch it or pay for it.

Read the paper
15 Fast Inference from Transformers via Speculative Decoding
TL;DR

A small draft model proposes tokens the large model verifies in parallel, provably preserving its distribution.

Why read this

Now standard in serving stacks; a rare free-lunch latency win.

Prerequisites

Autoregressive decoding, rejection sampling

Key takeaway

Verification is parallelisable even when generation isn't.

Read the paper
16 Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) MVRP
TL;DR

Virtual-memory-style paging for KV caches eliminates fragmentation and boosts serving throughput 2–4×.

Why read this

The paper behind vLLM — the default open-source LLM serving engine.

Prerequisites

Transformer KV cache, OS paging

Key takeaway

Classic OS ideas transplant directly into ML serving.

Read the paper