Machine Learning Systems
From mixed-precision arithmetic to trillion-parameter training runs and paged KV caches — the frameworks, parallelism strategies, and serving engines that turn models into products.
Minimum viable reading path
The 5 papers that give you most of the field's mental model, in reading order.
- Mixed Precision Training
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM)
01 Mixed Precision Training MVRP
Trains in FP16 with an FP32 master copy and loss scaling, halving memory with no accuracy loss.
The numerics recipe under every modern training run — and the reason tensor cores exist.
Floating-point representation, backprop
Most of training tolerates low precision if you protect the accumulations.
02 Horovod: Fast and Easy Distributed Deep Learning in TensorFlow
Ring-allreduce data parallelism as a drop-in library, replacing parameter servers.
Made multi-GPU training accessible and popularised the allreduce pattern now inside every framework.
SGD, basic networking
Bandwidth-optimal collectives, not central servers, are the right shape for gradient exchange.
03 TensorFlow: A System for Large-Scale Machine Learning
A dataflow-graph execution engine spanning mobile to datacenter, with automatic differentiation.
The first industrial-scale ML framework paper; defines the static-graph design point.
Backprop, dataflow concepts
Representing computation as a graph enables distribution, optimisation, and heterogeneity.
04 PyTorch: An Imperative Style, High-Performance Deep Learning Library MVRP
Eager, Pythonic execution with dynamic autograd, engineered to stay fast via async dispatch.
The framework that won research; the design argument for usability as a systems property.
TensorFlow, Python
Developer experience is a legitimate systems-design objective, not a trade-off.
05 TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Compiles models to diverse hardware using learned cost models to search schedule spaces.
Defines the ML-compiler layer between frameworks and silicon; ancestor of today's kernel autotuners.
Compilers basics, GPU programming
Kernel optimisation is a search problem — and ML can drive the search.
06 GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Splits a model into stages across accelerators and pipelines micro-batches through them.
The canonical pipeline-parallelism paper — one of the three axes of 3D parallelism.
Data parallelism, backprop memory costs
Micro-batching keeps pipeline stages busy and makes re-materialisation affordable.
07 Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism MVRP
Splits attention and MLP matrix multiplies across GPUs with just two communication points per layer.
Tensor parallelism as used in practically every large training run since.
Transformer, collective communication
Partition the matmuls, not the layers, and communication stays cheap.
08 ZeRO: Memory Optimizations Toward Training Trillion Parameter Models MVRP
Shards optimizer states, gradients, and parameters across data-parallel workers instead of replicating them.
The memory analysis every practitioner needs; the engine inside DeepSpeed and FSDP.
Adam, data parallelism, mixed precision
Optimizer state, not parameters, dominates training memory — so shard it.
09 GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Annotation-driven compiler sharding plus mixture-of-experts, demonstrated on a 600B model.
Where sparse scaling met compiler-managed parallelism; read alongside Switch Transformer.
Megatron-LM, MoE basics
Let the compiler place the computation; the programmer only annotates intent.
10 Ray: A Distributed Framework for Emerging AI Applications
A unified task-and-actor engine with a distributed object store for heterogeneous ML workloads.
The substrate under much of modern RL, hyperparameter search, and LLM serving (vLLM runs on it).
Distributed systems basics, futures/actors
ML workloads need dynamic task graphs, not just static dataflow.
11 Pathways: Asynchronous Distributed Dataflow for ML
Single-controller orchestration of thousands of accelerators with gang-scheduled asynchronous dataflow.
The infrastructure paper behind PaLM; the design point for multi-pod training.
TensorFlow, GPipe
A single logical controller can scale if dispatch is asynchronous and sharded.
12 In-Datacenter Performance Analysis of a Tensor Processing Unit
Retrospective on the TPUv1's systolic array delivering 15–30× perf/watt over contemporary GPUs.
The hardware end of the ML-systems stack; read with Roofline in mind.
Roofline model, matrix multiply
Domain-specific silicon wins when the workload is one operator.
13 FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reworks FlashAttention's loop structure and warp partitioning to reach ~70% of peak FLOPs.
The production attention kernel; a masterclass in mapping algorithms to GPU execution hierarchy.
FlashAttention, GPU architecture
Getting from 30% to 70% of peak is about work partitioning, not new math.
14 Efficiently Scaling Transformer Inference
An analytical cost model for partitioning LLM inference, separating prefill and decode regimes.
The clearest thinking on why inference economics differ from training economics.
Megatron-LM, Roofline model
Decode is memory-bandwidth-bound; batch it or pay for it.
15 Fast Inference from Transformers via Speculative Decoding
A small draft model proposes tokens the large model verifies in parallel, provably preserving its distribution.
Now standard in serving stacks; a rare free-lunch latency win.
Autoregressive decoding, rejection sampling
Verification is parallelisable even when generation isn't.
16 Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) MVRP
Virtual-memory-style paging for KV caches eliminates fragmentation and boosts serving throughput 2–4×.
The paper behind vLLM — the default open-source LLM serving engine.
Transformer KV cache, OS paging
Classic OS ideas transplant directly into ML serving.