Reinforcement Learning
From dynamic programming and TD learning to AlphaGo and human-feedback fine-tuning — the policies, values, and self-play recipes that built modern decision-making systems.
Minimum viable reading path
The 4 papers that give you most of the field's mental model, in reading order.
01 Learning to Predict by the Methods of Temporal Differences
Introduces TD learning, which updates predictions based on later predictions rather than waiting for outcomes.
The conceptual cornerstone of every value-based RL algorithm — Q-learning, DQN, AlphaGo all build on this.
Calculus, Markov chains
Bootstrapping — learning a guess from a guess — is the central trick of value-based RL.
02 Learning from Delayed Rewards (Q-Learning)
Defines Q-learning: an off-policy TD algorithm that converges to the optimal action-value function.
The single most important classical RL algorithm; the conceptual ancestor of DQN.
TD learning, basic MDPs
You can learn the optimal policy without knowing the environment's dynamics.
03 Simple Statistical Gradient-Following Algorithms (REINFORCE)
Derives the policy-gradient theorem and proposes REINFORCE for parameterised stochastic policies.
The starting point of policy-gradient methods; you can't understand PPO without it.
Calculus, basic probability
You can do gradient descent on expected reward by sampling and weighting log-probabilities.
04 Temporal Difference Learning and TD-Gammon
A neural network trained by self-play TD reaches championship-level Backgammon.
The proof-of-concept that RL + neural nets + self-play can match human experts.
TD learning, neural network basics
Self-play teaches the agent things no human dataset could.
05 Policy Gradient Methods for Reinforcement Learning with Function Approximation
Formalises the policy gradient theorem with function approximation, including actor–critic.
The theoretical foundation under modern actor–critic methods like A3C, PPO, SAC.
REINFORCE, multivariable calculus
You can scale policy gradients to neural networks without breaking convergence guarantees.
06 Playing Atari with Deep Reinforcement Learning (DQN) MVRP
A CNN trained with Q-learning, experience replay, and a target network plays Atari from raw pixels.
The paper that launched deep RL. Read it carefully — every modern algorithm reuses these tricks.
Q-learning, AlexNet
Replay buffer + target network are the duct tape that makes deep value-based RL stable.
07 Human-Level Control through Deep Reinforcement Learning (Nature DQN)
The polished, full-detail DQN paper across 49 Atari games, with rigorous evaluation.
The reference implementation document; far more thorough than the workshop version.
DQN
A single neural architecture, hyperparameters fixed, can reach human level across diverse tasks.
08 Deep Reinforcement Learning with Double Q-learning
Decouples action selection from action evaluation in the Bellman target to fix DQN's overestimation bias.
A clean lesson in how a small mathematical fix can dramatically improve a real algorithm.
DQN
The max in Q-learning is biased upward; you can fix it almost for free.
09 Prioritized Experience Replay
Samples replay transitions in proportion to their TD error rather than uniformly.
Important practical building block; understanding it sharpens your intuition for replay buffers.
DQN
Not all transitions teach equally — sample the surprising ones more.
10 Dueling Network Architectures for Deep RL
Splits the Q-network into separate state-value and advantage streams, recombined at the output.
A small architectural change that disproportionately helps; classic example of inductive bias in RL.
DQN
Decompose Q into how good is this state plus how much better is this action than average.
11 Rainbow: Combining Improvements in Deep RL
Combines six DQN extensions (double, dueling, PER, noisy nets, distributional, n-step) into one agent.
The cleanest study of which DQN tricks actually compose, and which don't.
DQN, Double DQN, PER
Most DQN improvements are complementary, not redundant.
12 Asynchronous Methods for Deep Reinforcement Learning (A3C)
Multiple parallel actor–critic workers update a shared model asynchronously, with no replay buffer.
Showed that on-policy parallelism is a viable alternative to off-policy replay.
Policy gradient, actor–critic
Diversity from parallel actors substitutes for the diversity of a replay buffer.
13 Trust Region Policy Optimization (TRPO)
Constrains policy updates to a trust region defined by KL divergence, with monotonic-improvement guarantees.
The theoretical underpinning of PPO — read TRPO to understand what PPO is approximating.
Policy gradient, KL divergence
Limit how far the policy moves per update and you avoid catastrophic collapses.
14 Proximal Policy Optimization (PPO) MVRP
A simple clipped-objective approximation of TRPO that is far easier to implement and tune.
The default policy gradient method everywhere — including in RLHF for LLMs.
TRPO, basic optimisation
A clipped surrogate objective gives 90% of TRPO's benefit at 10% of the complexity.
26 High-Dimensional Continuous Control Using Generalized Advantage Estimation (GAE)
An exponentially-weighted advantage estimator trading bias against variance with two knobs.
The variance-reduction machinery inside PPO implementations — read it to understand what lambda does.
TD learning, policy gradient
Advantage estimation is a bias–variance dial, and tuning it matters as much as the algorithm.
15 Continuous Control with Deep Reinforcement Learning (DDPG)
Combines DQN-style replay with a deterministic policy gradient for continuous-action control.
The first deep-RL algorithm to scale cleanly to continuous robotic-style action spaces.
DQN, deterministic policy gradient
Off-policy + actor–critic + replay works in continuous spaces too.
16 Soft Actor-Critic (SAC)
A stochastic actor–critic that maximises reward plus policy entropy, dramatically improving robustness.
The current default for continuous-control benchmarks.
DDPG, entropy regularisation
Encouraging entropy in the policy prevents premature convergence to brittle solutions.
17 Mastering the Game of Go with Deep Neural Networks and Tree Search (AlphaGo) MVRP
Combines policy and value networks with Monte Carlo tree search and human-game pretraining.
A turning point for AI in the public mind, and a beautiful integration of deep learning with classical search.
TD learning, CNNs, basic search algorithms
Learned heuristics + planning is far more powerful than either alone.
18 Mastering the Game of Go without Human Knowledge (AlphaGo Zero)
Drops human expert games entirely; learns from random self-play only and surpasses every prior version.
Pure self-play turns out to be sufficient — and produces better play than human-bootstrapped systems.
AlphaGo
Human data was a crutch, not a foundation.
19 A General Reinforcement Learning Algorithm that Masters Chess, Shogi, and Go (AlphaZero)
A single algorithm masters three games tabula-rasa; no game-specific tricks.
Generalises AlphaGo Zero into a general algorithm — a pivotal step toward "general" RL.
AlphaGo Zero
The same algorithm should work for any two-player perfect-information game.
20 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero)
Learns its own model of the environment dynamics in latent space and plans within it.
Removes the requirement for a known simulator — model-based RL that actually scales.
AlphaZero
You don't need the real rules of the game if you can learn a useful enough surrogate.
21 Decision Transformer: Reinforcement Learning via Sequence Modeling
Casts offline RL as conditional sequence modelling: predict next action given past states, actions, and target return.
A radical reframing of RL through the language-model lens.
Transformers, offline RL
You can solve some RL problems by just doing supervised learning on past trajectories.
22 Deep Reinforcement Learning from Human Preferences (RLHF) MVRP
Learns a reward model from human comparisons of trajectory pairs and optimises it with PPO.
The conceptual blueprint behind InstructGPT and ChatGPT — read it once, recognise it everywhere.
PPO, basic preference learning
Humans can label preferences far more reliably than they can specify reward functions.
23 Mastering Diverse Domains through World Models (DreamerV3)
A single model-based agent with fixed hyperparameters mastering 150+ tasks, including Minecraft diamonds.
The strongest evidence to date that world-model RL is finally robust enough to be a default approach.
MuZero, recurrent world models
A well-trained world model can replace a lot of hand-tuned RL machinery.
24 Grandmaster Level in StarCraft II (AlphaStar)
A multi-agent league produces grandmaster-level play in a real-time strategy game with imperfect information.
A frontier-scale demonstration of imperfect-information self-play.
AlphaZero, imitation learning
Population-based training avoids self-play exploits and produces robust strategies.
25 Discovering Faster Matrix Multiplication Algorithms with RL (AlphaTensor)
Casts algorithm discovery as a single-player game and finds new matrix multiplication algorithms.
A glimpse of RL applied to mathematics rather than control.
AlphaZero, basic linear algebra
Game-playing RL agents can rediscover and invent algorithms.
27 Conservative Q-Learning for Offline Reinforcement Learning (CQL)
Regularises Q-learning to lower-bound values of out-of-distribution actions in fixed datasets.
The standard baseline for offline RL — the setting where you learn from logs, not interaction.
Q-learning, SAC
Offline RL fails from overestimating unseen actions; pessimism is the fix.