Papers in Order

Reinforcement Learning

From dynamic programming and TD learning to AlphaGo and human-feedback fine-tuning — the policies, values, and self-play recipes that built modern decision-making systems.

27 papers 8 levels Included papers 1957 – 2024 MVRP 4 papers
0 of 27 papers read. Progress stays in this browser.

Minimum viable reading path

The 4 papers that give you most of the field's mental model, in reading order.

  1. Playing Atari with Deep Reinforcement Learning (DQN)
  2. Proximal Policy Optimization (PPO)
  3. Mastering the Game of Go with Deep Neural Networks and Tree Search (AlphaGo)
  4. Deep Reinforcement Learning from Human Preferences (RLHF)
Level 0 Foundations
01 Learning to Predict by the Methods of Temporal Differences
TL;DR

Introduces TD learning, which updates predictions based on later predictions rather than waiting for outcomes.

Why read this

The conceptual cornerstone of every value-based RL algorithm — Q-learning, DQN, AlphaGo all build on this.

Prerequisites

Calculus, Markov chains

Key takeaway

Bootstrapping — learning a guess from a guess — is the central trick of value-based RL.

Read the paper
02 Learning from Delayed Rewards (Q-Learning)
TL;DR

Defines Q-learning: an off-policy TD algorithm that converges to the optimal action-value function.

Why read this

The single most important classical RL algorithm; the conceptual ancestor of DQN.

Prerequisites

TD learning, basic MDPs

Key takeaway

You can learn the optimal policy without knowing the environment's dynamics.

Read the paper
03 Simple Statistical Gradient-Following Algorithms (REINFORCE)
TL;DR

Derives the policy-gradient theorem and proposes REINFORCE for parameterised stochastic policies.

Why read this

The starting point of policy-gradient methods; you can't understand PPO without it.

Prerequisites

Calculus, basic probability

Key takeaway

You can do gradient descent on expected reward by sampling and weighting log-probabilities.

Read the paper
Level 1 Classical algorithms
04 Temporal Difference Learning and TD-Gammon
TL;DR

A neural network trained by self-play TD reaches championship-level Backgammon.

Why read this

The proof-of-concept that RL + neural nets + self-play can match human experts.

Prerequisites

TD learning, neural network basics

Key takeaway

Self-play teaches the agent things no human dataset could.

Read the paper
05 Policy Gradient Methods for Reinforcement Learning with Function Approximation
TL;DR

Formalises the policy gradient theorem with function approximation, including actor–critic.

Why read this

The theoretical foundation under modern actor–critic methods like A3C, PPO, SAC.

Prerequisites

REINFORCE, multivariable calculus

Key takeaway

You can scale policy gradients to neural networks without breaking convergence guarantees.

Read the paper
Level 2 Deep RL
06 Playing Atari with Deep Reinforcement Learning (DQN) MVRP
TL;DR

A CNN trained with Q-learning, experience replay, and a target network plays Atari from raw pixels.

Why read this

The paper that launched deep RL. Read it carefully — every modern algorithm reuses these tricks.

Prerequisites

Q-learning, AlexNet

Key takeaway

Replay buffer + target network are the duct tape that makes deep value-based RL stable.

Read the paper
07 Human-Level Control through Deep Reinforcement Learning (Nature DQN)
TL;DR

The polished, full-detail DQN paper across 49 Atari games, with rigorous evaluation.

Why read this

The reference implementation document; far more thorough than the workshop version.

Prerequisites

DQN

Key takeaway

A single neural architecture, hyperparameters fixed, can reach human level across diverse tasks.

Read the paper
Level 3 Value-based methods
08 Deep Reinforcement Learning with Double Q-learning
TL;DR

Decouples action selection from action evaluation in the Bellman target to fix DQN's overestimation bias.

Why read this

A clean lesson in how a small mathematical fix can dramatically improve a real algorithm.

Prerequisites

DQN

Key takeaway

The max in Q-learning is biased upward; you can fix it almost for free.

Read the paper
09 Prioritized Experience Replay
TL;DR

Samples replay transitions in proportion to their TD error rather than uniformly.

Why read this

Important practical building block; understanding it sharpens your intuition for replay buffers.

Prerequisites

DQN

Key takeaway

Not all transitions teach equally — sample the surprising ones more.

Read the paper
10 Dueling Network Architectures for Deep RL
TL;DR

Splits the Q-network into separate state-value and advantage streams, recombined at the output.

Why read this

A small architectural change that disproportionately helps; classic example of inductive bias in RL.

Prerequisites

DQN

Key takeaway

Decompose Q into how good is this state plus how much better is this action than average.

Read the paper
11 Rainbow: Combining Improvements in Deep RL
TL;DR

Combines six DQN extensions (double, dueling, PER, noisy nets, distributional, n-step) into one agent.

Why read this

The cleanest study of which DQN tricks actually compose, and which don't.

Prerequisites

DQN, Double DQN, PER

Key takeaway

Most DQN improvements are complementary, not redundant.

Read the paper
Level 4 Policy gradient methods
12 Asynchronous Methods for Deep Reinforcement Learning (A3C)
TL;DR

Multiple parallel actor–critic workers update a shared model asynchronously, with no replay buffer.

Why read this

Showed that on-policy parallelism is a viable alternative to off-policy replay.

Prerequisites

Policy gradient, actor–critic

Key takeaway

Diversity from parallel actors substitutes for the diversity of a replay buffer.

Read the paper
13 Trust Region Policy Optimization (TRPO)
TL;DR

Constrains policy updates to a trust region defined by KL divergence, with monotonic-improvement guarantees.

Why read this

The theoretical underpinning of PPO — read TRPO to understand what PPO is approximating.

Prerequisites

Policy gradient, KL divergence

Key takeaway

Limit how far the policy moves per update and you avoid catastrophic collapses.

Read the paper
14 Proximal Policy Optimization (PPO) MVRP
TL;DR

A simple clipped-objective approximation of TRPO that is far easier to implement and tune.

Why read this

The default policy gradient method everywhere — including in RLHF for LLMs.

Prerequisites

TRPO, basic optimisation

Key takeaway

A clipped surrogate objective gives 90% of TRPO's benefit at 10% of the complexity.

Read the paper
26 High-Dimensional Continuous Control Using Generalized Advantage Estimation (GAE)
TL;DR

An exponentially-weighted advantage estimator trading bias against variance with two knobs.

Why read this

The variance-reduction machinery inside PPO implementations — read it to understand what lambda does.

Prerequisites

TD learning, policy gradient

Key takeaway

Advantage estimation is a bias–variance dial, and tuning it matters as much as the algorithm.

Read the paper
Level 5 Continuous control
15 Continuous Control with Deep Reinforcement Learning (DDPG)
TL;DR

Combines DQN-style replay with a deterministic policy gradient for continuous-action control.

Why read this

The first deep-RL algorithm to scale cleanly to continuous robotic-style action spaces.

Prerequisites

DQN, deterministic policy gradient

Key takeaway

Off-policy + actor–critic + replay works in continuous spaces too.

Read the paper
16 Soft Actor-Critic (SAC)
TL;DR

A stochastic actor–critic that maximises reward plus policy entropy, dramatically improving robustness.

Why read this

The current default for continuous-control benchmarks.

Prerequisites

DDPG, entropy regularisation

Key takeaway

Encouraging entropy in the policy prevents premature convergence to brittle solutions.

Read the paper
Level 6 Self-play & planning
17 Mastering the Game of Go with Deep Neural Networks and Tree Search (AlphaGo) MVRP
TL;DR

Combines policy and value networks with Monte Carlo tree search and human-game pretraining.

Why read this

A turning point for AI in the public mind, and a beautiful integration of deep learning with classical search.

Prerequisites

TD learning, CNNs, basic search algorithms

Key takeaway

Learned heuristics + planning is far more powerful than either alone.

Read the paper
18 Mastering the Game of Go without Human Knowledge (AlphaGo Zero)
TL;DR

Drops human expert games entirely; learns from random self-play only and surpasses every prior version.

Why read this

Pure self-play turns out to be sufficient — and produces better play than human-bootstrapped systems.

Prerequisites

AlphaGo

Key takeaway

Human data was a crutch, not a foundation.

Read the paper
19 A General Reinforcement Learning Algorithm that Masters Chess, Shogi, and Go (AlphaZero)
TL;DR

A single algorithm masters three games tabula-rasa; no game-specific tricks.

Why read this

Generalises AlphaGo Zero into a general algorithm — a pivotal step toward "general" RL.

Prerequisites

AlphaGo Zero

Key takeaway

The same algorithm should work for any two-player perfect-information game.

Read the paper
20 Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero)
TL;DR

Learns its own model of the environment dynamics in latent space and plans within it.

Why read this

Removes the requirement for a known simulator — model-based RL that actually scales.

Prerequisites

AlphaZero

Key takeaway

You don't need the real rules of the game if you can learn a useful enough surrogate.

Read the paper
Level 7 Frontier
21 Decision Transformer: Reinforcement Learning via Sequence Modeling
TL;DR

Casts offline RL as conditional sequence modelling: predict next action given past states, actions, and target return.

Why read this

A radical reframing of RL through the language-model lens.

Prerequisites

Transformers, offline RL

Key takeaway

You can solve some RL problems by just doing supervised learning on past trajectories.

Read the paper
22 Deep Reinforcement Learning from Human Preferences (RLHF) MVRP
TL;DR

Learns a reward model from human comparisons of trajectory pairs and optimises it with PPO.

Why read this

The conceptual blueprint behind InstructGPT and ChatGPT — read it once, recognise it everywhere.

Prerequisites

PPO, basic preference learning

Key takeaway

Humans can label preferences far more reliably than they can specify reward functions.

Read the paper
23 Mastering Diverse Domains through World Models (DreamerV3)
TL;DR

A single model-based agent with fixed hyperparameters mastering 150+ tasks, including Minecraft diamonds.

Why read this

The strongest evidence to date that world-model RL is finally robust enough to be a default approach.

Prerequisites

MuZero, recurrent world models

Key takeaway

A well-trained world model can replace a lot of hand-tuned RL machinery.

Read the paper
24 Grandmaster Level in StarCraft II (AlphaStar)
TL;DR

A multi-agent league produces grandmaster-level play in a real-time strategy game with imperfect information.

Why read this

A frontier-scale demonstration of imperfect-information self-play.

Prerequisites

AlphaZero, imitation learning

Key takeaway

Population-based training avoids self-play exploits and produces robust strategies.

Read the paper
25 Discovering Faster Matrix Multiplication Algorithms with RL (AlphaTensor)
TL;DR

Casts algorithm discovery as a single-player game and finds new matrix multiplication algorithms.

Why read this

A glimpse of RL applied to mathematics rather than control.

Prerequisites

AlphaZero, basic linear algebra

Key takeaway

Game-playing RL agents can rediscover and invent algorithms.

Read the paper
27 Conservative Q-Learning for Offline Reinforcement Learning (CQL)
TL;DR

Regularises Q-learning to lower-bound values of out-of-distribution actions in fixed datasets.

Why read this

The standard baseline for offline RL — the setting where you learn from logs, not interaction.

Prerequisites

Q-learning, SAC

Key takeaway

Offline RL fails from overestimating unseen actions; pessimism is the fix.

Read the paper