Papers in Order

LLMs & Transformers

From the first dense word vectors to frontier reasoning models — the architectures, the scaling laws, and the alignment recipes that produced ChatGPT and its descendants.

37 papers 9 levels Included papers 2013 – 2025 MVRP 11 papers
0 of 37 papers read. Progress stays in this browser.

Minimum viable reading path

The 11 papers that give you most of the field's mental model, in reading order.

  1. Attention Is All You Need
  2. Improving Language Understanding by Generative Pre-Training (GPT-1)
  3. Language Models are Unsupervised Multitask Learners (GPT-2)
  4. Language Models are Few-Shot Learners (GPT-3)
  5. Training Language Models to Follow Instructions with Human Feedback (InstructGPT)
  6. LLaMA: Open and Efficient Foundation Language Models
  7. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
  8. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG)
  9. LoRA: Low-Rank Adaptation of Large Language Models
  10. The Llama 3 Herd of Models
  11. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Level 0 Foundations word vectors & seq2seq
01 Efficient Estimation of Word Representations in Vector Space (Word2Vec)
TL;DR

Learns dense word vectors via Skip-gram and CBOW objectives over huge text corpora.

Why read this

The seed paper for distributed semantic representations — every embedding-based system traces back here.

Prerequisites

Basic neural networks, softmax classification

Key takeaway

Words appearing in similar contexts end up close in vector space, and you can do arithmetic on meaning.

Read the paper
02 GloVe: Global Vectors for Word Representation
TL;DR

Factorises the global word co-occurrence matrix to learn vectors that capture linear semantic structure.

Why read this

The other half of the embedding canon — useful contrast to Word2Vec's local-context approach.

Prerequisites

Word2Vec, matrix factorisation

Key takeaway

Co-occurrence ratios, not raw counts, are what encode meaning.

Read the paper
03 Sequence to Sequence Learning with Neural Networks
TL;DR

Two stacked LSTMs encode a source sequence into a vector and decode the target sequence from it.

Why read this

Establishes the encoder–decoder template that attention and Transformers later supercharged.

Prerequisites

LSTMs, backprop through time

Key takeaway

You can pose almost any structured prediction task as sequence-to-sequence.

Read the paper
Level 1 The Transformer
04 Attention Is All You Need MVRP
TL;DR

Introduces the Transformer: an attention-only architecture with no recurrence or convolution.

Why read this

The architectural blueprint underlying every modern LLM. Read it for the original framing.

Prerequisites

Seq2seq with attention, matrix multiplication

Key takeaway

Self-attention is fully parallelisable, captures long-range dependencies, and scales beautifully.

Read the paper
Level 2 Pretraining era
05 Deep Contextualized Word Representations (ELMo)
TL;DR

Trains a bidirectional LSTM language model and uses its hidden states as context-sensitive word features.

Why read this

The first widely-used "context-aware" embedding, bridging static vectors and full pretraining.

Prerequisites

Word2Vec, LSTMs

Key takeaway

The same word should have different vectors in different sentences.

Read the paper
06 Improving Language Understanding by Generative Pre-Training (GPT-1) MVRP
TL;DR

Decoder-only Transformer pretrained as a language model, then fine-tuned for downstream tasks.

Why read this

The origin point of the GPT line and of "pretrain then fine-tune" as the dominant NLP recipe.

Prerequisites

Transformer architecture, cross-entropy loss

Key takeaway

A single generative pretraining objective transfers to almost any NLP task.

Read the paper
07 BERT: Pre-training of Deep Bidirectional Transformers
TL;DR

Encoder-only Transformer pretrained with masked-language-modeling and next-sentence prediction.

Why read this

Read against GPT-1 to understand the encoder-vs-decoder split that dominated 2018–2020 NLP.

Prerequisites

Transformer architecture, GPT-1

Key takeaway

Bidirectional context (mask-fill) makes a far stronger representation than left-to-right alone.

Read the paper
Level 3 Scaling up
08 Language Models are Unsupervised Multitask Learners (GPT-2) MVRP
TL;DR

Scales GPT-1 to 1.5B parameters on a 40GB web corpus and shows surprising zero-shot abilities.

Why read this

The first hint that scale alone unlocks new capabilities without task-specific training.

Prerequisites

GPT-1, basic information theory

Key takeaway

A big enough language model is implicitly a multi-task learner.

Read the paper
09 RoBERTa: A Robustly Optimized BERT Pretraining Approach
TL;DR

Trains BERT longer, on more data, with bigger batches and no NSP — and beats every BERT variant.

Why read this

A masterclass in how much pretraining hyperparameters matter.

Prerequisites

BERT

Key takeaway

Many architectural "improvements" are really just under-trained baselines.

Read the paper
10 Language Models are Few-Shot Learners (GPT-3) MVRP
TL;DR

A 175B-parameter Transformer that performs many tasks via in-context examples, no gradient updates.

Why read this

The paper that launched the modern era; in-context learning is the central abstraction of the LLM stack.

Prerequisites

GPT-2, basic prompting

Key takeaway

At sufficient scale, "show, don't train" becomes a viable interface to a model.

Read the paper
11 Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)
TL;DR

Casts every NLP task as text-to-text and pretrains a giant encoder-decoder on a clean web corpus (C4).

Why read this

The clearest empirical study of pretraining design choices, plus the C4 dataset most LLMs still use.

Prerequisites

BERT, seq2seq

Key takeaway

Unify your tasks as text-to-text and almost everything becomes simpler.

Read the paper
Level 4 Scaling laws & compute
12 Scaling Laws for Neural Language Models
TL;DR

Empirically derives smooth power-law relationships between loss, parameters, data, and compute.

Why read this

Establishes that LLM progress is predictable and that you can budget compute analytically.

Prerequisites

GPT-2, basic statistics

Key takeaway

Loss falls as a power law in compute — and the exponent tells you exactly how to spend it.

Read the paper
13 Training Compute-Optimal Large Language Models (Chinchilla)
TL;DR

Re-runs scaling laws and shows existing models are massively under-trained relative to their size.

Why read this

Corrects Kaplan's recipe; sets the modern data-to-parameters ratio used by Llama and beyond.

Prerequisites

Kaplan scaling laws

Key takeaway

For a fixed compute budget, you should scale parameters and tokens roughly equally.

Read the paper
14 PaLM: Scaling Language Modeling with Pathways
TL;DR

A 540B dense Transformer trained on Google's Pathways system, surfacing emergent reasoning capabilities.

Why read this

The high-water mark for dense models pre-Chinchilla and a careful study of capability emergence.

Prerequisites

GPT-3, scaling laws

Key takeaway

Some abilities appear suddenly and only at scale — they aren't visible at smaller sizes.

Read the paper
30 Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
TL;DR

Routes each token to a single expert FFN, training trillion-parameter sparse models with dense-model stability.

Why read this

The paper that made mixture-of-experts practical at scale — read it before Mixtral or any modern MoE.

Prerequisites

T5, scaling laws

Key takeaway

Sparsity lets parameter count and compute cost scale independently.

Read the paper
31 Emergent Abilities of Large Language Models
TL;DR

Catalogues abilities that appear abruptly at scale rather than improving smoothly.

Why read this

Frames the central scientific puzzle of scaling — and the debate over whether emergence is real or a metric artefact.

Prerequisites

Scaling laws, GPT-3

Key takeaway

Some capabilities are invisible at small scale, which makes extrapolation genuinely hard.

Read the paper
32 GPT-4 Technical Report
TL;DR

Reports GPT-4's capabilities, evals, and safety work — while withholding architecture and training details.

Why read this

A milestone in capability and a turning point in (non-)openness; also introduces predictable-scaling loss forecasts.

Prerequisites

GPT-3, InstructGPT

Key takeaway

Frontier performance became predictable from small-scale runs — and frontier papers stopped disclosing how.

Read the paper
Level 5 Alignment
15 Training Language Models to Follow Instructions with Human Feedback (InstructGPT) MVRP
TL;DR

Fine-tunes GPT-3 with supervised demonstrations and PPO against a learned reward model (RLHF).

Why read this

The recipe that turned a raw language model into a usable assistant — the real ChatGPT precursor.

Prerequisites

GPT-3, basic RL (policy gradients)

Key takeaway

A small amount of human feedback beats a huge amount of next-token training for usefulness.

Read the paper
16 Constitutional AI: Harmlessness from AI Feedback
TL;DR

Replaces human preference labels with model-generated critiques guided by a written "constitution".

Why read this

A scalable alternative to RLHF that frames alignment as supervisable principles, not just preferences.

Prerequisites

InstructGPT, RLHF

Key takeaway

You can train safety using the model itself as the labeler — if the principles are explicit.

Read the paper
17 Direct Preference Optimization (DPO)
TL;DR

Derives a closed-form loss that fits a preference dataset directly, no reward model or PPO required.

Why read this

Drastically simpler than RLHF; now the default for open-source preference tuning.

Prerequisites

InstructGPT, KL divergence

Key takeaway

Preference learning can be a one-stage classification problem rather than two-stage RL.

Read the paper
Level 6 Open weights
18 LLaMA: Open and Efficient Foundation Language Models MVRP
TL;DR

A family of 7B–65B Chinchilla-style models trained on public data, released to researchers.

Why read this

The leak-and-release that catalyzed the modern open-weight LLM ecosystem and birthed the open-weight ecosystem.

Prerequisites

GPT-3, Chinchilla

Key takeaway

A well-trained 7B model can rival GPT-3 — open weights matter.

Read the paper
19 Llama 2: Open Foundation and Fine-Tuned Chat Models
TL;DR

Successor to LLaMA, with full RLHF chat variants, more pretraining data, and a permissive license.

Why read this

The first commercially-usable competitive open model — a watershed for industry adoption.

Prerequisites

LLaMA, InstructGPT

Key takeaway

Open + commercial-friendly + RLHF'd turned out to be the unlock for adoption.

Read the paper
20 Mistral 7B
TL;DR

A 7B model with grouped-query and sliding-window attention that beats Llama 2 13B on most benchmarks.

Why read this

Demonstrates how much performance is left on the table by architectural choices at small scale.

Prerequisites

Llama 2, attention variants

Key takeaway

Smarter attention beats bigger model size, dollar for dollar.

Read the paper
21 Mixtral of Experts
TL;DR

A sparse mixture-of-experts model where each token routes to 2 of 8 expert FFNs per layer.

Why read this

The clearest example of MoE at production scale; shows where dense scaling is heading.

Prerequisites

Mistral 7B, Switch Transformer

Key takeaway

Activate only a fraction of parameters per token and you get bigger-model quality at smaller-model cost.

Read the paper
33 DeepSeek-V3 Technical Report
TL;DR

A 671B-parameter MoE (37B active) trained for ~$5.5M using MLA attention, FP8, and aggressive systems co-design.

Why read this

Rewrote assumptions about frontier training costs; the strongest open-weight base model of its moment.

Prerequisites

Mixtral, FlashAttention

Key takeaway

Model architecture and training systems co-designed together beat either optimised alone.

Read the paper
Level 7 Reasoning & retrieval
22 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models MVRP
TL;DR

Adding "let's think step by step" exemplars to prompts massively improves multi-step reasoning.

Why read this

The most influential prompting paper; reframes inference-time compute as a useful lever.

Prerequisites

GPT-3, few-shot prompting

Key takeaway

Letting the model show its work makes it correct more often — and only at scale.

Read the paper
23 ReAct: Synergizing Reasoning and Acting in Language Models
TL;DR

Interleaves chain-of-thought with tool calls, letting the model think and act in alternating steps.

Why read this

The conceptual spine of every "agent" framework — read this before any agent paper.

Prerequisites

Chain-of-Thought, basic API/tool concepts

Key takeaway

Reasoning and acting are stronger together than either alone.

Read the paper
24 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG) MVRP
TL;DR

Combines a Transformer generator with a learned dense retriever over a non-parametric corpus.

Why read this

The original RAG; the foundation of every "talk to your docs" system.

Prerequisites

Transformers, basic retrieval (BM25 or DPR)

Key takeaway

External memory beats trying to cram every fact into the weights.

Read the paper
25 Self-Consistency Improves Chain of Thought Reasoning
TL;DR

Sample many CoT trajectories and majority-vote the final answer, instead of greedy decoding.

Why read this

A simple, free win on top of any CoT model — and the conceptual ancestor of test-time scaling.

Prerequisites

Chain-of-Thought

Key takeaway

Many noisy answers can be aggregated into one good answer.

Read the paper
34 Toolformer: Language Models Can Teach Themselves to Use Tools
TL;DR

Self-supervised annotation teaches an LM when and how to call APIs like calculators and search.

Why read this

The cleanest early statement of tool use as a learned capability rather than a prompt hack.

Prerequisites

GPT-3, ReAct

Key takeaway

A model can bootstrap its own tool-use training data by checking which calls reduce its loss.

Read the paper
35 Tree of Thoughts: Deliberate Problem Solving with Large Language Models
TL;DR

Generalises chain-of-thought into a search tree with lookahead, backtracking, and self-evaluation.

Why read this

A conceptual bridge from prompting tricks to genuine inference-time search.

Prerequisites

Chain-of-Thought, Self-Consistency

Key takeaway

Reasoning improves when the model can explore and prune alternatives, not just sample forward.

Read the paper
36 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning MVRP
TL;DR

Pure RL on verifiable rewards elicits long chain-of-thought reasoning, distilled into open models.

Why read this

The open counterpart to o1-style reasoning models — made test-time-compute training reproducible.

Prerequisites

Chain-of-Thought, PPO/GRPO, DeepSeek-V3

Key takeaway

Reasoning behaviour can emerge from RL against checkable answers, without human reasoning traces.

Read the paper
Level 8 Efficiency
26 LoRA: Low-Rank Adaptation of Large Language Models MVRP
TL;DR

Freezes the base model and trains low-rank update matrices added to its weights.

Why read this

The technique that made fine-tuning huge models accessible to anyone with one GPU.

Prerequisites

Linear algebra, fine-tuning basics

Key takeaway

Most fine-tuning lives in a low-rank subspace; you don't need to touch every weight.

Read the paper
27 QLoRA: Efficient Finetuning of Quantized LLMs
TL;DR

Combines 4-bit base model quantization with LoRA adapters and paged optimizers.

Why read this

Brings 65B-parameter fine-tuning to a single consumer GPU.

Prerequisites

LoRA, quantization basics

Key takeaway

4-bit + low-rank is enough to match full-precision fine-tuning quality.

Read the paper
28 FlashAttention: Fast and Memory-Efficient Exact Attention
TL;DR

A tiling-based, IO-aware attention kernel that reduces memory reads/writes without changing the math.

Why read this

The kernel that's now under almost every Transformer in production.

Prerequisites

Transformer attention, GPU memory hierarchy

Key takeaway

Optimising for memory bandwidth, not FLOPs, is the right move on modern GPUs.

Read the paper
29 The Llama 3 Herd of Models MVRP
TL;DR

A detailed engineering report on training Llama 3 (8B, 70B, 405B) on 15T tokens.

Why read this

The most thorough open description of a frontier-grade training run — the new state of the art for open weights.

Prerequisites

Llama 2, Chinchilla, modern alignment

Key takeaway

Frontier models are mostly an engineering problem now: data, infra, and post-training quality.

Read the paper
37 Mamba: Linear-Time Sequence Modeling with Selective State Spaces
TL;DR

A selective state-space model matching Transformer quality with linear-time, constant-memory inference.

Why read this

The strongest post-Transformer architecture candidate; defines the attention-alternative research line.

Prerequisites

Transformer, state-space models

Key takeaway

Input-dependent state selection recovers the content-based routing that made attention win.

Read the paper