LLMs & Transformers
From the first dense word vectors to frontier reasoning models — the architectures, the scaling laws, and the alignment recipes that produced ChatGPT and its descendants.
Minimum viable reading path
The 11 papers that give you most of the field's mental model, in reading order.
- Attention Is All You Need
- Improving Language Understanding by Generative Pre-Training (GPT-1)
- Language Models are Unsupervised Multitask Learners (GPT-2)
- Language Models are Few-Shot Learners (GPT-3)
- Training Language Models to Follow Instructions with Human Feedback (InstructGPT)
- LLaMA: Open and Efficient Foundation Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG)
- LoRA: Low-Rank Adaptation of Large Language Models
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
01 Efficient Estimation of Word Representations in Vector Space (Word2Vec)
Learns dense word vectors via Skip-gram and CBOW objectives over huge text corpora.
The seed paper for distributed semantic representations — every embedding-based system traces back here.
Basic neural networks, softmax classification
Words appearing in similar contexts end up close in vector space, and you can do arithmetic on meaning.
02 GloVe: Global Vectors for Word Representation
Factorises the global word co-occurrence matrix to learn vectors that capture linear semantic structure.
The other half of the embedding canon — useful contrast to Word2Vec's local-context approach.
Word2Vec, matrix factorisation
Co-occurrence ratios, not raw counts, are what encode meaning.
03 Sequence to Sequence Learning with Neural Networks
Two stacked LSTMs encode a source sequence into a vector and decode the target sequence from it.
Establishes the encoder–decoder template that attention and Transformers later supercharged.
LSTMs, backprop through time
You can pose almost any structured prediction task as sequence-to-sequence.
04 Attention Is All You Need MVRP
Introduces the Transformer: an attention-only architecture with no recurrence or convolution.
The architectural blueprint underlying every modern LLM. Read it for the original framing.
Seq2seq with attention, matrix multiplication
Self-attention is fully parallelisable, captures long-range dependencies, and scales beautifully.
05 Deep Contextualized Word Representations (ELMo)
Trains a bidirectional LSTM language model and uses its hidden states as context-sensitive word features.
The first widely-used "context-aware" embedding, bridging static vectors and full pretraining.
Word2Vec, LSTMs
The same word should have different vectors in different sentences.
06 Improving Language Understanding by Generative Pre-Training (GPT-1) MVRP
Decoder-only Transformer pretrained as a language model, then fine-tuned for downstream tasks.
The origin point of the GPT line and of "pretrain then fine-tune" as the dominant NLP recipe.
Transformer architecture, cross-entropy loss
A single generative pretraining objective transfers to almost any NLP task.
07 BERT: Pre-training of Deep Bidirectional Transformers
Encoder-only Transformer pretrained with masked-language-modeling and next-sentence prediction.
Read against GPT-1 to understand the encoder-vs-decoder split that dominated 2018–2020 NLP.
Transformer architecture, GPT-1
Bidirectional context (mask-fill) makes a far stronger representation than left-to-right alone.
08 Language Models are Unsupervised Multitask Learners (GPT-2) MVRP
Scales GPT-1 to 1.5B parameters on a 40GB web corpus and shows surprising zero-shot abilities.
The first hint that scale alone unlocks new capabilities without task-specific training.
GPT-1, basic information theory
A big enough language model is implicitly a multi-task learner.
09 RoBERTa: A Robustly Optimized BERT Pretraining Approach
Trains BERT longer, on more data, with bigger batches and no NSP — and beats every BERT variant.
A masterclass in how much pretraining hyperparameters matter.
BERT
Many architectural "improvements" are really just under-trained baselines.
10 Language Models are Few-Shot Learners (GPT-3) MVRP
A 175B-parameter Transformer that performs many tasks via in-context examples, no gradient updates.
The paper that launched the modern era; in-context learning is the central abstraction of the LLM stack.
GPT-2, basic prompting
At sufficient scale, "show, don't train" becomes a viable interface to a model.
11 Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)
Casts every NLP task as text-to-text and pretrains a giant encoder-decoder on a clean web corpus (C4).
The clearest empirical study of pretraining design choices, plus the C4 dataset most LLMs still use.
BERT, seq2seq
Unify your tasks as text-to-text and almost everything becomes simpler.
12 Scaling Laws for Neural Language Models
Empirically derives smooth power-law relationships between loss, parameters, data, and compute.
Establishes that LLM progress is predictable and that you can budget compute analytically.
GPT-2, basic statistics
Loss falls as a power law in compute — and the exponent tells you exactly how to spend it.
13 Training Compute-Optimal Large Language Models (Chinchilla)
Re-runs scaling laws and shows existing models are massively under-trained relative to their size.
Corrects Kaplan's recipe; sets the modern data-to-parameters ratio used by Llama and beyond.
Kaplan scaling laws
For a fixed compute budget, you should scale parameters and tokens roughly equally.
14 PaLM: Scaling Language Modeling with Pathways
A 540B dense Transformer trained on Google's Pathways system, surfacing emergent reasoning capabilities.
The high-water mark for dense models pre-Chinchilla and a careful study of capability emergence.
GPT-3, scaling laws
Some abilities appear suddenly and only at scale — they aren't visible at smaller sizes.
30 Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Routes each token to a single expert FFN, training trillion-parameter sparse models with dense-model stability.
The paper that made mixture-of-experts practical at scale — read it before Mixtral or any modern MoE.
T5, scaling laws
Sparsity lets parameter count and compute cost scale independently.
31 Emergent Abilities of Large Language Models
Catalogues abilities that appear abruptly at scale rather than improving smoothly.
Frames the central scientific puzzle of scaling — and the debate over whether emergence is real or a metric artefact.
Scaling laws, GPT-3
Some capabilities are invisible at small scale, which makes extrapolation genuinely hard.
32 GPT-4 Technical Report
Reports GPT-4's capabilities, evals, and safety work — while withholding architecture and training details.
A milestone in capability and a turning point in (non-)openness; also introduces predictable-scaling loss forecasts.
GPT-3, InstructGPT
Frontier performance became predictable from small-scale runs — and frontier papers stopped disclosing how.
15 Training Language Models to Follow Instructions with Human Feedback (InstructGPT) MVRP
Fine-tunes GPT-3 with supervised demonstrations and PPO against a learned reward model (RLHF).
The recipe that turned a raw language model into a usable assistant — the real ChatGPT precursor.
GPT-3, basic RL (policy gradients)
A small amount of human feedback beats a huge amount of next-token training for usefulness.
16 Constitutional AI: Harmlessness from AI Feedback
Replaces human preference labels with model-generated critiques guided by a written "constitution".
A scalable alternative to RLHF that frames alignment as supervisable principles, not just preferences.
InstructGPT, RLHF
You can train safety using the model itself as the labeler — if the principles are explicit.
17 Direct Preference Optimization (DPO)
Derives a closed-form loss that fits a preference dataset directly, no reward model or PPO required.
Drastically simpler than RLHF; now the default for open-source preference tuning.
InstructGPT, KL divergence
Preference learning can be a one-stage classification problem rather than two-stage RL.
18 LLaMA: Open and Efficient Foundation Language Models MVRP
A family of 7B–65B Chinchilla-style models trained on public data, released to researchers.
The leak-and-release that catalyzed the modern open-weight LLM ecosystem and birthed the open-weight ecosystem.
GPT-3, Chinchilla
A well-trained 7B model can rival GPT-3 — open weights matter.
19 Llama 2: Open Foundation and Fine-Tuned Chat Models
Successor to LLaMA, with full RLHF chat variants, more pretraining data, and a permissive license.
The first commercially-usable competitive open model — a watershed for industry adoption.
LLaMA, InstructGPT
Open + commercial-friendly + RLHF'd turned out to be the unlock for adoption.
20 Mistral 7B
A 7B model with grouped-query and sliding-window attention that beats Llama 2 13B on most benchmarks.
Demonstrates how much performance is left on the table by architectural choices at small scale.
Llama 2, attention variants
Smarter attention beats bigger model size, dollar for dollar.
21 Mixtral of Experts
A sparse mixture-of-experts model where each token routes to 2 of 8 expert FFNs per layer.
The clearest example of MoE at production scale; shows where dense scaling is heading.
Mistral 7B, Switch Transformer
Activate only a fraction of parameters per token and you get bigger-model quality at smaller-model cost.
33 DeepSeek-V3 Technical Report
A 671B-parameter MoE (37B active) trained for ~$5.5M using MLA attention, FP8, and aggressive systems co-design.
Rewrote assumptions about frontier training costs; the strongest open-weight base model of its moment.
Mixtral, FlashAttention
Model architecture and training systems co-designed together beat either optimised alone.
22 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models MVRP
Adding "let's think step by step" exemplars to prompts massively improves multi-step reasoning.
The most influential prompting paper; reframes inference-time compute as a useful lever.
GPT-3, few-shot prompting
Letting the model show its work makes it correct more often — and only at scale.
23 ReAct: Synergizing Reasoning and Acting in Language Models
Interleaves chain-of-thought with tool calls, letting the model think and act in alternating steps.
The conceptual spine of every "agent" framework — read this before any agent paper.
Chain-of-Thought, basic API/tool concepts
Reasoning and acting are stronger together than either alone.
24 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG) MVRP
Combines a Transformer generator with a learned dense retriever over a non-parametric corpus.
The original RAG; the foundation of every "talk to your docs" system.
Transformers, basic retrieval (BM25 or DPR)
External memory beats trying to cram every fact into the weights.
25 Self-Consistency Improves Chain of Thought Reasoning
Sample many CoT trajectories and majority-vote the final answer, instead of greedy decoding.
A simple, free win on top of any CoT model — and the conceptual ancestor of test-time scaling.
Chain-of-Thought
Many noisy answers can be aggregated into one good answer.
34 Toolformer: Language Models Can Teach Themselves to Use Tools
Self-supervised annotation teaches an LM when and how to call APIs like calculators and search.
The cleanest early statement of tool use as a learned capability rather than a prompt hack.
GPT-3, ReAct
A model can bootstrap its own tool-use training data by checking which calls reduce its loss.
35 Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Generalises chain-of-thought into a search tree with lookahead, backtracking, and self-evaluation.
A conceptual bridge from prompting tricks to genuine inference-time search.
Chain-of-Thought, Self-Consistency
Reasoning improves when the model can explore and prune alternatives, not just sample forward.
36 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning MVRP
Pure RL on verifiable rewards elicits long chain-of-thought reasoning, distilled into open models.
The open counterpart to o1-style reasoning models — made test-time-compute training reproducible.
Chain-of-Thought, PPO/GRPO, DeepSeek-V3
Reasoning behaviour can emerge from RL against checkable answers, without human reasoning traces.
26 LoRA: Low-Rank Adaptation of Large Language Models MVRP
Freezes the base model and trains low-rank update matrices added to its weights.
The technique that made fine-tuning huge models accessible to anyone with one GPU.
Linear algebra, fine-tuning basics
Most fine-tuning lives in a low-rank subspace; you don't need to touch every weight.
27 QLoRA: Efficient Finetuning of Quantized LLMs
Combines 4-bit base model quantization with LoRA adapters and paged optimizers.
Brings 65B-parameter fine-tuning to a single consumer GPU.
LoRA, quantization basics
4-bit + low-rank is enough to match full-precision fine-tuning quality.
28 FlashAttention: Fast and Memory-Efficient Exact Attention
A tiling-based, IO-aware attention kernel that reduces memory reads/writes without changing the math.
The kernel that's now under almost every Transformer in production.
Transformer attention, GPU memory hierarchy
Optimising for memory bandwidth, not FLOPs, is the right move on modern GPUs.
29 The Llama 3 Herd of Models MVRP
A detailed engineering report on training Llama 3 (8B, 70B, 405B) on 15T tokens.
The most thorough open description of a frontier-grade training run — the new state of the art for open weights.
Llama 2, Chinchilla, modern alignment
Frontier models are mostly an engineering problem now: data, infra, and post-training quality.
37 Mamba: Linear-Time Sequence Modeling with Selective State Spaces
A selective state-space model matching Transformer quality with linear-time, constant-memory inference.
The strongest post-Transformer architecture candidate; defines the attention-alternative research line.
Transformer, state-space models
Input-dependent state selection recovers the content-based routing that made attention win.