Papers in Order

Computer Vision

From hand-engineered convnets on MNIST to general-purpose vision foundation models — the network designs, the loss functions, and the diffusion processes that built modern visual AI.

32 papers 8 levels Included papers 1998 – 2023 MVRP 6 papers
0 of 32 papers read. Progress stays in this browser.

Minimum viable reading path

The 6 papers that give you most of the field's mental model, in reading order.

  1. ImageNet Classification with Deep Convolutional Neural Networks (AlexNet)
  2. Deep Residual Learning for Image Recognition (ResNet)
  3. U-Net: Convolutional Networks for Biomedical Image Segmentation
  4. Learning Transferable Visual Models From Natural Language Supervision (CLIP)
  5. Denoising Diffusion Probabilistic Models (DDPM)
  6. Segment Anything (SAM)
Level 0 Foundations
01 Gradient-Based Learning Applied to Document Recognition (LeNet)
TL;DR

A 7-layer convolutional network reads handwritten digits with weight sharing, pooling, and end-to-end training.

Why read this

The blueprint for every modern CNN; everything that follows is a refinement of these ideas.

Prerequisites

Backpropagation, basic image processing

Key takeaway

Convolution + pooling + non-linearities is the right inductive bias for images.

Read the paper
02 ImageNet Classification with Deep Convolutional Neural Networks (AlexNet) MVRP
TL;DR

A deep CNN with ReLU, dropout, and dual-GPU training crushes ImageNet 2012 by a 10-point margin.

Why read this

The landmark result that accelerated deep learning adoption in vision — the moment deep learning took over computer vision.

Prerequisites

LeNet, basic CUDA awareness

Key takeaway

Big data + GPUs + ReLU + dropout was the unlock; the architecture was the easy part.

Read the paper
Level 1 Backbones & training tricks
03 Very Deep Convolutional Networks for Large-Scale Image Recognition (VGG)
TL;DR

Stacks 3×3 convolutions to depths of 16 and 19 layers, showing depth itself is a strong prior.

Why read this

The cleanest demonstration that depth matters and a perennial backbone for transfer learning.

Prerequisites

AlexNet, basic CNN math

Key takeaway

Small kernels stacked deep beat large kernels stacked shallow.

Read the paper
04 Going Deeper with Convolutions (GoogLeNet / Inception)
TL;DR

Introduces the Inception module, mixing parallel kernels of multiple sizes within each block.

Why read this

A counterpoint to VGG showing that smarter blocks beat brute-force depth.

Prerequisites

VGG, computational complexity of conv

Key takeaway

Multi-scale features can live inside a single layer.

Read the paper
05 Deep Residual Learning for Image Recognition (ResNet) MVRP
TL;DR

Adds identity skip connections so the network only has to learn residuals; trains 152-layer nets.

Why read this

The single most important architectural idea after the convolution itself; appears in every modern net.

Prerequisites

VGG, vanishing gradients

Key takeaway

If you can't train it deeper, give the gradient a shortcut.

Read the paper
06 Densely Connected Convolutional Networks (DenseNet)
TL;DR

Each layer receives inputs from every preceding layer, encouraging feature reuse.

Why read this

Useful contrast to ResNet; clarifies what skip connections are actually doing.

Prerequisites

ResNet

Key takeaway

Feature reuse, not just gradient flow, explains why skip connections help.

Read the paper
07 Batch Normalization: Accelerating Deep Network Training
TL;DR

Normalises layer activations within each mini-batch, dramatically stabilising deep network training.

Why read this

Pretty much every CNN since uses BN; understanding its dynamics is foundational.

Prerequisites

CNN training basics

Key takeaway

Stabilising activation distributions during training is half the battle.

Read the paper
Level 2 Object detection
08 Rich Feature Hierarchies for Accurate Object Detection (R-CNN)
TL;DR

Region proposals fed into a CNN and an SVM produce object detections.

Why read this

The starting point of the modern detection pipeline.

Prerequisites

AlexNet, selective search

Key takeaway

You can repurpose a classifier into a detector by feeding it proposals.

Read the paper
09 Fast R-CNN
TL;DR

Reuses one CNN forward pass for all proposals via a region-of-interest pooling layer.

Why read this

A clean lesson in how one architectural idea (RoI pooling) can collapse a whole pipeline.

Prerequisites

R-CNN

Key takeaway

Compute features once, then crop them — don't re-run the CNN per proposal.

Read the paper
10 Faster R-CNN: Towards Real-Time Detection with Region Proposal Networks
TL;DR

Replaces external proposals with a learned region proposal network sharing the detector's features.

Why read this

The end-to-end detection paper; basis of two-stage detectors for years.

Prerequisites

Fast R-CNN

Key takeaway

Make every component of the pipeline learnable.

Read the paper
11 You Only Look Once (YOLO): Unified, Real-Time Object Detection
TL;DR

Treats detection as a single regression problem from image to bounding boxes and classes.

Why read this

The other branch of the detection family — single-stage, fast, deployable.

Prerequisites

CNN basics

Key takeaway

Sometimes you should just predict the answer directly.

Read the paper
12 Focal Loss for Dense Object Detection (RetinaNet)
TL;DR

A modified cross-entropy that down-weights easy examples, fixing class imbalance in dense detectors.

Why read this

A self-contained lesson in loss design; the focal loss now lives in many other domains too.

Prerequisites

YOLO, cross-entropy

Key takeaway

When most examples are trivial, weight by difficulty.

Read the paper
Level 3 Segmentation
13 Fully Convolutional Networks for Semantic Segmentation
TL;DR

Replaces the final FC layers of a classification CNN with conv layers to produce per-pixel labels.

Why read this

The simple, beautiful idea that turned classifiers into segmenters.

Prerequisites

VGG, transposed conv

Key takeaway

A classifier without dense layers is a segmenter.

Read the paper
14 U-Net: Convolutional Networks for Biomedical Image Segmentation MVRP
TL;DR

A symmetric encoder–decoder with skip connections that fuses low-level detail and high-level context.

Why read this

Far more than a medical paper — U-Net is the backbone of modern diffusion models too.

Prerequisites

FCN

Key takeaway

Symmetric encoder–decoder with skips is a remarkably general architecture.

Read the paper
15 Mask R-CNN
TL;DR

Adds a small fully-conv mask branch to Faster R-CNN, producing instance segmentation alongside detection.

Why read this

Shows how a tiny architectural addition unlocks a whole new task.

Prerequisites

Faster R-CNN

Key takeaway

Multi-task heads atop a shared backbone are absurdly effective.

Read the paper
Level 4 Generative models
16 Generative Adversarial Networks
TL;DR

A generator and discriminator play a minimax game; samples come from the generator at convergence.

Why read this

A foundational generative-modeling idea; understanding GANs sharpens your intuition for diffusion.

Prerequisites

Game theory basics, CNNs

Key takeaway

You can train a generator without a likelihood by pitting it against a learned critic.

Read the paper
17 Unsupervised Representation Learning with DCGAN
TL;DR

A set of architectural guidelines (strided conv, batch norm, no FC) that finally make GANs train stably.

Why read this

The first pretty GAN samples and a study in how much architectural detail matters in adversarial training.

Prerequisites

GANs

Key takeaway

GAN training is mostly a stability problem.

Read the paper
18 A Style-Based Generator Architecture for GANs (StyleGAN)
TL;DR

Disentangles style and content via adaptive instance norm and a learned latent W space.

Why read this

The pinnacle of GAN-era image quality and a beautiful study in disentanglement.

Prerequisites

DCGAN

Key takeaway

Architectural inductive biases can produce semantically meaningful latent spaces for free.

Read the paper
Level 5 Vision Transformers
19 An Image is Worth 16×16 Words: Transformers for Image Recognition (ViT)
TL;DR

Splits an image into patches, treats them as tokens, and feeds them into a vanilla Transformer encoder.

Why read this

The convergence point of NLP and vision — proves convolutions aren't architecturally necessary.

Prerequisites

Transformer, ResNet

Key takeaway

With enough data, attention beats convolution at images.

Read the paper
20 End-to-End Object Detection with Transformers (DETR)
TL;DR

Detection as set prediction: a Transformer outputs a fixed set of boxes, matched via Hungarian loss.

Why read this

A radical re-framing of detection that removed anchors, NMS, and most heuristics.

Prerequisites

Transformer, Faster R-CNN

Key takeaway

Many "necessary" hand-tuned components turn out to be replaceable by attention + the right loss.

Read the paper
21 Swin Transformer: Hierarchical Vision Transformer with Shifted Windows
TL;DR

Restricts self-attention to local shifted windows, restoring CNN-like multi-scale efficiency.

Why read this

The bridge between ViT's purity and a CNN's practical efficiency.

Prerequisites

ViT

Key takeaway

Locality and hierarchy can be reintroduced into Transformers when compute is tight.

Read the paper
29 Masked Autoencoders Are Scalable Vision Learners (MAE)
TL;DR

Masks 75% of image patches and reconstructs them with an asymmetric encoder-decoder ViT.

Why read this

Brought BERT-style masked pretraining to vision — now a default self-supervised recipe.

Prerequisites

ViT, BERT

Key takeaway

Images are redundant enough that aggressive masking makes a great pretext task.

Read the paper
30 A ConvNet for the 2020s (ConvNeXt)
TL;DR

Modernises a ResNet step-by-step with ViT-era training tricks until it matches Swin Transformer.

Why read this

A careful ablation showing how much of the ViT gap was training recipe, not architecture.

Prerequisites

ResNet, Swin Transformer

Key takeaway

Attribute gains carefully — recipes and architectures are easy to conflate.

Read the paper
Level 6 Multimodal
22 Learning Transferable Visual Models From Natural Language Supervision (CLIP) MVRP
TL;DR

Trains image and text encoders contrastively on 400M web image-caption pairs.

Why read this

The bedrock of modern multimodal AI; underlies retrieval, classification, and image-generation guidance.

Prerequisites

Transformers, contrastive learning

Key takeaway

Natural language is a far more general supervision signal than fixed label sets.

Read the paper
23 Zero-Shot Text-to-Image Generation (DALL-E)
TL;DR

A 12B-parameter Transformer trained autoregressively over discrete image-token + text-token sequences.

Why read this

The first compelling text-to-image system and a clean conceptual ancestor of all diffusion T2I.

Prerequisites

GPT-3, VQ-VAE

Key takeaway

You can model images as discrete tokens and generate them with a language model.

Read the paper
Level 7 Diffusion & foundation models
24 Denoising Diffusion Probabilistic Models (DDPM) MVRP
TL;DR

Trains a network to reverse a fixed Gaussian noising process, generating images by iterative denoising.

Why read this

The paper that revived and re-formalised diffusion models. Read this before any T2I paper.

Prerequisites

GANs, basic stochastic processes

Key takeaway

Reversing noise step-by-step is a simpler, more stable generative recipe than adversarial training.

Read the paper
25 High-Resolution Image Synthesis with Latent Diffusion (Stable Diffusion)
TL;DR

Runs the diffusion process in a compressed VAE latent space, slashing compute by an order of magnitude.

Why read this

The paper that put open-source text-to-image into millions of hands.

Prerequisites

DDPM, VAE

Key takeaway

Don't diffuse pixels — diffuse a small, perceptually-equivalent latent representation.

Read the paper
26 Photorealistic Text-to-Image Diffusion Models (Imagen)
TL;DR

Pairs a frozen large language model text encoder with a cascaded pixel-space diffusion decoder.

Why read this

A clean, practical study showing that text-encoder quality dominates image-decoder size.

Prerequisites

DDPM, CLIP

Key takeaway

A bigger text encoder beats a bigger image decoder.

Read the paper
27 Segment Anything (SAM) MVRP
TL;DR

A promptable segmentation foundation model trained on 1B masks across 11M images.

Why read this

The "major breakthrough moment" for segmentation — a single model that segments anything from any prompt.

Prerequisites

ViT, segmentation basics

Key takeaway

Foundation models work in vision too — given the right task and a billion labels.

Read the paper
28 DINOv2: Learning Robust Visual Features without Supervision
TL;DR

A self-supervised ViT trained on a curated 142M image dataset producing all-purpose features.

Why read this

The strongest pure-vision foundation model — a counterpoint to CLIP's language-supervised approach.

Prerequisites

ViT, self-supervised learning

Key takeaway

You can build vision foundation models without text — careful data and self-supervision suffice.

Read the paper
31 NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
TL;DR

Encodes a scene as an MLP mapping 5D coordinates to colour and density, rendered by ray marching.

Why read this

Founded neural 3D representation — a whole subfield traces to this paper.

Prerequisites

Volume rendering basics, MLPs

Key takeaway

A neural network can be the scene representation itself, not just a processor of one.

Read the paper
32 3D Gaussian Splatting for Real-Time Radiance Field Rendering
TL;DR

Represents scenes as millions of optimised anisotropic 3D Gaussians, rasterised in real time.

Why read this

Displaced NeRF for many practical uses within a year — a lesson in explicit vs implicit representations.

Prerequisites

NeRF, rasterisation basics

Key takeaway

An explicit, differentiable primitive can beat an implicit network on both speed and quality.

Read the paper