Computer Vision
From hand-engineered convnets on MNIST to general-purpose vision foundation models — the network designs, the loss functions, and the diffusion processes that built modern visual AI.
Minimum viable reading path
The 6 papers that give you most of the field's mental model, in reading order.
- ImageNet Classification with Deep Convolutional Neural Networks (AlexNet)
- Deep Residual Learning for Image Recognition (ResNet)
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Learning Transferable Visual Models From Natural Language Supervision (CLIP)
- Denoising Diffusion Probabilistic Models (DDPM)
- Segment Anything (SAM)
01 Gradient-Based Learning Applied to Document Recognition (LeNet)
A 7-layer convolutional network reads handwritten digits with weight sharing, pooling, and end-to-end training.
The blueprint for every modern CNN; everything that follows is a refinement of these ideas.
Backpropagation, basic image processing
Convolution + pooling + non-linearities is the right inductive bias for images.
02 ImageNet Classification with Deep Convolutional Neural Networks (AlexNet) MVRP
A deep CNN with ReLU, dropout, and dual-GPU training crushes ImageNet 2012 by a 10-point margin.
The landmark result that accelerated deep learning adoption in vision — the moment deep learning took over computer vision.
LeNet, basic CUDA awareness
Big data + GPUs + ReLU + dropout was the unlock; the architecture was the easy part.
03 Very Deep Convolutional Networks for Large-Scale Image Recognition (VGG)
Stacks 3×3 convolutions to depths of 16 and 19 layers, showing depth itself is a strong prior.
The cleanest demonstration that depth matters and a perennial backbone for transfer learning.
AlexNet, basic CNN math
Small kernels stacked deep beat large kernels stacked shallow.
04 Going Deeper with Convolutions (GoogLeNet / Inception)
Introduces the Inception module, mixing parallel kernels of multiple sizes within each block.
A counterpoint to VGG showing that smarter blocks beat brute-force depth.
VGG, computational complexity of conv
Multi-scale features can live inside a single layer.
05 Deep Residual Learning for Image Recognition (ResNet) MVRP
Adds identity skip connections so the network only has to learn residuals; trains 152-layer nets.
The single most important architectural idea after the convolution itself; appears in every modern net.
VGG, vanishing gradients
If you can't train it deeper, give the gradient a shortcut.
06 Densely Connected Convolutional Networks (DenseNet)
Each layer receives inputs from every preceding layer, encouraging feature reuse.
Useful contrast to ResNet; clarifies what skip connections are actually doing.
ResNet
Feature reuse, not just gradient flow, explains why skip connections help.
07 Batch Normalization: Accelerating Deep Network Training
Normalises layer activations within each mini-batch, dramatically stabilising deep network training.
Pretty much every CNN since uses BN; understanding its dynamics is foundational.
CNN training basics
Stabilising activation distributions during training is half the battle.
08 Rich Feature Hierarchies for Accurate Object Detection (R-CNN)
Region proposals fed into a CNN and an SVM produce object detections.
The starting point of the modern detection pipeline.
AlexNet, selective search
You can repurpose a classifier into a detector by feeding it proposals.
09 Fast R-CNN
Reuses one CNN forward pass for all proposals via a region-of-interest pooling layer.
A clean lesson in how one architectural idea (RoI pooling) can collapse a whole pipeline.
R-CNN
Compute features once, then crop them — don't re-run the CNN per proposal.
10 Faster R-CNN: Towards Real-Time Detection with Region Proposal Networks
Replaces external proposals with a learned region proposal network sharing the detector's features.
The end-to-end detection paper; basis of two-stage detectors for years.
Fast R-CNN
Make every component of the pipeline learnable.
11 You Only Look Once (YOLO): Unified, Real-Time Object Detection
Treats detection as a single regression problem from image to bounding boxes and classes.
The other branch of the detection family — single-stage, fast, deployable.
CNN basics
Sometimes you should just predict the answer directly.
12 Focal Loss for Dense Object Detection (RetinaNet)
A modified cross-entropy that down-weights easy examples, fixing class imbalance in dense detectors.
A self-contained lesson in loss design; the focal loss now lives in many other domains too.
YOLO, cross-entropy
When most examples are trivial, weight by difficulty.
13 Fully Convolutional Networks for Semantic Segmentation
Replaces the final FC layers of a classification CNN with conv layers to produce per-pixel labels.
The simple, beautiful idea that turned classifiers into segmenters.
VGG, transposed conv
A classifier without dense layers is a segmenter.
14 U-Net: Convolutional Networks for Biomedical Image Segmentation MVRP
A symmetric encoder–decoder with skip connections that fuses low-level detail and high-level context.
Far more than a medical paper — U-Net is the backbone of modern diffusion models too.
FCN
Symmetric encoder–decoder with skips is a remarkably general architecture.
15 Mask R-CNN
Adds a small fully-conv mask branch to Faster R-CNN, producing instance segmentation alongside detection.
Shows how a tiny architectural addition unlocks a whole new task.
Faster R-CNN
Multi-task heads atop a shared backbone are absurdly effective.
16 Generative Adversarial Networks
A generator and discriminator play a minimax game; samples come from the generator at convergence.
A foundational generative-modeling idea; understanding GANs sharpens your intuition for diffusion.
Game theory basics, CNNs
You can train a generator without a likelihood by pitting it against a learned critic.
17 Unsupervised Representation Learning with DCGAN
A set of architectural guidelines (strided conv, batch norm, no FC) that finally make GANs train stably.
The first pretty GAN samples and a study in how much architectural detail matters in adversarial training.
GANs
GAN training is mostly a stability problem.
18 A Style-Based Generator Architecture for GANs (StyleGAN)
Disentangles style and content via adaptive instance norm and a learned latent W space.
The pinnacle of GAN-era image quality and a beautiful study in disentanglement.
DCGAN
Architectural inductive biases can produce semantically meaningful latent spaces for free.
19 An Image is Worth 16×16 Words: Transformers for Image Recognition (ViT)
Splits an image into patches, treats them as tokens, and feeds them into a vanilla Transformer encoder.
The convergence point of NLP and vision — proves convolutions aren't architecturally necessary.
Transformer, ResNet
With enough data, attention beats convolution at images.
20 End-to-End Object Detection with Transformers (DETR)
Detection as set prediction: a Transformer outputs a fixed set of boxes, matched via Hungarian loss.
A radical re-framing of detection that removed anchors, NMS, and most heuristics.
Transformer, Faster R-CNN
Many "necessary" hand-tuned components turn out to be replaceable by attention + the right loss.
21 Swin Transformer: Hierarchical Vision Transformer with Shifted Windows
Restricts self-attention to local shifted windows, restoring CNN-like multi-scale efficiency.
The bridge between ViT's purity and a CNN's practical efficiency.
ViT
Locality and hierarchy can be reintroduced into Transformers when compute is tight.
29 Masked Autoencoders Are Scalable Vision Learners (MAE)
Masks 75% of image patches and reconstructs them with an asymmetric encoder-decoder ViT.
Brought BERT-style masked pretraining to vision — now a default self-supervised recipe.
ViT, BERT
Images are redundant enough that aggressive masking makes a great pretext task.
30 A ConvNet for the 2020s (ConvNeXt)
Modernises a ResNet step-by-step with ViT-era training tricks until it matches Swin Transformer.
A careful ablation showing how much of the ViT gap was training recipe, not architecture.
ResNet, Swin Transformer
Attribute gains carefully — recipes and architectures are easy to conflate.
22 Learning Transferable Visual Models From Natural Language Supervision (CLIP) MVRP
Trains image and text encoders contrastively on 400M web image-caption pairs.
The bedrock of modern multimodal AI; underlies retrieval, classification, and image-generation guidance.
Transformers, contrastive learning
Natural language is a far more general supervision signal than fixed label sets.
23 Zero-Shot Text-to-Image Generation (DALL-E)
A 12B-parameter Transformer trained autoregressively over discrete image-token + text-token sequences.
The first compelling text-to-image system and a clean conceptual ancestor of all diffusion T2I.
GPT-3, VQ-VAE
You can model images as discrete tokens and generate them with a language model.
24 Denoising Diffusion Probabilistic Models (DDPM) MVRP
Trains a network to reverse a fixed Gaussian noising process, generating images by iterative denoising.
The paper that revived and re-formalised diffusion models. Read this before any T2I paper.
GANs, basic stochastic processes
Reversing noise step-by-step is a simpler, more stable generative recipe than adversarial training.
25 High-Resolution Image Synthesis with Latent Diffusion (Stable Diffusion)
Runs the diffusion process in a compressed VAE latent space, slashing compute by an order of magnitude.
The paper that put open-source text-to-image into millions of hands.
DDPM, VAE
Don't diffuse pixels — diffuse a small, perceptually-equivalent latent representation.
26 Photorealistic Text-to-Image Diffusion Models (Imagen)
Pairs a frozen large language model text encoder with a cascaded pixel-space diffusion decoder.
A clean, practical study showing that text-encoder quality dominates image-decoder size.
DDPM, CLIP
A bigger text encoder beats a bigger image decoder.
27 Segment Anything (SAM) MVRP
A promptable segmentation foundation model trained on 1B masks across 11M images.
The "major breakthrough moment" for segmentation — a single model that segments anything from any prompt.
ViT, segmentation basics
Foundation models work in vision too — given the right task and a billion labels.
28 DINOv2: Learning Robust Visual Features without Supervision
A self-supervised ViT trained on a curated 142M image dataset producing all-purpose features.
The strongest pure-vision foundation model — a counterpoint to CLIP's language-supervised approach.
ViT, self-supervised learning
You can build vision foundation models without text — careful data and self-supervision suffice.
31 NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
Encodes a scene as an MLP mapping 5D coordinates to colour and density, rendered by ray marching.
Founded neural 3D representation — a whole subfield traces to this paper.
Volume rendering basics, MLPs
A neural network can be the scene representation itself, not just a processor of one.
32 3D Gaussian Splatting for Real-Time Radiance Field Rendering
Represents scenes as millions of optimised anisotropic 3D Gaussians, rasterised in real time.
Displaced NeRF for many practical uses within a year — a lesson in explicit vs implicit representations.
NeRF, rasterisation basics
An explicit, differentiable primitive can beat an implicit network on both speed and quality.