Papers in Order

Computational Biology

From dynamic-programming alignment to AlphaFold and protein design — how algorithms went from finding motifs in DNA to predicting and engineering biological structure.

20 papers 6 levels Included papers 1970 – 2024 MVRP 5 papers
0 of 20 papers read. Progress stays in this browser.

Minimum viable reading path

The 5 papers that give you most of the field's mental model, in reading order.

  1. A General Method Applicable to the Search for Similarities (Needleman–Wunsch)
  2. Basic Local Alignment Search Tool (BLAST)
  3. Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform (BWA)
  4. Highly Accurate Protein Structure Prediction with AlphaFold
  5. Robust Deep Learning–Based Protein Sequence Design Using ProteinMPNN
Level 0 Sequence alignment
01 A General Method Applicable to the Search for Similarities (Needleman–Wunsch) MVRP
TL;DR

A dynamic-programming algorithm for global alignment of two biological sequences.

Why read this

The first proper algorithm of computational biology; conceptual ancestor of every aligner.

Prerequisites

Dynamic programming basics

Key takeaway

Optimal sequence alignment is solvable in O(mn) time with a single DP table.

Read the paper
02 Identification of Common Molecular Subsequences (Smith–Waterman)
TL;DR

Adapts Needleman–Wunsch to local alignment by allowing zero scores in the DP table.

Why read this

A small modification with huge consequences for finding conserved motifs and homology.

Prerequisites

Needleman–Wunsch

Key takeaway

Local alignment is just global alignment with a non-negativity constraint.

Read the paper
03 Basic Local Alignment Search Tool (BLAST) MVRP
TL;DR

A heuristic algorithm trading some sensitivity for orders-of-magnitude speed over Smith–Waterman.

Why read this

The most-used algorithm in all of biology; the workhorse of every database search.

Prerequisites

Smith–Waterman, basic indexing

Key takeaway

A statistically-grounded heuristic can be the right answer when exact is too slow.

Read the paper
Level 1 Probabilistic models
04 Profile HMMs for Biological Sequence Analysis (HMMER)
TL;DR

Builds a hidden Markov model per protein family for sensitive remote-homology searches.

Why read this

The standard probabilistic framework for protein family detection.

Prerequisites

BLAST, HMM basics

Key takeaway

A probabilistic model of a family beats a single best representative.

Read the paper
Level 2 Sequencing & assembly
05 Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform (BWA) MVRP
TL;DR

Indexes the reference genome via the Burrows–Wheeler transform for fast NGS read alignment.

Why read this

One of the most-used aligners in modern genomics — the BWT made it possible.

Prerequisites

BLAST, suffix arrays

Key takeaway

A clever string-data structure makes a previously intractable problem trivial.

Read the paper
06 Bowtie: Ultrafast and Memory-Efficient Alignment of Short DNA Sequences
TL;DR

Another BWT-based aligner with a focus on memory footprint and speed.

Why read this

A useful contrast to BWA; both papers together define the modern read-alignment toolbox.

Prerequisites

BWA, suffix arrays

Key takeaway

There's often more than one good way to apply a powerful new data structure.

Read the paper
07 Velvet: Algorithms for De Novo Short Read Assembly
TL;DR

Uses de Bruijn graphs to assemble genomes from millions of short sequencing reads.

Why read this

The de Bruijn graph idea is the conceptual core of every modern assembler.

Prerequisites

Graph algorithms, basic genomics

Key takeaway

Genome assembly is an Eulerian path problem on a graph of k-mers.

Read the paper
08 STAR: Ultrafast Universal RNA-seq Aligner
TL;DR

A spliced aligner that handles RNA-seq reads spanning exon-exon junctions efficiently.

Why read this

The de facto standard RNA-seq aligner — read it before any expression analysis paper.

Prerequisites

BWA, basic transcriptomics

Key takeaway

RNA-seq needs splice-aware alignment, not just genome-aware alignment.

Read the paper
Level 3 Functional & single-cell analysis
09 Moderated Estimation of Fold Change and Dispersion for RNA-Seq (DESeq2)
TL;DR

A statistical framework for differential expression with shrinkage estimators for low-count noise.

Why read this

The most-cited paper in modern transcriptomics — methodologically influential beyond its field.

Prerequisites

Negative binomial distribution, basic GLMs

Key takeaway

Empirical-Bayes shrinkage gives stable inferences from noisy count data.

Read the paper
10 UMAP: Uniform Manifold Approximation and Projection
TL;DR

A dimensionality-reduction algorithm with a manifold-learning theoretical foundation.

Why read this

Quietly displaced t-SNE in single-cell biology and elsewhere — a major workflow shift.

Prerequisites

PCA, basic topology useful

Key takeaway

Better assumptions about the data manifold give better low-dimensional embeddings.

Read the paper
11 Comprehensive Integration of Single-Cell Data (Seurat v3)
TL;DR

A statistical framework for integrating single-cell datasets across conditions and modalities.

Why read this

The standard analysis pipeline for one of the most-active areas of biology.

Prerequisites

PCA, basic single-cell biology

Key takeaway

Anchor-based integration lets you compare cells across very different experiments.

Read the paper
12 Massively Parallel Digital Transcriptional Profiling of Single Cells (Cell Ranger / 10x)
TL;DR

Demonstrates droplet-based single-cell RNA sequencing at hundreds of thousands of cells per run.

Why read this

The technology paper that made single-cell biology a mainstream field.

Prerequisites

RNA-seq basics

Key takeaway

Massive throughput on cells, not just reads, opens new biological questions.

Read the paper
Level 4 Protein structure
13 Highly Accurate Protein Structure Prediction with AlphaFold MVRP
TL;DR

A geometry-aware Transformer architecture (Evoformer + structure module) predicting 3D protein structure to atomic accuracy.

Why read this

A scientific landmark of historic magnitude; the protein-folding problem largely solved.

Prerequisites

Transformers, basic structural biology

Key takeaway

Hard scientific problems crack when you bake the right geometric and biological priors into the model.

Read the paper
14 Evolutionary-Scale Prediction of Atomic-Level Protein Structure (ESMFold)
TL;DR

Uses a 15B-parameter protein language model (ESM-2) to predict structure directly from sequence.

Why read this

Faster, simpler alternative to AlphaFold for many use cases; demonstrates the power of protein LLMs.

Prerequisites

AlphaFold, protein language models

Key takeaway

A big enough language model implicitly learns structure from sequence alone.

Read the paper
15 Accurate Prediction of Protein Structures and Interactions Using a 3-Track Network (RoseTTAFold)
TL;DR

A three-track network reasoning over sequence, distance, and structure simultaneously.

Why read this

A useful counterpoint to AlphaFold; same problem, related but distinct architecture.

Prerequisites

AlphaFold, basic graph nets

Key takeaway

Structure prediction benefits from reasoning over multiple representations in parallel.

Read the paper
Level 5 Generative biology & frontier
16 Robust Deep Learning–Based Protein Sequence Design Using ProteinMPNN MVRP
TL;DR

A graph neural network that designs protein sequences to fold into a desired backbone.

Why read this

A modern reference for inverse folding and the design half of computational protein engineering.

Prerequisites

AlphaFold, GNNs

Key takeaway

Inverse design (give me a sequence for this shape) is the dual problem to folding.

Read the paper
17 De Novo Design of Protein Structure and Function with RFdiffusion
TL;DR

A diffusion model that generates novel protein backbones for specified geometric constraints.

Why read this

The "Stable Diffusion of biology" — generative protein design now actually works.

Prerequisites

Diffusion models, protein design basics

Key takeaway

Diffusion is a remarkably general way to generate structured biological objects.

Read the paper
18 Transfer Learning Enables Predictions in Network Biology (Geneformer)
TL;DR

A Transformer pretrained on 30M single-cell transcriptomes and fine-tuned for downstream biology.

Why read this

Foundation-model thinking finally hitting cell biology.

Prerequisites

Transformers, single-cell RNA-seq

Key takeaway

Pretraining on rank-encoded gene expression yields strong biological transferability.

Read the paper
19 Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3
TL;DR

Extends AlphaFold to predict protein–nucleic-acid, protein–ligand, and protein–protein complexes.

Why read this

The current state of the art; redefines what computational structural biology can do.

Prerequisites

AlphaFold, diffusion models

Key takeaway

A single foundation model can handle proteins, RNA, DNA, and small-molecule interactions.

Read the paper
20 Language Models Generalize Beyond Natural Proteins
TL;DR

A protein language model designs novel functional proteins beyond known evolutionary distributions.

Why read this

A glimpse of foundation-model-driven exploration of biological design space.

Prerequisites

ESM, protein design basics

Key takeaway

Language models trained on natural proteins can imagine functional unnatural ones.

Read the paper