Computational Biology
From dynamic-programming alignment to AlphaFold and protein design — how algorithms went from finding motifs in DNA to predicting and engineering biological structure.
Minimum viable reading path
The 5 papers that give you most of the field's mental model, in reading order.
- A General Method Applicable to the Search for Similarities (Needleman–Wunsch)
- Basic Local Alignment Search Tool (BLAST)
- Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform (BWA)
- Highly Accurate Protein Structure Prediction with AlphaFold
- Robust Deep Learning–Based Protein Sequence Design Using ProteinMPNN
01 A General Method Applicable to the Search for Similarities (Needleman–Wunsch) MVRP
A dynamic-programming algorithm for global alignment of two biological sequences.
The first proper algorithm of computational biology; conceptual ancestor of every aligner.
Dynamic programming basics
Optimal sequence alignment is solvable in O(mn) time with a single DP table.
02 Identification of Common Molecular Subsequences (Smith–Waterman)
Adapts Needleman–Wunsch to local alignment by allowing zero scores in the DP table.
A small modification with huge consequences for finding conserved motifs and homology.
Needleman–Wunsch
Local alignment is just global alignment with a non-negativity constraint.
03 Basic Local Alignment Search Tool (BLAST) MVRP
A heuristic algorithm trading some sensitivity for orders-of-magnitude speed over Smith–Waterman.
The most-used algorithm in all of biology; the workhorse of every database search.
Smith–Waterman, basic indexing
A statistically-grounded heuristic can be the right answer when exact is too slow.
04 Profile HMMs for Biological Sequence Analysis (HMMER)
Builds a hidden Markov model per protein family for sensitive remote-homology searches.
The standard probabilistic framework for protein family detection.
BLAST, HMM basics
A probabilistic model of a family beats a single best representative.
05 Fast and Accurate Short Read Alignment with Burrows–Wheeler Transform (BWA) MVRP
Indexes the reference genome via the Burrows–Wheeler transform for fast NGS read alignment.
One of the most-used aligners in modern genomics — the BWT made it possible.
BLAST, suffix arrays
A clever string-data structure makes a previously intractable problem trivial.
06 Bowtie: Ultrafast and Memory-Efficient Alignment of Short DNA Sequences
Another BWT-based aligner with a focus on memory footprint and speed.
A useful contrast to BWA; both papers together define the modern read-alignment toolbox.
BWA, suffix arrays
There's often more than one good way to apply a powerful new data structure.
07 Velvet: Algorithms for De Novo Short Read Assembly
Uses de Bruijn graphs to assemble genomes from millions of short sequencing reads.
The de Bruijn graph idea is the conceptual core of every modern assembler.
Graph algorithms, basic genomics
Genome assembly is an Eulerian path problem on a graph of k-mers.
08 STAR: Ultrafast Universal RNA-seq Aligner
A spliced aligner that handles RNA-seq reads spanning exon-exon junctions efficiently.
The de facto standard RNA-seq aligner — read it before any expression analysis paper.
BWA, basic transcriptomics
RNA-seq needs splice-aware alignment, not just genome-aware alignment.
09 Moderated Estimation of Fold Change and Dispersion for RNA-Seq (DESeq2)
A statistical framework for differential expression with shrinkage estimators for low-count noise.
The most-cited paper in modern transcriptomics — methodologically influential beyond its field.
Negative binomial distribution, basic GLMs
Empirical-Bayes shrinkage gives stable inferences from noisy count data.
10 UMAP: Uniform Manifold Approximation and Projection
A dimensionality-reduction algorithm with a manifold-learning theoretical foundation.
Quietly displaced t-SNE in single-cell biology and elsewhere — a major workflow shift.
PCA, basic topology useful
Better assumptions about the data manifold give better low-dimensional embeddings.
11 Comprehensive Integration of Single-Cell Data (Seurat v3)
A statistical framework for integrating single-cell datasets across conditions and modalities.
The standard analysis pipeline for one of the most-active areas of biology.
PCA, basic single-cell biology
Anchor-based integration lets you compare cells across very different experiments.
12 Massively Parallel Digital Transcriptional Profiling of Single Cells (Cell Ranger / 10x)
Demonstrates droplet-based single-cell RNA sequencing at hundreds of thousands of cells per run.
The technology paper that made single-cell biology a mainstream field.
RNA-seq basics
Massive throughput on cells, not just reads, opens new biological questions.
13 Highly Accurate Protein Structure Prediction with AlphaFold MVRP
A geometry-aware Transformer architecture (Evoformer + structure module) predicting 3D protein structure to atomic accuracy.
A scientific landmark of historic magnitude; the protein-folding problem largely solved.
Transformers, basic structural biology
Hard scientific problems crack when you bake the right geometric and biological priors into the model.
14 Evolutionary-Scale Prediction of Atomic-Level Protein Structure (ESMFold)
Uses a 15B-parameter protein language model (ESM-2) to predict structure directly from sequence.
Faster, simpler alternative to AlphaFold for many use cases; demonstrates the power of protein LLMs.
AlphaFold, protein language models
A big enough language model implicitly learns structure from sequence alone.
15 Accurate Prediction of Protein Structures and Interactions Using a 3-Track Network (RoseTTAFold)
A three-track network reasoning over sequence, distance, and structure simultaneously.
A useful counterpoint to AlphaFold; same problem, related but distinct architecture.
AlphaFold, basic graph nets
Structure prediction benefits from reasoning over multiple representations in parallel.
16 Robust Deep Learning–Based Protein Sequence Design Using ProteinMPNN MVRP
A graph neural network that designs protein sequences to fold into a desired backbone.
A modern reference for inverse folding and the design half of computational protein engineering.
AlphaFold, GNNs
Inverse design (give me a sequence for this shape) is the dual problem to folding.
17 De Novo Design of Protein Structure and Function with RFdiffusion
A diffusion model that generates novel protein backbones for specified geometric constraints.
The "Stable Diffusion of biology" — generative protein design now actually works.
Diffusion models, protein design basics
Diffusion is a remarkably general way to generate structured biological objects.
18 Transfer Learning Enables Predictions in Network Biology (Geneformer)
A Transformer pretrained on 30M single-cell transcriptomes and fine-tuned for downstream biology.
Foundation-model thinking finally hitting cell biology.
Transformers, single-cell RNA-seq
Pretraining on rank-encoded gene expression yields strong biological transferability.
19 Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3
Extends AlphaFold to predict protein–nucleic-acid, protein–ligand, and protein–protein complexes.
The current state of the art; redefines what computational structural biology can do.
AlphaFold, diffusion models
A single foundation model can handle proteins, RNA, DNA, and small-molecule interactions.
20 Language Models Generalize Beyond Natural Proteins
A protein language model designs novel functional proteins beyond known evolutionary distributions.
A glimpse of foundation-model-driven exploration of biological design space.
ESM, protein design basics
Language models trained on natural proteins can imagine functional unnatural ones.