Information Retrieval
From tf-idf to neural retrievers — how systems learned to find the right documents (and now the right passages) for a query, and how that became the foundation of modern search and RAG.
Minimum viable reading path
The 6 papers that give you most of the field's mental model, in reading order.
- A Statistical Interpretation of Term Specificity and its Application in Retrieval
- Okapi at TREC-3 (BM25)
- The Anatomy of a Large-Scale Hypertextual Web Search Engine (Google)
- Passage Re-ranking with BERT
- Dense Passage Retrieval for Open-Domain Question Answering (DPR)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG)
01 A Statistical Interpretation of Term Specificity and its Application in Retrieval MVRP
Introduces inverse document frequency: rare terms carry more retrieval weight than common ones.
The idf in tf-idf comes from this short paper — IR's foundational quantitative idea.
Basic statistics
How specific a term is matters at least as much as how often it appears.
02 A Vector Space Model for Automatic Indexing
Represents documents and queries as vectors in a term space and ranks by cosine similarity.
The geometric framing of IR that everyone still uses; conceptual ancestor of every embedding-based retriever.
Linear algebra, basic NLP
Documents are points in space; relevance is geometry.
03 Relevance Weighting of Search Terms
Derives optimal term weights for retrieval under a probabilistic model of relevance.
The probabilistic foundation under BM25 and most classical IR systems.
tf-idf, probability theory
Optimal weighting falls out of treating relevance as a probability.
04 Indexing by Latent Semantic Analysis (LSI)
Uses singular value decomposition on the term-document matrix to capture latent topics.
The first attempt at semantic, dimensionality-reduced retrieval; the lineage to embeddings.
Vector space model, SVD
Retrieval can happen in a low-dimensional latent space, not the original term space.
05 Okapi at TREC-3 (BM25) MVRP
A probabilistic ranking function with saturating term frequency and length normalisation.
Still the strongest non-neural baseline you can run; under-rated relative to its impact.
Probabilistic relevance
A small handful of well-chosen heuristics gives a baseline that's remarkably hard to beat.
06 A Language Modeling Approach to Information Retrieval
Models each document as a language model and ranks by the likelihood it generated the query.
A clean re-framing of IR that connects retrieval to statistical NLP.
Probability, language modeling basics
Documents that "could have generated" the query are the relevant ones.
07 The Anatomy of a Large-Scale Hypertextual Web Search Engine (Google) MVRP
Describes the original Google: PageRank, anchor text, distributed crawling, repository design.
A founding document of the modern web. Still the clearest introduction to how a real search engine is built.
Vector space model, basic graph theory
Web search needs both content signals and link signals; one isn't enough.
08 The PageRank Citation Ranking: Bringing Order to the Web
Ranks pages by the stationary distribution of a random walk over the link graph.
A timeless trick for ranking by structural authority on any graph.
Linear algebra, Markov chains
A page's importance is recursively the importance of pages that link to it.
09 Authoritative Sources in a Hyperlinked Environment (HITS)
Identifies hubs and authorities via mutually-reinforcing link structure.
A useful conceptual contrast to PageRank — different framings of "authority on the web".
PageRank, eigenvectors
Authority and hub-ness are dual concepts that emerge from link structure.
10 Learning to Rank using Gradient Descent (RankNet)
A pairwise neural ranker trained on cross-entropy of relevance comparisons.
The first paper of the modern learning-to-rank era; conceptual seed of LambdaMART and beyond.
Neural network basics
Treat ranking as a pairwise classification problem and you can learn it with gradient descent.
11 From RankNet to LambdaRank to LambdaMART: An Overview
Walks through the evolution of LTR algorithms used in production search engines.
A clean, self-contained narrative through three classic algorithms; the practical LTR canon.
RankNet, gradient boosting
A principled non-smooth IR metric becomes trainable through pairwise lambda gradients.
12 Learning Deep Structured Semantic Models (DSSM)
Learns matching score between query and document via deep neural networks over character n-grams.
The first widely-influential neural retrieval paper; predates BERT-era methods by half a decade.
Neural networks, vector space model
Even small dense networks beat lexical retrieval when trained on click data.
13 A Deep Relevance Matching Model for Ad-hoc Retrieval (DRMM)
Argues that retrieval is fundamentally about exact matching, not just semantic matching.
A useful corrective to early neural-IR enthusiasm; sharpens the design space.
DSSM, basic IR
Strong neural retrievers must combine semantic and exact-match signals, not just one.
14 Passage Re-ranking with BERT MVRP
Fine-tunes BERT to score query-passage pairs, dramatically improving MS MARCO scores.
The "BERT just works" moment for IR; the start of the BERT-for-search wave.
BERT, basic IR metrics
Cross-encoders trivially beat lexical baselines for re-ranking.
15 Dense Passage Retrieval for Open-Domain Question Answering (DPR) MVRP
Trains dual BERT encoders to embed queries and passages into a shared space for nearest-neighbour retrieval.
The most-cited dense-retrieval paper; the dual-encoder pattern still dominates.
BERT, contrastive learning
A dual encoder + contrastive loss gives you a strong, scalable dense retriever.
16 ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction
Encodes each token separately and scores with late MaxSim interaction at retrieval time.
A practical sweet spot between cross-encoder accuracy and dual-encoder speed.
BERT, DPR
Late, fine-grained token-level interaction is much cheaper than full cross-attention.
17 SPLADE: Sparse Lexical and Expansion Model for First-Stage Ranking
Produces sparse expansion vectors over the BERT vocabulary, retrievable with classical inverted indexes.
The most pragmatic neural-IR design — neural quality with infrastructure that already exists.
BM25, BERT
You can stay on inverted indexes and still get most of the neural-IR benefit.
21 Efficient and Robust Approximate Nearest Neighbor Search Using HNSW Graphs
A multi-layer navigable small-world graph giving logarithmic-time approximate nearest-neighbour search.
The index inside most production vector databases — the other half of dense retrieval.
Vector space model, basic graph algorithms
A hierarchy of long- and short-range links makes greedy graph search fast and accurate.
22 BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models
An 18-dataset benchmark showing dense retrievers often lose to BM25 out of domain.
The reality check for neural IR — defines how retrieval generalisation is measured.
BM25, DPR
In-domain wins don't transfer; evaluate retrieval zero-shot before believing it.
18 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG) MVRP
Combines a Transformer generator with a learned dense retriever over a non-parametric memory.
The paper that defines RAG and the foundation of every "talk to your docs" system.
BERT, DPR
External memory beats trying to cram every fact into model weights.
19 Transformer Memory as a Differentiable Search Index (DSI)
Trains a single Transformer to map queries directly to document identifiers — no retriever, no index.
A radical reframing of search as pure sequence generation; influence is just starting.
BERT, T5
A model can memorise an entire corpus into its weights and still retrieve from it.
20 Text Embeddings by Weakly-Supervised Contrastive Pre-training (E5)
Trains general-purpose text embeddings with multi-stage contrastive training on weakly-supervised pairs.
A strong baseline for modern open-source embeddings; widely adopted.
DPR, contrastive learning
Carefully-mined weak supervision is enough to train a top-tier general embedding model.