Papers in Order

Information Retrieval

From tf-idf to neural retrievers — how systems learned to find the right documents (and now the right passages) for a query, and how that became the foundation of modern search and RAG.

22 papers 6 levels Included papers 1972 – 2024 MVRP 6 papers
0 of 22 papers read. Progress stays in this browser.

Minimum viable reading path

The 6 papers that give you most of the field's mental model, in reading order.

  1. A Statistical Interpretation of Term Specificity and its Application in Retrieval
  2. Okapi at TREC-3 (BM25)
  3. The Anatomy of a Large-Scale Hypertextual Web Search Engine (Google)
  4. Passage Re-ranking with BERT
  5. Dense Passage Retrieval for Open-Domain Question Answering (DPR)
  6. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG)
Level 0 Classical IR
01 A Statistical Interpretation of Term Specificity and its Application in Retrieval MVRP
TL;DR

Introduces inverse document frequency: rare terms carry more retrieval weight than common ones.

Why read this

The idf in tf-idf comes from this short paper — IR's foundational quantitative idea.

Prerequisites

Basic statistics

Key takeaway

How specific a term is matters at least as much as how often it appears.

Read the paper
02 A Vector Space Model for Automatic Indexing
TL;DR

Represents documents and queries as vectors in a term space and ranks by cosine similarity.

Why read this

The geometric framing of IR that everyone still uses; conceptual ancestor of every embedding-based retriever.

Prerequisites

Linear algebra, basic NLP

Key takeaway

Documents are points in space; relevance is geometry.

Read the paper
03 Relevance Weighting of Search Terms
TL;DR

Derives optimal term weights for retrieval under a probabilistic model of relevance.

Why read this

The probabilistic foundation under BM25 and most classical IR systems.

Prerequisites

tf-idf, probability theory

Key takeaway

Optimal weighting falls out of treating relevance as a probability.

Read the paper
Level 1 Probabilistic & language models
04 Indexing by Latent Semantic Analysis (LSI)
TL;DR

Uses singular value decomposition on the term-document matrix to capture latent topics.

Why read this

The first attempt at semantic, dimensionality-reduced retrieval; the lineage to embeddings.

Prerequisites

Vector space model, SVD

Key takeaway

Retrieval can happen in a low-dimensional latent space, not the original term space.

Read the paper
05 Okapi at TREC-3 (BM25) MVRP
TL;DR

A probabilistic ranking function with saturating term frequency and length normalisation.

Why read this

Still the strongest non-neural baseline you can run; under-rated relative to its impact.

Prerequisites

Probabilistic relevance

Key takeaway

A small handful of well-chosen heuristics gives a baseline that's remarkably hard to beat.

Read the paper
06 A Language Modeling Approach to Information Retrieval
TL;DR

Models each document as a language model and ranks by the likelihood it generated the query.

Why read this

A clean re-framing of IR that connects retrieval to statistical NLP.

Prerequisites

Probability, language modeling basics

Key takeaway

Documents that "could have generated" the query are the relevant ones.

Read the paper
Level 2 Web-scale & link analysis
07 The Anatomy of a Large-Scale Hypertextual Web Search Engine (Google) MVRP
TL;DR

Describes the original Google: PageRank, anchor text, distributed crawling, repository design.

Why read this

A founding document of the modern web. Still the clearest introduction to how a real search engine is built.

Prerequisites

Vector space model, basic graph theory

Key takeaway

Web search needs both content signals and link signals; one isn't enough.

Read the paper
08 The PageRank Citation Ranking: Bringing Order to the Web
TL;DR

Ranks pages by the stationary distribution of a random walk over the link graph.

Why read this

A timeless trick for ranking by structural authority on any graph.

Prerequisites

Linear algebra, Markov chains

Key takeaway

A page's importance is recursively the importance of pages that link to it.

Read the paper
09 Authoritative Sources in a Hyperlinked Environment (HITS)
TL;DR

Identifies hubs and authorities via mutually-reinforcing link structure.

Why read this

A useful conceptual contrast to PageRank — different framings of "authority on the web".

Prerequisites

PageRank, eigenvectors

Key takeaway

Authority and hub-ness are dual concepts that emerge from link structure.

Read the paper
Level 3 Learning to rank
10 Learning to Rank using Gradient Descent (RankNet)
TL;DR

A pairwise neural ranker trained on cross-entropy of relevance comparisons.

Why read this

The first paper of the modern learning-to-rank era; conceptual seed of LambdaMART and beyond.

Prerequisites

Neural network basics

Key takeaway

Treat ranking as a pairwise classification problem and you can learn it with gradient descent.

Read the paper
11 From RankNet to LambdaRank to LambdaMART: An Overview
TL;DR

Walks through the evolution of LTR algorithms used in production search engines.

Why read this

A clean, self-contained narrative through three classic algorithms; the practical LTR canon.

Prerequisites

RankNet, gradient boosting

Key takeaway

A principled non-smooth IR metric becomes trainable through pairwise lambda gradients.

Read the paper
Level 4 Neural IR
12 Learning Deep Structured Semantic Models (DSSM)
TL;DR

Learns matching score between query and document via deep neural networks over character n-grams.

Why read this

The first widely-influential neural retrieval paper; predates BERT-era methods by half a decade.

Prerequisites

Neural networks, vector space model

Key takeaway

Even small dense networks beat lexical retrieval when trained on click data.

Read the paper
13 A Deep Relevance Matching Model for Ad-hoc Retrieval (DRMM)
TL;DR

Argues that retrieval is fundamentally about exact matching, not just semantic matching.

Why read this

A useful corrective to early neural-IR enthusiasm; sharpens the design space.

Prerequisites

DSSM, basic IR

Key takeaway

Strong neural retrievers must combine semantic and exact-match signals, not just one.

Read the paper
14 Passage Re-ranking with BERT MVRP
TL;DR

Fine-tunes BERT to score query-passage pairs, dramatically improving MS MARCO scores.

Why read this

The "BERT just works" moment for IR; the start of the BERT-for-search wave.

Prerequisites

BERT, basic IR metrics

Key takeaway

Cross-encoders trivially beat lexical baselines for re-ranking.

Read the paper
15 Dense Passage Retrieval for Open-Domain Question Answering (DPR) MVRP
TL;DR

Trains dual BERT encoders to embed queries and passages into a shared space for nearest-neighbour retrieval.

Why read this

The most-cited dense-retrieval paper; the dual-encoder pattern still dominates.

Prerequisites

BERT, contrastive learning

Key takeaway

A dual encoder + contrastive loss gives you a strong, scalable dense retriever.

Read the paper
16 ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction
TL;DR

Encodes each token separately and scores with late MaxSim interaction at retrieval time.

Why read this

A practical sweet spot between cross-encoder accuracy and dual-encoder speed.

Prerequisites

BERT, DPR

Key takeaway

Late, fine-grained token-level interaction is much cheaper than full cross-attention.

Read the paper
17 SPLADE: Sparse Lexical and Expansion Model for First-Stage Ranking
TL;DR

Produces sparse expansion vectors over the BERT vocabulary, retrievable with classical inverted indexes.

Why read this

The most pragmatic neural-IR design — neural quality with infrastructure that already exists.

Prerequisites

BM25, BERT

Key takeaway

You can stay on inverted indexes and still get most of the neural-IR benefit.

Read the paper
21 Efficient and Robust Approximate Nearest Neighbor Search Using HNSW Graphs
TL;DR

A multi-layer navigable small-world graph giving logarithmic-time approximate nearest-neighbour search.

Why read this

The index inside most production vector databases — the other half of dense retrieval.

Prerequisites

Vector space model, basic graph algorithms

Key takeaway

A hierarchy of long- and short-range links makes greedy graph search fast and accurate.

Read the paper
22 BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models
TL;DR

An 18-dataset benchmark showing dense retrievers often lose to BM25 out of domain.

Why read this

The reality check for neural IR — defines how retrieval generalisation is measured.

Prerequisites

BM25, DPR

Key takeaway

In-domain wins don't transfer; evaluate retrieval zero-shot before believing it.

Read the paper
Level 5 Generative & retrieval-augmented
18 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG) MVRP
TL;DR

Combines a Transformer generator with a learned dense retriever over a non-parametric memory.

Why read this

The paper that defines RAG and the foundation of every "talk to your docs" system.

Prerequisites

BERT, DPR

Key takeaway

External memory beats trying to cram every fact into model weights.

Read the paper
19 Transformer Memory as a Differentiable Search Index (DSI)
TL;DR

Trains a single Transformer to map queries directly to document identifiers — no retriever, no index.

Why read this

A radical reframing of search as pure sequence generation; influence is just starting.

Prerequisites

BERT, T5

Key takeaway

A model can memorise an entire corpus into its weights and still retrieve from it.

Read the paper
20 Text Embeddings by Weakly-Supervised Contrastive Pre-training (E5)
TL;DR

Trains general-purpose text embeddings with multi-stage contrastive training on weakly-supervised pairs.

Why read this

A strong baseline for modern open-source embeddings; widely adopted.

Prerequisites

DPR, contrastive learning

Key takeaway

Carefully-mined weak supervision is enough to train a top-tier general embedding model.

Read the paper