Papers in Order

Database Systems

From the relational model to learned indexes — the data structures, transaction protocols, and storage architectures that organise the world's information.

24 papers 7 levels Included papers 1970 – 2024 MVRP 3 papers
0 of 24 papers read. Progress stays in this browser.

Minimum viable reading path

The 3 papers that give you most of the field's mental model, in reading order.

  1. A Relational Model of Data for Large Shared Data Banks
  2. Bigtable: A Distributed Storage System for Structured Data
  3. Spanner: Google's Globally-Distributed Database
Level 0 Foundations
01 A Relational Model of Data for Large Shared Data Banks MVRP
TL;DR

Proposes the relational model: data as relations (tables), with set-theoretic operators for querying.

Why read this

The single founding document of modern databases. Read it to see SQL's entire intellectual scaffolding in 11 pages.

Prerequisites

Set theory, basic logic

Key takeaway

Separate the logical structure of data from how it's physically stored.

Read the paper
02 System R: Relational Approach to Database Management
TL;DR

Reports on the first full implementation of a relational system, including SQL, optimisation, and transactions.

Why read this

The blueprint of every commercial RDBMS that followed.

Prerequisites

Codd's relational model

Key takeaway

Almost every modern RDBMS architectural choice was made first in System R.

Read the paper
03 The Notions of Consistency and Predicate Locks
TL;DR

Defines serialisability and lays out the locking protocols for transactions.

Why read this

The transaction-theory paper underlying every concurrency-control implementation.

Prerequisites

Concurrency basics

Key takeaway

Two-phase locking + serialisability is the recipe for transactional correctness.

Read the paper
Level 1 Recovery & concurrency
04 ARIES: A Transaction Recovery Method
TL;DR

Write-ahead logging with redo, undo, and physiological logging — the basis of every modern crash-recovery system.

Why read this

The reference for how databases survive crashes; foundational for storage engineers.

Prerequisites

Two-phase locking, B-trees

Key takeaway

Write-ahead logging plus carefully ordered redo/undo gives durability without giving up performance.

Read the paper
05 Hekaton: SQL Server's Memory-Optimized OLTP Engine
TL;DR

A latch-free, lock-free in-memory OLTP engine that uses MVCC and compiled query plans.

Why read this

A clean, modern industry account of in-memory transaction processing.

Prerequisites

ARIES, concurrency control

Key takeaway

When everything fits in RAM, the entire database design changes.

Read the paper
Level 2 Column stores & analytics
06 C-Store: A Column-Oriented DBMS
TL;DR

Argues that storing data by column, not row, is the right design for analytical workloads.

Why read this

Read this against System R to see how analytical workloads pulled DB design in a new direction.

Prerequisites

System R

Key takeaway

OLTP and analytics want fundamentally different storage layouts.

Read the paper
07 The End of an Architectural Era (H-Store)
TL;DR

A polemic against general-purpose RDBMSs, arguing for specialised in-memory engines.

Why read this

A bracing critique that anticipates Hekaton, MemSQL, and most modern OLTP engines.

Prerequisites

System R, basic OLTP

Key takeaway

General-purpose database architectures leave too much performance on the table.

Read the paper
08 The Vertica Analytic Database: C-Store 7 Years Later
TL;DR

A retrospective on what worked and didn't in commercialising C-Store.

Why read this

A rare, candid post-mortem on translating a research prototype into a production system.

Prerequisites

C-Store

Key takeaway

Real workloads break clean research designs in interesting and instructive ways.

Read the paper
Level 3 NoSQL & scale-out
09 Bigtable: A Distributed Storage System for Structured Data MVRP
TL;DR

A sparse, distributed, persistent multi-dimensional sorted map built atop GFS and Chubby.

Why read this

A foundational system; deserves to be read once for distributed systems and once for databases.

Prerequisites

GFS, LSM-trees

Key takeaway

A simple sorted-key data model plus an LSM scales to huge datasets.

Read the paper
10 Dynamo: Amazon's Highly Available Key-value Store
TL;DR

A masterless, eventually-consistent KV store that prioritises availability over consistency.

Why read this

The conceptual ancestor of every AP-side NoSQL store.

Prerequisites

Distributed systems basics, consistent hashing

Key takeaway

Tuneable consistency at the operation level beats a global setting.

Read the paper
11 PNUTS: Yahoo!'s Hosted Data Serving Platform
TL;DR

A geo-replicated database with per-record timeline consistency between strong and eventual.

Why read this

A pragmatic middle-ground design between Dynamo and Spanner — clarifies the consistency spectrum.

Prerequisites

Distributed systems basics

Key takeaway

Real applications usually want a consistency knob, not a fixed setting.

Read the paper
Level 4 Storage engines
12 The Log-Structured Merge-Tree (LSM-Tree)
TL;DR

A write-optimised data structure that batches updates in memory and merges them to disk in sorted runs.

Why read this

The foundational data structure of every modern KV store: RocksDB, LevelDB, Cassandra.

Prerequisites

B-trees, basic algorithms

Key takeaway

Trade some read amplification for huge wins on write throughput.

Read the paper
13 The Case for Learned Index Structures
TL;DR

Replaces B-trees and Bloom filters with neural networks that learn the data's CDF.

Why read this

A speculative-but-influential paper redefining what a data structure can be.

Prerequisites

B-trees, basic ML

Key takeaway

A data structure is just a function; you might as well learn it.

Read the paper
Level 5 Cloud-native databases
14 Spanner: Google's Globally-Distributed Database MVRP
TL;DR

A globally-distributed SQL database with externally-consistent transactions powered by TrueTime.

Why read this

The watershed paper that made global, strongly-consistent transactions practical.

Prerequisites

Bigtable, Paxos

Key takeaway

Bound your clock skew, wait it out, and you can serialise transactions across continents.

Read the paper
15 F1: A Distributed SQL Database That Scales
TL;DR

The SQL layer atop Spanner that runs Google's AdWords business.

Why read this

Companion piece to Spanner; shows the SQL engineering needed on top of a distributed KV layer.

Prerequisites

Spanner, SQL internals

Key takeaway

Most apps want SQL even when they need a distributed KV — so build the SQL layer well.

Read the paper
16 The Snowflake Elastic Data Warehouse
TL;DR

Decouples compute and storage in the cloud, with elastic, billed-by-the-second virtual warehouses.

Why read this

A clean modern reference on the cloud-native warehouse architecture now industry-standard.

Prerequisites

C-Store, cloud architecture basics

Key takeaway

Decoupled compute and storage is the right shape for the cloud.

Read the paper
17 CockroachDB: The Resilient Geo-Distributed SQL Database
TL;DR

An open-source SQL database that approximates Spanner's guarantees without atomic clocks.

Why read this

A clear case study in re-implementing a research system as production-grade open source.

Prerequisites

Spanner, Raft

Key takeaway

Hybrid logical clocks and careful engineering can substitute for TrueTime.

Read the paper
Level 6 Modern frontiers
18 DuckDB: An Embeddable Analytical Database
TL;DR

An in-process columnar SQL engine designed for analytical queries on a single machine.

Why read this

The "SQLite of analytics" — represents a real movement back toward single-node power.

Prerequisites

C-Store

Key takeaway

Most analytical workloads fit on one big modern machine, and shouldn't pay distributed overhead.

Read the paper
19 Photon: A Fast Query Engine for Lakehouse Systems
TL;DR

A vectorised C++ execution engine for Apache Spark that beats hand-tuned warehouses on TPC-DS.

Why read this

A modern reference on vectorised execution and the lakehouse architecture.

Prerequisites

Spark, vectorised execution

Key takeaway

You can replace a slow JVM execution layer without breaking the rest of the stack.

Read the paper
20 Self-Driving Database Management Systems
TL;DR

Argues for autonomous DBMSs that tune themselves using ML over their workloads.

Why read this

A useful manifesto for where database operations is heading.

Prerequisites

Database internals

Key takeaway

The DBMS itself is a control system that should learn its own configuration.

Read the paper
21 FAISS: A Library for Efficient Similarity Search
TL;DR

GPU-accelerated billion-scale nearest-neighbour search via inverted indexes and product quantisation.

Why read this

The reference work for vector search; foundation of every "vector database".

Prerequisites

Linear algebra, basic IR

Key takeaway

Approximate nearest-neighbour search is mostly an indexing and quantisation problem.

Read the paper
22 Lakehouse: A New Generation of Open Platforms
TL;DR

Argues for unified data architectures combining the openness of lakes with the management of warehouses.

Why read this

Frames the data-platform debate of the early 2020s.

Prerequisites

Snowflake, data lakes basics

Key takeaway

Schema, transactions, and governance can live atop open file formats.

Read the paper
23 RocksDB: Evolution of Development Priorities in a Key-Value Store
TL;DR

A retrospective on a decade of running RocksDB as a foundational storage engine inside Meta.

Why read this

Honest reflections on what really matters in a long-lived storage engine.

Prerequisites

LSM-tree

Key takeaway

Real-world priorities (efficiency, then features, then performance) shift dramatically over time.

Read the paper
24 Vectorized Query Execution at Microsoft (SQL Server Batch Mode)
TL;DR

Adds batch-mode columnstore execution into SQL Server, mixing column and row processing in one engine.

Why read this

Underrated production retrofit paper; useful contrast with the column-only DuckDB and Photon approaches.

Prerequisites

C-Store, query execution basics

Key takeaway

You don't have to ditch your row engine — batch mode and column stores can coexist.

Read the paper