Papers in Order

Computer Architecture

From the von Neumann report to AI accelerators and side-channel attacks — the design choices, performance walls, and security boundaries shaping what software actually runs on.

22 papers 7 levels Included papers 1945 – 2023 MVRP 5 papers
0 of 22 papers read. Progress stays in this browser.

Minimum viable reading path

The 5 papers that give you most of the field's mental model, in reading order.

  1. First Draft of a Report on the EDVAC
  2. The Case for the Reduced Instruction Set Computer (RISC)
  3. Dark Silicon and the End of Multicore Scaling
  4. In-Datacenter Performance Analysis of a Tensor Processing Unit (TPU)
  5. Spectre Attacks: Exploiting Speculative Execution
Level 0 Foundations
01 First Draft of a Report on the EDVAC MVRP
TL;DR

Describes the stored-program computer architecture — instructions and data in the same memory.

Why read this

Eighty years on, almost every machine you touch is still a von Neumann machine.

Prerequisites

Basic logic and arithmetic

Key takeaway

Programs are data. The whole field flows from this one observation.

Read the paper
02 Architecture of the IBM System/360
TL;DR

Defines the System/360 ISA family with binary compatibility across multiple price-performance points.

Why read this

The paper that invented the idea of an ISA as an enduring contract; foundation of x86, ARM, RISC-V.

Prerequisites

Basic computer organisation

Key takeaway

Separating architecture from implementation lets you reuse software across decades of hardware.

Read the paper
Level 1 Pipelining & ILP
03 Slave Memories and Dynamic Storage Allocation (Caches)
TL;DR

Proposes a small fast "slave memory" between the CPU and main memory — the cache.

Why read this

A single 1965 idea that still shapes every modern processor.

Prerequisites

Memory hierarchy basics

Key takeaway

A small fast memory in front of a large slow one is the right shape of the storage system.

Read the paper
04 An Efficient Algorithm for Exploiting Multiple Arithmetic Units (Tomasulo)
TL;DR

A dynamic instruction scheduling algorithm with reservation stations and register renaming.

Why read this

The conceptual core of every modern out-of-order processor.

Prerequisites

Pipelining, dependencies

Key takeaway

Renaming destination registers eliminates false dependencies and unlocks ILP.

Read the paper
05 The Case for the Reduced Instruction Set Computer (RISC) MVRP
TL;DR

Argues that simple, regular instruction sets enable faster pipelines and better compilers.

Why read this

A short, sharp polemic that reshaped processor design for decades.

Prerequisites

Pipelining, basic ISAs

Key takeaway

Simpler ISAs beat complex ones once you have a good compiler.

Read the paper
Level 2 ISAs & cache coherence
06 A Study of Branch Prediction Strategies
TL;DR

Surveys static and dynamic branch predictors and shows how much performance hangs on getting them right.

Why read this

Foundational for understanding why modern CPUs spend so many transistors on prediction.

Prerequisites

Pipelining basics

Key takeaway

A few bits of state per branch can outperform any compiler-only static prediction.

Read the paper
07 Cache Coherence Protocols (MESI)
TL;DR

Defines the four-state cache coherence protocol used by most modern multiprocessors.

Why read this

Required reading for understanding how shared-memory multiprocessing actually works.

Prerequisites

Caches, basic concurrency

Key takeaway

A small state machine per cache line can make multi-core memory look like single-core memory.

Read the paper
08 The RISC-V Instruction Set Manual
TL;DR

A clean, modular, royalty-free open ISA designed for academic and industrial use.

Why read this

The first ISA in decades that's actually open — already reshaping the chip industry.

Prerequisites

RISC, ISA basics

Key takeaway

An ISA can be a public good — and the ecosystem benefits enormously from it.

Read the paper
Level 3 Performance models
09 Roofline: An Insightful Visual Performance Model
TL;DR

Plots achievable performance vs arithmetic intensity, exposing where memory or compute bounds you.

Why read this

The single best mental model for "is this fast enough on this hardware".

Prerequisites

Memory bandwidth, FLOPs

Key takeaway

Performance is bounded by the lower of compute roof and memory bandwidth roof.

Read the paper
10 Computer Architecture: A Quantitative Approach (Hennessy & Patterson, Ch. 1 excerpt)
TL;DR

The foundational textbook chapter on quantitative architecture analysis and Amdahl's law.

Why read this

The discipline's standard reference; the chapter on principles is irreplaceable.

Prerequisites

Basic computer organisation

Key takeaway

Architecture is a quantitative discipline; design choices need numbers attached.

Read the paper
Level 4 Power wall & multicore
11 Dark Silicon and the End of Multicore Scaling MVRP
TL;DR

Argues that thermal limits force chips to leave huge fractions of silicon dark at any one moment.

Why read this

The paper that named the wall. Required for understanding the post-Dennard-scaling era.

Prerequisites

Dennard scaling, basic VLSI

Key takeaway

You can build the transistors but you can't power them all on at once.

Read the paper
12 The Datacenter as a Computer
TL;DR

Reframes the datacenter as a single warehouse-scale machine to be designed, not a building of servers.

Why read this

Reset the conversation about cloud and warehouse computing; still the cleanest framing.

Prerequisites

Computer organisation, networking basics

Key takeaway

Treat the entire datacenter as one machine when you design it.

Read the paper
13 A New Golden Age for Computer Architecture
TL;DR

Argues domain-specific architectures and open ISAs are reviving the field after a fallow period.

Why read this

A short, accessible state-of-the-field essay; the best modern intro for non-architects.

Prerequisites

Roofline, ISA basics

Key takeaway

Domain-specific accelerators are the way out of the post-Dennard slump.

Read the paper
Level 5 Specialised hardware
14 NVIDIA Tesla: A Unified Graphics and Computing Architecture (CUDA)
TL;DR

Describes the unified-shader architecture and the CUDA programming model.

Why read this

The hardware and abstraction that made deep learning possible.

Prerequisites

Computer organisation, SIMD

Key takeaway

GPUs aren't graphics cards anymore — they're a parallel computing primitive.

Read the paper
15 In-Datacenter Performance Analysis of a Tensor Processing Unit (TPU) MVRP
TL;DR

A retrospective on Google's first-generation TPU and its 15–30× perf/watt over GPUs for inference.

Why read this

A landmark domain-specific architecture paper; the most influential industry-architecture paper of the 2010s.

Prerequisites

Roofline, matrix multiply

Key takeaway

A purpose-built matrix multiplier annihilates a general-purpose accelerator on its target task.

Read the paper
16 Eyeriss: A Spatial Architecture for Energy-Efficient CNN Inference
TL;DR

A spatial accelerator with row-stationary dataflow that minimises off-chip memory accesses.

Why read this

A canonical academic accelerator paper; teaches you how to think about accelerator dataflow.

Prerequisites

CNN basics, memory hierarchy

Key takeaway

Where data flows on the chip matters more than how fast you can compute on it.

Read the paper
17 Cerebras Wafer-Scale Engine
TL;DR

A single-chip processor occupying an entire 300mm wafer with 850k cores and 2.6T transistors.

Why read this

A radical alternative to chiplets; useful for understanding the design space at the extreme.

Prerequisites

TPU, multi-core architecture

Key takeaway

Sometimes the right answer is to make the chip way bigger, not smaller.

Read the paper
Level 6 Security & modern era
18 Spectre Attacks: Exploiting Speculative Execution MVRP
TL;DR

Speculative execution leaks data across security boundaries via cache side channels.

Why read this

A class of vulnerabilities every modern CPU shares; remade hardware-security thinking.

Prerequisites

Branch prediction, caches, basic OS

Key takeaway

Performance optimisations create observable side channels — even when "no instruction commits".

Read the paper
19 Meltdown: Reading Kernel Memory from User Space
TL;DR

A specific micro-architectural attack reading kernel memory from user space on most Intel CPUs.

Why read this

The companion to Spectre; together they redefined hardware security.

Prerequisites

Spectre, OS memory protection

Key takeaway

Out-of-order execution can briefly violate isolation in measurable ways.

Read the paper
20 Chiplet Heterogeneous Integration
TL;DR

Industry account of AMD's shift to chiplets and what it bought them in cost, yield, and flexibility.

Why read this

A clear modern reference on a major industry trend.

Prerequisites

VLSI basics, computer organisation

Key takeaway

Smaller dies, packaged together, often beat one big monolithic die.

Read the paper
21 Scaling DNNs to Trillions of Parameters with GShard
TL;DR

A compiler and parallelism framework for sparsely-activated giant models, demonstrated on a 600B-parameter MoE.

Why read this

A bridge between systems and architecture; underlies modern multi-thousand-GPU LLM training.

Prerequisites

CUDA, collective communication

Key takeaway

For trillion-parameter models, the architecture is the compiler, the network, and the partitioning.

Read the paper
22 Computer Architecture Performance Evaluation Methods
TL;DR

A guide to simulators, benchmarks, and statistical rigour in evaluating processor designs.

Why read this

Less famous than the architecture papers, but invaluable for actually doing the work.

Prerequisites

Basic statistics, computer organisation

Key takeaway

Most architecture papers are wrong about something; methodology is how you avoid being one of them.

Read the paper