Computer Architecture
From the von Neumann report to AI accelerators and side-channel attacks — the design choices, performance walls, and security boundaries shaping what software actually runs on.
Minimum viable reading path
The 5 papers that give you most of the field's mental model, in reading order.
01 First Draft of a Report on the EDVAC MVRP
Describes the stored-program computer architecture — instructions and data in the same memory.
Eighty years on, almost every machine you touch is still a von Neumann machine.
Basic logic and arithmetic
Programs are data. The whole field flows from this one observation.
02 Architecture of the IBM System/360
Defines the System/360 ISA family with binary compatibility across multiple price-performance points.
The paper that invented the idea of an ISA as an enduring contract; foundation of x86, ARM, RISC-V.
Basic computer organisation
Separating architecture from implementation lets you reuse software across decades of hardware.
03 Slave Memories and Dynamic Storage Allocation (Caches)
Proposes a small fast "slave memory" between the CPU and main memory — the cache.
A single 1965 idea that still shapes every modern processor.
Memory hierarchy basics
A small fast memory in front of a large slow one is the right shape of the storage system.
04 An Efficient Algorithm for Exploiting Multiple Arithmetic Units (Tomasulo)
A dynamic instruction scheduling algorithm with reservation stations and register renaming.
The conceptual core of every modern out-of-order processor.
Pipelining, dependencies
Renaming destination registers eliminates false dependencies and unlocks ILP.
05 The Case for the Reduced Instruction Set Computer (RISC) MVRP
Argues that simple, regular instruction sets enable faster pipelines and better compilers.
A short, sharp polemic that reshaped processor design for decades.
Pipelining, basic ISAs
Simpler ISAs beat complex ones once you have a good compiler.
06 A Study of Branch Prediction Strategies
Surveys static and dynamic branch predictors and shows how much performance hangs on getting them right.
Foundational for understanding why modern CPUs spend so many transistors on prediction.
Pipelining basics
A few bits of state per branch can outperform any compiler-only static prediction.
07 Cache Coherence Protocols (MESI)
Defines the four-state cache coherence protocol used by most modern multiprocessors.
Required reading for understanding how shared-memory multiprocessing actually works.
Caches, basic concurrency
A small state machine per cache line can make multi-core memory look like single-core memory.
08 The RISC-V Instruction Set Manual
A clean, modular, royalty-free open ISA designed for academic and industrial use.
The first ISA in decades that's actually open — already reshaping the chip industry.
RISC, ISA basics
An ISA can be a public good — and the ecosystem benefits enormously from it.
09 Roofline: An Insightful Visual Performance Model
Plots achievable performance vs arithmetic intensity, exposing where memory or compute bounds you.
The single best mental model for "is this fast enough on this hardware".
Memory bandwidth, FLOPs
Performance is bounded by the lower of compute roof and memory bandwidth roof.
10 Computer Architecture: A Quantitative Approach (Hennessy & Patterson, Ch. 1 excerpt)
The foundational textbook chapter on quantitative architecture analysis and Amdahl's law.
The discipline's standard reference; the chapter on principles is irreplaceable.
Basic computer organisation
Architecture is a quantitative discipline; design choices need numbers attached.
11 Dark Silicon and the End of Multicore Scaling MVRP
Argues that thermal limits force chips to leave huge fractions of silicon dark at any one moment.
The paper that named the wall. Required for understanding the post-Dennard-scaling era.
Dennard scaling, basic VLSI
You can build the transistors but you can't power them all on at once.
12 The Datacenter as a Computer
Reframes the datacenter as a single warehouse-scale machine to be designed, not a building of servers.
Reset the conversation about cloud and warehouse computing; still the cleanest framing.
Computer organisation, networking basics
Treat the entire datacenter as one machine when you design it.
13 A New Golden Age for Computer Architecture
Argues domain-specific architectures and open ISAs are reviving the field after a fallow period.
A short, accessible state-of-the-field essay; the best modern intro for non-architects.
Roofline, ISA basics
Domain-specific accelerators are the way out of the post-Dennard slump.
14 NVIDIA Tesla: A Unified Graphics and Computing Architecture (CUDA)
Describes the unified-shader architecture and the CUDA programming model.
The hardware and abstraction that made deep learning possible.
Computer organisation, SIMD
GPUs aren't graphics cards anymore — they're a parallel computing primitive.
15 In-Datacenter Performance Analysis of a Tensor Processing Unit (TPU) MVRP
A retrospective on Google's first-generation TPU and its 15–30× perf/watt over GPUs for inference.
A landmark domain-specific architecture paper; the most influential industry-architecture paper of the 2010s.
Roofline, matrix multiply
A purpose-built matrix multiplier annihilates a general-purpose accelerator on its target task.
16 Eyeriss: A Spatial Architecture for Energy-Efficient CNN Inference
A spatial accelerator with row-stationary dataflow that minimises off-chip memory accesses.
A canonical academic accelerator paper; teaches you how to think about accelerator dataflow.
CNN basics, memory hierarchy
Where data flows on the chip matters more than how fast you can compute on it.
17 Cerebras Wafer-Scale Engine
A single-chip processor occupying an entire 300mm wafer with 850k cores and 2.6T transistors.
A radical alternative to chiplets; useful for understanding the design space at the extreme.
TPU, multi-core architecture
Sometimes the right answer is to make the chip way bigger, not smaller.
18 Spectre Attacks: Exploiting Speculative Execution MVRP
Speculative execution leaks data across security boundaries via cache side channels.
A class of vulnerabilities every modern CPU shares; remade hardware-security thinking.
Branch prediction, caches, basic OS
Performance optimisations create observable side channels — even when "no instruction commits".
19 Meltdown: Reading Kernel Memory from User Space
A specific micro-architectural attack reading kernel memory from user space on most Intel CPUs.
The companion to Spectre; together they redefined hardware security.
Spectre, OS memory protection
Out-of-order execution can briefly violate isolation in measurable ways.
20 Chiplet Heterogeneous Integration
Industry account of AMD's shift to chiplets and what it bought them in cost, yield, and flexibility.
A clear modern reference on a major industry trend.
VLSI basics, computer organisation
Smaller dies, packaged together, often beat one big monolithic die.
21 Scaling DNNs to Trillions of Parameters with GShard
A compiler and parallelism framework for sparsely-activated giant models, demonstrated on a 600B-parameter MoE.
A bridge between systems and architecture; underlies modern multi-thousand-GPU LLM training.
CUDA, collective communication
For trillion-parameter models, the architecture is the compiler, the network, and the partitioning.
22 Computer Architecture Performance Evaluation Methods
A guide to simulators, benchmarks, and statistical rigour in evaluating processor designs.
Less famous than the architecture papers, but invaluable for actually doing the work.
Basic statistics, computer organisation
Most architecture papers are wrong about something; methodology is how you avoid being one of them.