Papers in Order

Graph Neural Networks

From random-walk embeddings to molecular structure prediction — how neural networks learned to operate on graphs and unlocked science applications from drug discovery to AlphaFold.

23 papers 7 levels Included papers 2013 – 2024 MVRP 4 papers
0 of 23 papers read. Progress stays in this browser.

Minimum viable reading path

The 4 papers that give you most of the field's mental model, in reading order.

  1. node2vec: Scalable Feature Learning for Networks
  2. Semi-Supervised Classification with Graph Convolutional Networks (GCN)
  3. Graph Attention Networks (GAT)
  4. Highly Accurate Protein Structure Prediction with AlphaFold 2
Level 0 Node embeddings
01 DeepWalk: Online Learning of Social Representations
TL;DR

Treats random walks on graphs as sentences and applies Word2Vec to learn node embeddings.

Why read this

The transfer of NLP techniques to graphs that opened the door to neural graph methods.

Prerequisites

Word2Vec

Key takeaway

A graph and a corpus look the same once you sample walks.

Read the paper
02 LINE: Large-scale Information Network Embedding
TL;DR

Learns node embeddings preserving both first- and second-order proximity, with negative sampling.

Why read this

Refines DeepWalk by being explicit about which graph structure to preserve.

Prerequisites

DeepWalk

Key takeaway

Different proximity orders capture different aspects of network structure.

Read the paper
03 node2vec: Scalable Feature Learning for Networks MVRP
TL;DR

Generalises DeepWalk with biased random walks that interpolate between BFS and DFS exploration.

Why read this

The most-used random-walk embedding; conceptually clean way to tune local-vs-global emphasis.

Prerequisites

DeepWalk

Key takeaway

Two hyperparameters let you tune between structural roles and community membership.

Read the paper
Level 1 Spectral methods
04 Spectral Networks and Locally Connected Networks on Graphs
TL;DR

The first attempt to define convolution on graphs via the graph Laplacian and spectral filters.

Why read this

Where graph convolutions began; understand it to follow why later methods spatial-localise.

Prerequisites

Linear algebra, CNN basics, spectral graph theory

Key takeaway

Convolution on graphs is naturally defined in the frequency domain.

Read the paper
05 Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering (ChebNet)
TL;DR

Approximates spectral filters with K-th-order Chebyshev polynomials, giving K-hop locality.

Why read this

The bridge from pure spectral methods to spatial GCNs.

Prerequisites

Spectral GCN

Key takeaway

You can keep the spectral interpretation while making everything local and fast.

Read the paper
Level 2 Core GNN architectures
06 Semi-Supervised Classification with Graph Convolutional Networks (GCN) MVRP
TL;DR

A first-order approximation of ChebNet that simplifies to a single neighbourhood-averaging layer.

Why read this

The simplest, most-cited GNN; the conceptual baseline against which all others are measured.

Prerequisites

ChebNet

Key takeaway

Most of the value of spectral GCNs is captured by a one-line update rule.

Read the paper
07 Inductive Representation Learning on Large Graphs (GraphSAGE)
TL;DR

Samples and aggregates neighbour features so embeddings generalise to unseen nodes and graphs.

Why read this

Lets GNNs work on dynamic and out-of-sample graphs — practically essential.

Prerequisites

GCN

Key takeaway

Sampled neighbourhoods + learnable aggregators turn GNNs into inductive learners.

Read the paper
08 Graph Attention Networks (GAT) MVRP
TL;DR

Replaces uniform neighbourhood averaging with attention-weighted aggregation.

Why read this

The "attention is all you need" moment for graphs — now a default building block.

Prerequisites

GCN, attention mechanism

Key takeaway

Let nodes decide how much to weight each neighbour.

Read the paper
Level 3 Theory & expressivity
09 Neural Message Passing for Quantum Chemistry (MPNN)
TL;DR

Unifies prior GNN variants under a generic message-passing framework and applies it to molecular properties.

Why read this

The framework paper that gives every GNN a common language; foundational for chemistry applications.

Prerequisites

GCN, basic chemistry useful

Key takeaway

Most GNNs are special cases of the same message-update-readout structure.

Read the paper
10 How Powerful are Graph Neural Networks? (GIN)
TL;DR

Proves message-passing GNNs are at most as expressive as the 1-WL graph isomorphism test.

Why read this

The cleanest theoretical lens on what GNNs can and cannot distinguish.

Prerequisites

MPNN, graph isomorphism basics

Key takeaway

Sum aggregators with MLPs are maximally expressive within the message-passing class.

Read the paper
11 Deep Graph Infomax (DGI)
TL;DR

Self-supervised GNN training by maximising mutual information between local and global graph summaries.

Why read this

A clean, label-free objective for graph representation learning.

Prerequisites

GCN, mutual information basics

Key takeaway

You can pretrain GNNs without labels by contrasting local and global views.

Read the paper
23 Relational Inductive Biases, Deep Learning, and Graph Networks
TL;DR

A position paper unifying GNN variants into a general graph-network block and arguing for relational reasoning.

Why read this

The best conceptual overview of why graphs are the right abstraction for structured computation.

Prerequisites

MPNN

Key takeaway

Architecture choices are inductive-bias choices — and graphs encode composability.

Read the paper
Level 4 Scaling
12 Graph Convolutional Neural Networks for Web-Scale Recommender Systems (PinSAGE)
TL;DR

A GCN variant deployed on 3B-node Pinterest graph using random-walk-based neighbour sampling.

Why read this

The clearest case study of GNNs at production web scale.

Prerequisites

GraphSAGE

Key takeaway

GNNs scale via clever sampling, not bigger machines.

Read the paper
13 Cluster-GCN: An Efficient Algorithm for Training Deep and Large GCNs
TL;DR

Uses graph clustering to form mini-batches that fit in memory while preserving local structure.

Why read this

A practical recipe for training deep GCNs on large graphs without subsampling artefacts.

Prerequisites

GCN, graph partitioning

Key takeaway

Cluster the graph first, then train per-cluster.

Read the paper
14 GraphSAINT: Graph Sampling Based Inductive Learning Method
TL;DR

Samples subgraphs (not just neighbourhoods) and trains a full GCN on each, with bias correction.

Why read this

A complementary scaling strategy to Cluster-GCN with cleaner statistical properties.

Prerequisites

GraphSAGE

Key takeaway

Sampling subgraphs gives a less biased gradient than sampling per-node neighbourhoods.

Read the paper
Level 5 Specialised graphs
15 Modeling Relational Data with Graph Convolutional Networks (R-GCN)
TL;DR

Extends GCNs to multi-relation graphs by parameterising different transforms per edge type.

Why read this

The natural step from homogeneous to heterogeneous graphs; underlies most KG-embedding GNNs.

Prerequisites

GCN, knowledge-graph basics

Key takeaway

Different edge types deserve different transformations.

Read the paper
16 Heterogeneous Graph Attention Network (HAN)
TL;DR

Two-level attention: node-level over neighbours and semantic-level over meta-paths.

Why read this

Standard reference for working with heterogeneous (typed) graphs in practice.

Prerequisites

GAT, meta-path concept

Key takeaway

In heterogeneous graphs, attend over both neighbours and relation types.

Read the paper
17 SchNet: A Continuous-Filter Convolutional Neural Network for Modeling Quantum Interactions
TL;DR

Uses continuous-filter convolutions over molecular distances to predict quantum-mechanical properties.

Why read this

A landmark in chemistry-aware GNNs; shows how physical priors integrate cleanly with message passing.

Prerequisites

MPNN, basic quantum chemistry helpful

Key takeaway

Distances and physical symmetries belong inside the architecture, not learned from data.

Read the paper
18 Directional Message Passing for Molecular Graphs (DimeNet)
TL;DR

Adds bond-angle information to molecular GNNs via embeddings of edges, not just nodes.

Why read this

A clean lesson in how the right inductive bias unlocks new tasks (geometry, energies).

Prerequisites

SchNet

Key takeaway

Geometric primitives like angles and dihedrals belong in the architecture.

Read the paper
Level 6 Frontier
19 Highly Accurate Protein Structure Prediction with AlphaFold 2 MVRP
TL;DR

A geometry-aware Transformer/GNN hybrid (Evoformer) that predicts protein 3D structure to atomic accuracy.

Why read this

A scientific result of historic magnitude — and a master class in domain-aware architecture design.

Prerequisites

Transformers, GNNs, basic structural biology

Key takeaway

Hard scientific problems crack when you bake the right geometric and biological priors into the model.

Read the paper
20 Do Transformers Really Perform Bad for Graph Representation? (Graphormer)
TL;DR

Adds graph-aware positional and edge encodings to a vanilla Transformer; wins OGB benchmarks.

Why read this

The cleanest demonstration that Transformers can be the right architecture for graphs too.

Prerequisites

Transformers, GNNs

Key takeaway

Most graph inductive biases can be encoded as biases on attention scores.

Read the paper
21 Recipe for a General, Powerful, Scalable Graph Transformer (GraphGPS)
TL;DR

A modular framework that combines local message passing with global attention and various positional encodings.

Why read this

A useful taxonomy of graph-Transformer design choices.

Prerequisites

GAT, Graphormer

Key takeaway

Local and global views of the graph are complementary; combine them explicitly.

Read the paper
22 GFlowNet Foundations
TL;DR

A new generative-modeling framework that samples diverse high-reward objects (e.g. molecules) proportionally to their reward.

Why read this

A radical alternative to RL for combinatorial generation, especially for science applications.

Prerequisites

RL, flow-based generative models

Key takeaway

Sometimes you want a sampler, not a maximiser — and the math is different.

Read the paper