Graph Neural Networks
From random-walk embeddings to molecular structure prediction — how neural networks learned to operate on graphs and unlocked science applications from drug discovery to AlphaFold.
Minimum viable reading path
The 4 papers that give you most of the field's mental model, in reading order.
01 DeepWalk: Online Learning of Social Representations
Treats random walks on graphs as sentences and applies Word2Vec to learn node embeddings.
The transfer of NLP techniques to graphs that opened the door to neural graph methods.
Word2Vec
A graph and a corpus look the same once you sample walks.
02 LINE: Large-scale Information Network Embedding
Learns node embeddings preserving both first- and second-order proximity, with negative sampling.
Refines DeepWalk by being explicit about which graph structure to preserve.
DeepWalk
Different proximity orders capture different aspects of network structure.
03 node2vec: Scalable Feature Learning for Networks MVRP
Generalises DeepWalk with biased random walks that interpolate between BFS and DFS exploration.
The most-used random-walk embedding; conceptually clean way to tune local-vs-global emphasis.
DeepWalk
Two hyperparameters let you tune between structural roles and community membership.
04 Spectral Networks and Locally Connected Networks on Graphs
The first attempt to define convolution on graphs via the graph Laplacian and spectral filters.
Where graph convolutions began; understand it to follow why later methods spatial-localise.
Linear algebra, CNN basics, spectral graph theory
Convolution on graphs is naturally defined in the frequency domain.
05 Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering (ChebNet)
Approximates spectral filters with K-th-order Chebyshev polynomials, giving K-hop locality.
The bridge from pure spectral methods to spatial GCNs.
Spectral GCN
You can keep the spectral interpretation while making everything local and fast.
06 Semi-Supervised Classification with Graph Convolutional Networks (GCN) MVRP
A first-order approximation of ChebNet that simplifies to a single neighbourhood-averaging layer.
The simplest, most-cited GNN; the conceptual baseline against which all others are measured.
ChebNet
Most of the value of spectral GCNs is captured by a one-line update rule.
07 Inductive Representation Learning on Large Graphs (GraphSAGE)
Samples and aggregates neighbour features so embeddings generalise to unseen nodes and graphs.
Lets GNNs work on dynamic and out-of-sample graphs — practically essential.
GCN
Sampled neighbourhoods + learnable aggregators turn GNNs into inductive learners.
08 Graph Attention Networks (GAT) MVRP
Replaces uniform neighbourhood averaging with attention-weighted aggregation.
The "attention is all you need" moment for graphs — now a default building block.
GCN, attention mechanism
Let nodes decide how much to weight each neighbour.
09 Neural Message Passing for Quantum Chemistry (MPNN)
Unifies prior GNN variants under a generic message-passing framework and applies it to molecular properties.
The framework paper that gives every GNN a common language; foundational for chemistry applications.
GCN, basic chemistry useful
Most GNNs are special cases of the same message-update-readout structure.
10 How Powerful are Graph Neural Networks? (GIN)
Proves message-passing GNNs are at most as expressive as the 1-WL graph isomorphism test.
The cleanest theoretical lens on what GNNs can and cannot distinguish.
MPNN, graph isomorphism basics
Sum aggregators with MLPs are maximally expressive within the message-passing class.
11 Deep Graph Infomax (DGI)
Self-supervised GNN training by maximising mutual information between local and global graph summaries.
A clean, label-free objective for graph representation learning.
GCN, mutual information basics
You can pretrain GNNs without labels by contrasting local and global views.
23 Relational Inductive Biases, Deep Learning, and Graph Networks
A position paper unifying GNN variants into a general graph-network block and arguing for relational reasoning.
The best conceptual overview of why graphs are the right abstraction for structured computation.
MPNN
Architecture choices are inductive-bias choices — and graphs encode composability.
12 Graph Convolutional Neural Networks for Web-Scale Recommender Systems (PinSAGE)
A GCN variant deployed on 3B-node Pinterest graph using random-walk-based neighbour sampling.
The clearest case study of GNNs at production web scale.
GraphSAGE
GNNs scale via clever sampling, not bigger machines.
13 Cluster-GCN: An Efficient Algorithm for Training Deep and Large GCNs
Uses graph clustering to form mini-batches that fit in memory while preserving local structure.
A practical recipe for training deep GCNs on large graphs without subsampling artefacts.
GCN, graph partitioning
Cluster the graph first, then train per-cluster.
14 GraphSAINT: Graph Sampling Based Inductive Learning Method
Samples subgraphs (not just neighbourhoods) and trains a full GCN on each, with bias correction.
A complementary scaling strategy to Cluster-GCN with cleaner statistical properties.
GraphSAGE
Sampling subgraphs gives a less biased gradient than sampling per-node neighbourhoods.
15 Modeling Relational Data with Graph Convolutional Networks (R-GCN)
Extends GCNs to multi-relation graphs by parameterising different transforms per edge type.
The natural step from homogeneous to heterogeneous graphs; underlies most KG-embedding GNNs.
GCN, knowledge-graph basics
Different edge types deserve different transformations.
16 Heterogeneous Graph Attention Network (HAN)
Two-level attention: node-level over neighbours and semantic-level over meta-paths.
Standard reference for working with heterogeneous (typed) graphs in practice.
GAT, meta-path concept
In heterogeneous graphs, attend over both neighbours and relation types.
17 SchNet: A Continuous-Filter Convolutional Neural Network for Modeling Quantum Interactions
Uses continuous-filter convolutions over molecular distances to predict quantum-mechanical properties.
A landmark in chemistry-aware GNNs; shows how physical priors integrate cleanly with message passing.
MPNN, basic quantum chemistry helpful
Distances and physical symmetries belong inside the architecture, not learned from data.
18 Directional Message Passing for Molecular Graphs (DimeNet)
Adds bond-angle information to molecular GNNs via embeddings of edges, not just nodes.
A clean lesson in how the right inductive bias unlocks new tasks (geometry, energies).
SchNet
Geometric primitives like angles and dihedrals belong in the architecture.
19 Highly Accurate Protein Structure Prediction with AlphaFold 2 MVRP
A geometry-aware Transformer/GNN hybrid (Evoformer) that predicts protein 3D structure to atomic accuracy.
A scientific result of historic magnitude — and a master class in domain-aware architecture design.
Transformers, GNNs, basic structural biology
Hard scientific problems crack when you bake the right geometric and biological priors into the model.
20 Do Transformers Really Perform Bad for Graph Representation? (Graphormer)
Adds graph-aware positional and edge encodings to a vanilla Transformer; wins OGB benchmarks.
The cleanest demonstration that Transformers can be the right architecture for graphs too.
Transformers, GNNs
Most graph inductive biases can be encoded as biases on attention scores.
21 Recipe for a General, Powerful, Scalable Graph Transformer (GraphGPS)
A modular framework that combines local message passing with global attention and various positional encodings.
A useful taxonomy of graph-Transformer design choices.
GAT, Graphormer
Local and global views of the graph are complementary; combine them explicitly.
22 GFlowNet Foundations
A new generative-modeling framework that samples diverse high-reward objects (e.g. molecules) proportionally to their reward.
A radical alternative to RL for combinatorial generation, especially for science applications.
RL, flow-based generative models
Sometimes you want a sampler, not a maximiser — and the math is different.