skip to content

Inductive Training at Scale

Whether a trained model works on nodes it never saw, and how you train when pulling two hops of neighbours already touches most of the graph. Interviewers probe the neighbour-explosion problem.

on this pageshow

questions

3

In graph neural networks, what separates transductive from inductive training?

level: juniorimportance: must knowfreq 74%

answer

  1. who was in the graph at fit time
  2. per-node parameters versus a shared function
  3. can a node created tonight be scored?
  4. new nodes, or an entirely new graph

basics

~20 s

Transductive training assumes every node you will ever score was already in the graph when the model was fitted. Inductive training learns a function of node features and neighbourhoods, so a node added later can be scored without refitting.

solid answer

~50 s

The split is about which nodes the model is allowed to have seen at fit time. A **transductive** setup trains on one fixed graph and only ever scores nodes inside it; if the model holds free parameters attached to specific node IDs, a node that appears later simply has no parameters and cannot be scored at all. An **inductive** setup learns shared weights that map a node's own features plus its sampled neighbourhood into a representation, so the same weights apply to a node created after training — or, in the strongest form, to an entirely different graph. This is an architectural property, not a data split: holding out labels for nodes that were still present in the training graph is transductive semi-supervised learning, and it will overstate how well the model handles genuinely new nodes.

go deeper

for a junior

Be ready to state the distinction in one sentence and give one concrete consequence: a node created after training cannot be scored by a model that stores one learned vector per node.

for a middle

Explain the mechanism — shared weight matrices applied to features and aggregated neighbour features generalise to new nodes, while free per-node parameters cannot — and name what the inference path must supply.

for a senior

An interviewer expects you to catch the evaluation leak: held-out nodes left inside the training graph make a transductive result look inductive. Say how you would rebuild the split to measure it honestly.

for a principal

Own the consequence for the platform: whether new entities can be scored between training runs decides your retraining cadence, your cold-start fallback, and whether serving has a hard dependency on the training pipeline.

## The question behind the words "Transductive" and "inductive" answer one thing: **which nodes was the model allowed to see while it was being fitted, and which nodes can it therefore score afterwards?** Everything else follows from that. ### Transductive training In the transductive setting there is one fixed graph. Training runs over the whole adjacency at once, labels exist for some nodes, and the goal is to fill in labels for the *other nodes of that same graph*. Crucially, the unlabelled nodes were present the entire time — their edges participated in every message-passing step, and their features (if any) flowed into their neighbours' representations. A model becomes hard-transductive when it carries **free parameters indexed by node identity**: a table with one trainable vector per node, learned directly by gradient descent. Nothing in that table generalises. A node that did not exist when the table was allocated has no row, and there is no forward pass that can produce one. The only remedy is to refit. A model can also be *softly* transductive through its training procedure rather than its parameters. Full-graph training that normalises by global quantities, or an evaluation that leaves test nodes and their edges inside the training graph, both assume the graph is closed. ### Inductive training In the inductive setting the model's parameters are **shared functions**, not per-node values: weight matrices that take a node's own feature vector plus aggregated feature vectors from its neighbours and return a representation. Because the same weights are applied at every node, they apply just as well to a node the model has never seen. Scoring a new node needs only two things at inference time: its features, and access to its current neighbours. That second requirement is the one people forget. An inductive model is not context-free — it still has to read the neighbourhood. A brand-new node with no edges yet falls back to whatever the architecture does with an empty neighbourhood, usually just its own transformed features, which is a much weaker signal. ### Three degrees of "unseen" It helps to grade the difficulty: 1. **Unlabelled but present.** The node was in the training graph; only its label was withheld. This is semi-supervised *transductive* node classification. It is the easiest case and the one most benchmark numbers report. 2. **New node, same graph.** The node was created after the training run — a new account, a new product, a new device. The model must produce a representation from features and current neighbours alone. 3. **New graph entirely.** The strongest form. A protein–protein interaction benchmark has exactly this shape: fit on around twenty tissue graphs, tune on a couple more, then run on a tissue graph the model has never touched, whose node identities mean nothing across graphs. Only shared, feature-driven weights survive that transfer. ### Why this bites in production Consider a feed-ranking graph rebuilt daily. Accounts created since last night's training run must be scored tonight; there is no retraining window before the traffic arrives. If the ranking model is transductive, those accounts are unscorable and fall into a cold-start default path, and the cost of being wrong compounds because new accounts are exactly the population you most want to rank correctly. If it is inductive, tonight's scoring job builds each new account's representation from its profile features and the neighbours it has accumulated today, using the same weights fitted last week. Retraining then becomes a periodic quality refresh rather than a hard dependency of serving. ### The evaluation trap The most common way to fool yourself is to *claim* inductive evaluation while measuring the transductive case. If your held-out nodes were present in the graph during training — their edges intact, their features feeding neighbours — then information about them leaked into the fitted weights and into every neighbour representation. A genuine inductive evaluation **removes the test nodes and all their incident edges from the training graph**, fits, then reinserts them only at inference. The gap between the two protocols is often large, and discovering it after launch is expensive. ### How to tell which one you have Ask one question of the architecture: *if a node ID appeared that was never in the training set, what would the forward pass do?* If the answer is "index a table and fail", it is transductive. If the answer is "transform its features and aggregate its neighbours' features", it is inductive. Mixed designs exist — feature-driven weights plus a learned per-node correction — and they inherit the transductive limitation for any node whose correction does not exist.

  • Does a model that can score unseen nodes automatically transfer to an entirely new graph?
    Not automatically. Scoring a new node in a familiar graph and scoring a graph you have never seen are different bars. Transfer across graphs also needs the feature semantics to line up — the same feature slots meaning the same thing — and the degree and homophily statistics to be comparable. The protein-interaction setup, where you fit on around twenty tissue graphs and test on an unseen tissue, is the benchmark shape that actually measures this.
  • A feed-ranking graph is rebuilt daily and accounts created today must be scored tonight. What has to be true of the model?
    It has to be inductive end to end: shared weights, node features available at scoring time, and a serving path that can fetch each new account's current neighbours. Retraining then becomes a periodic quality refresh, not a prerequisite for serving. You also need a defined behaviour for accounts with no edges yet, since their representation collapses to their own transformed features.
  • How would you design the split so that an inductive claim is actually measured?
    Remove the evaluation nodes and every edge incident to them from the graph used for fitting, train, then reinsert them only at inference. Leaving those nodes in place — even unlabelled — lets their features and edges shape the learned weights and their neighbours' representations, which inflates the reported score relative to what serving will see.

saying these in an interview costs you the question

  • Assumes any neural model on a graph is automatically inductive
  • Thinks inductive just means having a held-out test set
  • Says scoring a brand-new node requires retraining every time
  • Forgets that an inductive model still needs node features at inference
  • Evaluates on held-out nodes that stayed inside the training graph

context

open as a page

Why does mini-batching a 3-layer GNN on a high-degree graph cause neighbour explosion?

level: middleimportance: must knowfreq 63%

basics

~20 s

Each message-passing layer adds one hop to a node's receptive field, so a batch holds degree-to-the-power-of-layers nodes per seed. At average degree 100 with three layers that is about a million nodes for one seed, which fixed per-hop fanout caps.

open as a page

How does cluster-based subgraph mini-batching train a GNN on a 200-million-node graph?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Partition the graph once into dense clusters that minimise cut edges, then make each mini-batch one or a few whole clusters and propagate inside that subgraph only. Neighbourhoods stay bounded, but every cross-partition edge is dropped for that step.

open as a page