skip to content

How do you choose the number of message-passing rounds in a graph neural network?

level: seniorimportance: should knowfreq 46%

answer

  1. count hops, not layers
  2. how far away does the evidence sit
  3. diameter is a ceiling, not a target
  4. two is the usual landing point
  5. sweep k and read validation spread

basics

~20 s

k rounds give a node a receptive field of exactly k hops, so pick k to match the distance at which the label's evidence lives, then confirm it with a small validation sweep. Most production graph models land at two or three.

solid answer

~50 s

Rounds are hops: after `k` rounds a node's representation depends on every node within `k` edges and nothing further. So the choice is a modelling question — how far away does the evidence for the label sit? On a co-purchase graph, one round tells an item about products bought alongside it; a second reaches its neighbours' neighbours, which is usually where taste communities become visible. A third often adds much less, because in a small-world graph three hops already sweep in a large and mostly generic slice of the catalogue, and each round costs another pass over the edges. Fix everything else, sweep k = 1, 2, 3, 4 over several seeds, and keep the smallest k the next value does not meaningfully beat. Diameter is not the target: a label decided by direct partners needs one or two rounds even on a graph twelve hops wide.

go deeper

for a junior

Remember the one rule that matters: k rounds means a node sees exactly k hops out. Being able to say that two rounds reach neighbours of neighbours, and that most models use two or three, covers the screening version of this.

for a middle

Explain why the receptive field grows by exactly one hop per round and why depth here is about distance rather than capacity. Be ready to say what you would try when the evidence sits further away than your depth reaches.

for a senior

Show the diagnosis and the experiment: how you picked k for a real graph, what the sweep looked like, how you handled seed noise, and whether you broke the comparison down by degree or by distance to labelled nodes.

for a principal

Own the tradeoff between depth, serving cost and graph engineering — when the right answer is a shallower model on a graph you deliberately rewired with shortcut edges or a virtual node, and how you would justify that structural choice to a team that wants to just stack more rounds.

## Rounds are hops, exactly After round one, a node's state depends on itself and its immediate neighbours. After round two, those neighbours' states already carried *their* neighbours' information, so the node now depends on everything within two edges. Inductively, `k` rounds give a receptive field of exactly `k` hops — no more, and no less along any path. This single fact drives the whole decision: choosing `k` is choosing the radius of graph the model is allowed to look at. Note what that means for a node whose evidence sits three hops away: with two rounds, no amount of width, training time or regularisation will help. The information never physically arrives. Conversely, if the label is decided by direct neighbours, extra rounds do not add signal — they add context. ## Start from the task's radius, not the graph's size Ask where the evidence lives. - **One hop.** Fraud on a transaction graph where the label is driven by who a party transacted with directly. One round, sometimes two. - **Two hops.** Recommendation over a co-purchase graph. Round one tells a product about items bought in the same basket; round two reaches its neighbours' neighbours, which is where a coherent taste community shows up — the accessories of the accessories, the second book by an author you never bought directly. This is why two is the most common shipped answer. - **Three or more hops.** Genuinely long-range tasks: molecular properties that depend on distant substructure, or reachability-flavoured labels on sparse graphs. The graph's **diameter** is a ceiling, not a target. A social graph twelve hops across is still typically small-world: a third round on a hub-rich graph pulls in a slice of the graph so large and so generic that the marginal signal per node is small, because almost every node is in almost every other node's three-hop neighbourhood. Radius chosen by the task beats radius chosen by the topology. ## Make it an empirical decision The defensible procedure is boring and works: 1. Hold width, learning rate, aggregator and features fixed. 2. Train k = 1, 2, 3, 4, several seeds each, and read validation performance with its spread across seeds. 3. Take the **smallest** k whose score the next value does not beat by more than seed noise. Cheaper to serve, fewer parameters, easier to debug. 4. Break the aggregate metric down by node degree and by how far the node sits from labelled examples. A third round that does nothing on average may be exactly what rescues sparsely connected nodes, and an average hides that. Use a split that respects the graph. If the evaluation nodes' `k`-hop neighbourhoods overlap the training nodes, a larger k can look better simply because it reaches more labelled context — measure with the split you will actually face in production, including the case where new nodes arrive with few edges. ## Widening the field without adding rounds If the diagnosis is genuinely *the signal is further away than my depth reaches*, depth is not the only lever. Changing the graph changes distances: adding shortcut edges between nodes you have external reason to link, or attaching a virtual node connected to every node so any pair becomes two hops apart, brings distant context within reach of a shallow stack. Those are structural interventions with their own costs — a virtual node can act as a global averaging channel that washes out local detail — but they are the right thing to consider when the alternative is stacking rounds you do not otherwise want. ## What a strong answer sounds like Name the hop equivalence, reason about the task's radius rather than the graph's, give a number with the evidence that produced it, and be honest that two is the usual landing point for a reason. A weak answer treats rounds like layers in an image model — *more depth, more capacity* — and never connects depth to distance on the graph.

  • Your graph has a diameter of twelve, but the label depends only on a party's direct transaction partners. How many rounds?
    One, maybe two. The diameter tells you how far information *could* travel, not how far the evidence sits. Extra rounds mix in distant nodes that carry no information about this label, cost another pass over the edges each, and make debugging harder. Match depth to the task's radius and validate that a second round earns its place.
  • How would you tell empirically that a third round is not helping?
    Fix everything else, train k = 2 and k = 3 with several seeds each, and compare validation scores against the seed-to-seed spread. If k = 3 sits inside the noise band, keep k = 2. Then break the comparison down by node degree and by distance to labelled nodes — a third round can help sparse or peripheral nodes while looking flat on the overall average.
  • Can you widen a node's receptive field without adding rounds?
    Yes, by changing distances in the graph rather than the depth of the model: add shortcut edges between nodes you have independent reason to connect, or attach a virtual node adjacent to every node so any two nodes are two hops apart. Both are real interventions with costs — a virtual node in particular can act as a global averaging channel that dilutes local structure — so validate rather than assume.

saying these in an interview costs you the question

  • Treats rounds like image-model depth, more is more capacity
  • Sets the number of rounds equal to the graph diameter
  • Thinks one round already reaches neighbours of neighbours
  • Never validates k, just copies a number from a paper
  • Reports only an overall metric when comparing depths

context