In GraphSAGE, when does the max-pooling aggregator beat the mean over sampled neighbours?
answer
- typical neighbour versus any neighbour
- a rare signal gets averaged away
- learned detectors, then elementwise max
- extremes depend on how many you looked at
- a sampled maximum understates the true maximum
basics
~20 sMax-pooling wins when a rare, salient neighbour is the signal: each sampled neighbour passes through a small learned layer and the aggregate is an elementwise max, so one distinctive neighbour survives. A mean dilutes it.
solid answer
~50 sThe mean aggregator averages the sampled neighbours' vectors elementwise. The pooling aggregator first pushes each sampled neighbour through a small learned layer with a nonlinearity, then takes an elementwise max — a bank of learned detectors answering 'does any neighbour look like this?'. Pooling wins when presence beats prevalence: one fraudulent counterparty among two hundred honest ones moves a max but barely moves a mean. The mean wins when the neighbourhood's typical profile is the signal, when features are noisy, and when you want fewer parameters and a lower-variance aggregate. Sampling interacts with the choice: the mean of a uniform neighbour sample is an unbiased estimate of the full-neighbourhood mean, but a sample's max is biased low relative to the full neighbourhood's, and the bias grows with the true neighbourhood size — so keep the fanout consistent between training and inference when you use pooling.
go deeper
Know that a layer must turn a variable number of neighbour vectors into one fixed-size vector, and that an elementwise mean and an elementwise max are two order-independent ways of doing it.
Explain the pooling aggregator precisely — a shared small layer applied to every neighbour, then an elementwise max — and say why that construction is invariant to neighbour ordering.
Show you have operated this: choose the aggregator from the shape of the signal, keep the training and inference fanout aligned so pooled statistics do not shift, and check results per degree bucket.
Own the tradeoff between a detector-style aggregate that catches rare risk and a smooth one that is stable under graph drift and adversarial edges, and set a defensible default for the team.
## What each aggregator computes Both aggregators take a variable-sized set of neighbour vectors and return one fixed-size vector, and both must be invariant to the order the neighbours arrive in. **Mean.** `AGG_mean = (1/|S|) * sum over j in S of h_j`, where `S` is the sampled neighbour set. It is a linear, parameter-free summary: the average neighbour. **Max-pooling.** `AGG_pool = elementwise max over j in S of sigma(W_pool * h_j + b)`. Every sampled neighbour is first pushed through a small learned layer with a nonlinearity, and the result is reduced by an elementwise maximum. Each output dimension is a learned detector, and the max answers "did *any* sampled neighbour fire this detector, and how strongly". The learned layer is what makes it more than a max over raw features: the model chooses which directions in feature space are worth detecting. Both are permutation-invariant, which is mandatory — a node's neighbours have no canonical order. ## Prevalence versus presence This is the decision, and it is a statement about the domain. - **The mean encodes prevalence.** It answers "what fraction of my neighbourhood looks like X". If half a node's counterparties are new accounts, that shows up cleanly. If one in two hundred is, the contribution is scaled by 1/200 and can vanish under the noise of the other 199. - **The max encodes presence.** One neighbour that lights up a detector sets that output dimension regardless of how many others exist. In fraud, money laundering, or abuse detection, "is this account connected to a *known* bad actor" is precisely a presence question, and a mean is structurally bad at it. - **What the max throws away.** Multiplicity. One suspicious neighbour and fifty suspicious neighbours can produce the same pooled vector. If the difference between "one bad link" and "a whole bad neighbourhood" matters, a max alone will not tell you — which is why using both aggregates side by side is a reasonable design. - **Robustness.** The mean is smooth and forgiving of noisy individual features; the max is an extreme-value statistic and is therefore sensitive to a single corrupted, mislabelled, or adversarially-crafted neighbour. On a graph where an attacker can add edges, that sensitivity is an attack surface. ## How sampling changes the comparison The aggregate is computed over a uniformly sampled fixed-size neighbour set, not the whole neighbourhood, and the two aggregators respond to that very differently. **The sampled mean is unbiased.** For a uniform sample `S` of size `n` drawn from a neighbourhood of size `N`, `E[mean(S)] = mean(full neighbourhood)`, because the mean is linear in the sampled items. Sampling only adds variance, on the order of the neighbourhood's feature variance divided by `n`. Changing the sample size between training and inference changes how noisy the estimate is, not what it estimates. **The sampled max is biased low.** `E[max(S)] <= max(full neighbourhood)`, with equality only when the argmax is certain to be sampled. If a node has 3,000 neighbours and exactly one fires a detector, a 25-neighbour sample finds it with probability about 25/3000, under 1%. So the same node's pooled representation depends systematically on both the fanout and the node's true degree — a shift that looks like a distribution shift but is entirely an artifact of the estimator. Two practical consequences: 1. **Keep the fanout consistent between training and inference when you pool.** If you train with a small sample and evaluate over the full neighbourhood, pooled activations shift upward at evaluation time, and the layer above them was never trained on that regime. The mean does not suffer this: full-neighbourhood inference just gives it a lower-variance version of the same quantity. 2. **Watch the per-degree slices.** Bias in a pooled aggregate scales with true degree, so pooled models can silently degrade on exactly the high-degree nodes that often matter most. Evaluate by degree bucket, not only in aggregate. Sampling also has an upside for both: the noise it injects acts as a regularizer, much as dropping units does, so a model trained on sampled neighbourhoods is often more robust than one trained on the full set. ## Choosing between them Prefer pooling when the label depends on the existence of a distinctive neighbour, when neighbourhoods are heterogeneous, and when you can afford the extra parameters of the per-neighbour layer. Prefer the mean when the neighbourhood's overall composition is the signal, when features are noisy, when the labelled set is small, or when you need a stable aggregate under a changing graph. When you are unsure, run both under a matched fanout and an honest split — typically a time-based one on a transaction or interaction graph, so future edges cannot leak backwards — and compare per-degree as well as overall. ## Misconceptions worth avoiding - "Max-pooling takes the max of the raw neighbour features." It takes the max *after* a shared learned layer with a nonlinearity; that layer is the point. - "The aggregator choice is independent of the sampling scheme." The mean is sample-unbiased and the max is not, which is a first-order interaction. - "Max is strictly more expressive, so always use it." It discards multiplicity and is fragile to a single bad neighbour. - "Sampling only speeds things up." It also changes the statistic you are estimating when that statistic is an extremum.
- Why is a sampled mean unbiased while a sampled max is not?The mean is linear in the sampled items, so the expectation of the sample mean equals the population mean whatever the sample size. A maximum is not linear: the sample's largest value can never exceed the neighbourhood's, and falls short whenever the true argmax is not drawn, so its expectation sits strictly below the full-neighbourhood max.
- What does a max-pooling aggregate throw away that a mean keeps?Multiplicity. One neighbour firing a detector and fifty firing it can give the same pooled value, so 'a single risky link' and 'an entirely risky neighbourhood' become indistinguishable. The mean encodes that proportion directly, which is why concatenating a mean aggregate alongside a pooled one is a common, cheap hedge.
- How would you decide between them empirically on a transaction graph?Train both with the same fanout, the same depth, and a time-based split so future edges cannot leak into training. Compare on the metric that matches the cost of a miss, then break results out by node degree — pooled aggregates degrade preferentially on high-degree nodes, which an overall average will hide.
Screening a neighbourhood by average income tells you the typical household; screening by the single loudest smoke alarm tells you whether anything is on fire. Sampling a few houses estimates the average well and the loudest alarm badly.
saying these in an interview costs you the question
- Says max-pooling takes the maximum over raw neighbour features
- Treats the aggregator choice as independent of how neighbours are sampled
- Assumes a sampled maximum is an unbiased estimate of the true maximum
- Claims max-pooling is strictly more expressive so always preferable
- Ignores that a max discards how many neighbours carried the signal
- Trains with a small fanout then evaluates over full neighbourhoods with pooling