What hand-built graph features would you demand as a baseline before funding a GNN?
answer
- know what the cheap alternative already achieves
- degree, importance score, triangles, clustering
- any ordinary tabular classifier consumes them
- graph-wide computation leaks across the split
- fixed summaries cannot learn neighbour interactions
basics
~20 sA few per-node structural columns fed to a plain tabular classifier: in-degree and out-degree, a PageRank score, triangle count and clustering coefficient. It trains in minutes, stays interpretable, and gives the proposal a number to beat.
solid answer
~50 sInsist on a tabular baseline first. For hyperlink spam-site detection, compute per node: in-degree and out-degree, a PageRank score, triangle count, local clustering coefficient, and simple neighbour aggregates such as the mean degree of neighbours — then concatenate the node's own attributes and fit any ordinary tabular classifier. This costs hours, not a quarter; it is interpretable enough to defend to a risk team; it refits cheaply; and it frequently captures most of the available signal, because link spam is largely a degree and reciprocity anomaly. The GNN proposal then has a specific number to beat and a specific gap to explain, rather than being judged against nothing. Two cautions I would attach as conditions: compute the features only from edges available at prediction time, or the whole-graph computation leaks future structure into training; and keep the baseline in production afterwards as a monitoring reference and a cold-start fallback.
go deeper
Know that degree, a PageRank score, triangle count and clustering coefficient are numbers you can attach to each node and hand straight to an ordinary classifier, with no graph model involved.
Explain what each column measures and why they separate a genuinely embedded node from a manufactured one, and be able to define local clustering coefficient as triangles over possible triples among neighbours.
Show the evaluation discipline: split by time, compute features from the snapshot available at prediction time, and scrutinise label-derived neighbour aggregates for timing leaks before believing any offline score.
Own the funding decision itself — require a baseline number, a ceiling estimate and a serving cost, and be able to articulate the specific signal a fixed feature set cannot express that would justify the investment.
## Why the baseline is a governance question, not a modelling detail A GNN proposal is a request for a serving stack, a graph store, a training pipeline and a team's quarter. The decision of whether to grant it is not a modelling preference; it is a resource allocation with an opportunity cost. The only honest way to make it is to know what the cheap alternative already achieves. That is what the structural-feature baseline is for, and demanding one is a lead's job. ## The feature set Every one of these is a scalar computed per node, permutation-invariant, and consumable by any tabular model. - **In-degree and out-degree.** How many links point at the page, how many it emits. On a hyperlink graph, an enormous out-degree with negligible in-degree is already a spam tell. - **PageRank score.** A single column expressing global importance under the link structure. Used here as a *feature*, not as an algorithm to reimplement. - **Triangle count.** How many closed triples the node participates in — a measure of how embedded it is in a genuine cluster. - **Local clustering coefficient.** Triangle count normalised by the possible triples among the node's neighbours: `2T / (d * (d - 1))` for an undirected node of degree `d`. It separates a node in a real community from a node at the centre of a manufactured star. - **Neighbour aggregates.** Mean and max of neighbours' degrees or PageRank; the fraction of neighbours that are themselves flagged, where labels permit. - **The node's own attributes.** Whatever the domain already supplies. For link-spam detection this handful is unusually strong, because manufactured link structure has a characteristic shape: high reciprocal density inside a farm, low embedding into the rest of the web, degree distributions that do not look organic. ## What the baseline buys you - **A number.** Every subsequent claim is measured against it. - **Speed.** Feature computation plus a tabular fit is an afternoon, so you learn whether the graph carries signal at all before committing. - **Interpretability.** You can explain to a policy or risk team why a site was demoted in terms of degree and clustering, which matters when demotion has an appeals process. - **Operational simplicity.** No embedding table to version, no space that re-orients on refit, no cold-start hole — features are computed on demand for any node with edges, including brand-new ones. - **A permanent asset.** Even after a GNN wins, these columns remain useful: concatenated as inputs, retained as a monitoring reference that detects graph-level drift, and kept as the fallback path when the heavy model is unavailable. ## The leakage trap — the condition to attach Structural features are computed on a graph, and a graph is global. If you compute PageRank and triangle counts on the full graph and then split nodes into train and test, every test node's edges have already influenced the training features. Offline scores inflate and production disappoints. Two disciplines fix it: - **Split by time.** Compute features from the graph snapshot as it existed at the prediction moment, and evaluate on nodes and edges that appeared later. This is what production will actually have. - **Watch label-derived features.** Neighbour aggregates over labels are the strongest columns and the most dangerous ones: if a neighbour's label was assigned after the prediction time, or was derived from the same enforcement action, the feature is leaking the answer. The same discipline governs the refresh cadence: a feature that is cheap to compute in a batch job may be expensive to compute at request time for a node whose neighbourhood spans millions of edges, and that gap is a serving design question to settle before, not after, the model ships. ## What the baseline structurally cannot do Be fair to the proposal you are challenging. Hand-built features are fixed summaries chosen in advance. They cannot learn a combination of *neighbour attributes* — for example that a page is suspicious specifically when its high-degree neighbours share a registration pattern — because no column encodes neighbour attribute interactions, only pre-chosen scalar aggregates. Adding more hand-built columns hits diminishing returns and a combinatorial wall. If the hypothesis is that the useful signal lives in learned, multi-hop interactions among neighbour attributes, the baseline will plateau, and that plateau is exactly the evidence that funds the next step. ## How I would frame the decision Ask for three numbers: the baseline's metric on a time-split evaluation, an estimate of the ceiling from label noise or an audit sample, and the incremental serving cost of the proposed model. Fund the GNN when the baseline plateaus well below the ceiling *and* the gap is plausibly attributable to learned neighbour interactions. Decline, for now, when the baseline is close to the ceiling, when the graph is small enough that the features already saturate it, or when the team cannot yet operate a graph pipeline. That is not conservatism — it is refusing to spend a quarter to find out something an afternoon could have told you.
- How do you stop these features from leaking across a train/test split?Split by time and compute each node's features from the graph snapshot available at its prediction moment, not from the final graph. Otherwise test-time edges shape training features and offline scores inflate. Watch label-derived neighbour aggregates hardest: if a neighbour's flag was assigned after the prediction time or by the same enforcement action, the column is simply leaking the answer.
- What signal is this baseline structurally unable to capture?Learned interactions among neighbour attributes across multiple hops. The columns are scalar summaries chosen in advance, so the model can learn that high degree with low clustering is suspicious, but not that a page is suspicious specifically when its high-degree neighbours share an attribute pattern. Adding more hand-built columns hits a combinatorial wall, and that plateau is the evidence that justifies a learned model.
- Would you keep the baseline features once a graph model wins?Yes, in three roles. Concatenated as inputs, since cheap strong columns rarely hurt. As a monitoring reference — if degree and clustering distributions shift, the graph itself has drifted and the heavy model is about to degrade. And as a fallback path for nodes or moments when the graph model is unavailable, which also gives you a clean answer for brand-new nodes.
- What three numbers would you require before approving the graph model?The baseline's metric under a time-split evaluation, an estimate of the achievable ceiling from an audit sample or label-noise study, and the incremental serving and operational cost of the proposal. Fund it when the baseline plateaus well below the ceiling and the gap is plausibly explained by learned neighbour interactions; otherwise the cheap path is still the right one.
saying these in an interview costs you the question
- Skips the baseline because the graph model is obviously better
- Computes structural features on the full graph, then splits randomly
- Uses neighbour label aggregates without checking label timing
- Treats the baseline as disposable once the graph model ships
- Cannot name a single per-node structural feature