skip to content

Can a graph attention model's coefficients be handed to a clinician as the explanation for a prediction?

level: principalimportance: nice to knowfreq 30%

answer

  1. allocation, not causation
  2. the deeper layer weights blended vectors
  3. many heads, many competing stories
  4. a weight multiplies a vector that may be small
  5. delete the edge and measure the change

basics

~20 s

Attention coefficients are not a causal explanation. They are per-head, per-layer relative weights over already-mixed neighbour vectors, showing where the model allocated weight rather than what changed the prediction. Validate any such claim with edge-removal counterfactuals first.

solid answer

~50 s

The story "these three prior encounters drove the prediction" is weaker than it sounds. Four problems. First, in a two-layer model the second layer's coefficients weight representations that already blend each neighbour's own neighbourhood, so a large weight on an encounter is a weight on a blend, not on that encounter's raw record. Second, a coefficient is a relative share within one neighbourhood — a heavily weighted neighbour whose transformed vector is near zero contributes almost nothing, so weight is not effect size. Third, there are as many coefficient maps as heads times layers, and picking the one that reads convincingly to a clinician is cherry-picking. Fourth, models with substantially different coefficient maps can produce the same prediction, so the map is not uniquely identified as the reason. What you can defend is a counterfactual: delete the edge, re-run, and report the measured change in output — that is falsifiable, and it is what belongs in front of a clinician.

go deeper

for a junior

Take away the one-line caution: a big attention weight means the model put a large share of one node's budget on that neighbour, and that is not the same as saying the neighbour caused the prediction.

for a middle

Be able to explain two concrete reasons the reading fails — the weight multiplies a vector whose size you have not checked, and beyond the first layer the weighted vector is already a blend of other nodes.

for a senior

Show that you would test the claim rather than present it: an edge-removal counterfactual with a measured output change, plus the caveats about correlated edges and softmax renormalization.

for a principal

Own the deployment rule. Decide what may be shown to a clinician at all, require a faithfulness number before any explanation ships, and be ready to argue for shipping no explanation over an unvalidated one.

## What a coefficient literally is One number attached to one edge, in one head, at one layer, after a softmax over that target node's neighbourhood. It is non-negative, and across the neighbourhood the numbers sum to one. That is the whole of its definition. Everything beyond it — "this encounter drove the prediction", "the model relied on this relationship" — is an interpretation layered on top, and each of the following breaks a different link in that interpretation. ## Problem 1 — depth destroys the referent In a two-layer model, the coefficients you would show a clinician sit at the second layer and weight the *first layer's outputs*. A first-layer output for encounter `j` is already a mixture of `j` and everything one hop from `j`. So a large second-layer coefficient on `j` says the model leaned on a blend that includes `j`, not that `j`'s own record mattered. The deeper the stack, the more the referent dissolves — the same reason attention maps in deep stacks generally do not read as input attributions. ## Problem 2 — weight is not effect The update is `sum_j alpha_ij z_j`. A coefficient multiplies a vector. If `z_j` has small magnitude, a coefficient of 0.7 still contributes little, and a coefficient of 0.05 on a large-magnitude vector may move the output more. Coefficients also carry no direction: they are non-negative shares, so they cannot tell a clinician whether an encounter pushed the prediction toward risk or away from it. Reporting the top three coefficients answers "where did the softmax mass land" and silently substitutes it for "what changed the answer". ## Problem 3 — many maps, one story With 8 heads across 2 layers there are 16 coefficient maps per node, and they disagree. Whoever builds the clinical view chooses one, or averages them, and the choice is made by whoever is looking at the outputs — usually after seeing which version looks plausible. That is a garden of forking paths dressed up as transparency, and it is worse than no explanation because it launders a modelling artefact into clinical language. ## Problem 4 — the map is not identified The research literature on attention and explanation makes a sharper point: it is often possible to find substantially different attention distributions that leave the model's output essentially unchanged. If two different weightings produce the same prediction, neither can be *the* reason for it. The counter-argument in that same literature is fair — attention is not *nothing*, it is a real part of the computation and constrained by training — but the defensible conclusion is that attention is a plausible hypothesis about the model's behaviour, not evidence of it. ## What is defensible **Counterfactual edge ablation.** Remove the edge, re-run the model on the modified neighbourhood, measure the change in the output. This is faithful by construction: it reports what the model does, not what a weight suggests. Two caveats to state alongside it — correlated edges mask each other (delete one of two redundant encounters and the output barely moves, even though the pair matters), and removing an edge shifts every other coefficient because the softmax renormalizes, so an ablation is not a clean single-variable intervention either. **Learned edge masks.** Methods that fit a sparse mask over edges to preserve the prediction — the family GNNExplainer belongs to — produce a subgraph whose sufficiency you can then test by the ablation above. They are optimisation procedures with their own hyperparameters and instability, so treat their output as a hypothesis with a check attached, not as ground truth. **Honest language.** "The model concentrated its weight on these encounters" is true and defensible. "These encounters drove the prediction" is a causal claim you have not established. In a clinical setting the difference is not pedantry — it determines whether a clinician can reasonably act on it. ## The organisational call The pressure to ship coefficients as an explanation is real: they are free, they are already computed, they render beautifully, and every stakeholder wants a reason next to the score. A lead's job is to insist that any artefact placed in front of a clinician carries a faithfulness check — a measured relationship between the highlighted edges and the model's actual behaviour, on a held-out sample, reported as a number. If no such check exists, the honest options are to ship the prediction without an explanation, or to ship the counterfactual instead and accept that it costs a forward pass per edge you want to interrogate. There is also a governance angle worth naming. An explanation that looks authoritative and is not faithful changes clinical behaviour in ways nobody measured — it can manufacture confidence in a wrong prediction and suppress a correct override. That failure mode is not detected by any model metric, which is why the check has to be part of the deployment, not a research nicety.

  • What would make an edge-importance claim from this model falsifiable?
    A counterfactual with a number attached: remove the edge, re-run the model, and report the measured change in the output on a held-out sample. State the caveats too — correlated edges hide each other, and removing an edge renormalizes the remaining coefficients, so the intervention is not perfectly clean. Even so, it measures the model's behaviour rather than reading a weight.
  • Is a high coefficient on a neighbour enough to say that neighbour contributed strongly?
    No. The coefficient multiplies that neighbour's transformed vector, and if the vector has small magnitude the product is small regardless. Coefficients also carry no sign, so they cannot say whether the contribution pushed the prediction up or down. Magnitude of the weighted term, not the weight alone, is the quantity that relates to contribution.
  • A stakeholder insists on shipping a reason next to every score. What do you offer instead?
    Either the prediction alone with an explicit statement that no validated explanation exists, or a counterfactual view that costs a forward pass per interrogated edge and reports a measured effect. If coefficients are shown at all, label them as where the model placed weight, never as what caused the outcome, and attach the faithfulness number you measured on held-out data.

saying these in an interview costs you the question

  • Treats attention coefficients as causal importance
  • Reports the top coefficients without checking effect size
  • Picks the head whose map looks most convincing
  • Ignores that deeper-layer weights apply to already-mixed vectors
  • Ships an explanation with no faithfulness check at all

context