Why can a health-risk model be biased when its label is prior healthcare spend?
answer
- the target is a stand-in, not the thing
- spend records treatment, not illness
- no feature audit reaches the label
- compare the construct at equal score
- the fix changes the label, not the model
basics
~20 sSpend is a proxy for need, and the two come apart by group. A group with equal illness but less access spends less, so it scores as lower risk. The bias sits in the target definition, where no feature audit will find it.
solid answer
~50 sIn a widely reported care-management case, a model ranked patients for extra support by predicting future healthcare costs, on the reasoning that costlier patients are sicker patients. That reasoning holds on average and breaks by group: where access, trust, or insurance differ, an equally sick population generates less recorded spend, so it receives lower risk scores and less care at the same level of illness. Nothing is wrong with the algorithm — it predicts spend accurately. The construct the team wanted was *need*, and the label they used was *spend*. This is why a proxy audit on features is not enough: scan every input you like and you will not find it, because the substitution happened when the target was chosen. You detect it by comparing the proxy label against a direct measure of the construct — active chronic conditions, uncontrolled lab values — within score bands and across groups, and you fix it by changing the label.
go deeper
Remember that a model learns whatever target you hand it, and that a convenient column such as spend or clicks may not be the thing you actually care about. Ask what the label really measures.
Explain the difference between bias in the inputs and bias in the target, and why an audit of features cannot reach a distorted label no matter how thorough it is.
Show the detection method: compare a direct measure of the construct across groups at equal model score, and inspect the label's recording process for group-dependent measurement. Be able to argue for a label change over a model tweak.
Own the design-time decision. Who picks the target, where the construct is documented, what review question catches the substitution, who funds measuring the construct, and how often construct validity is rechecked as access and behaviour shift.
## Two places bias can live Most fairness tooling looks at features and at outcomes. There is a third place, and it is the one that survives every feature audit: the **label**. A supervised model learns the target you gave it. If that target is a **proxy** for the thing you actually care about — the **construct** — then any way the proxy diverges from the construct becomes something the model faithfully reproduces. Correct maths, correct code, wrong quantity. ## The case A care-management programme has capacity to enrol a limited number of patients for extra clinical support, and wants the sickest. Future healthcare spend is chosen as the target: sicker patients cost more, spend is recorded cleanly in claims data, and it is available for everyone. The model predicts next-year cost well. But spend is what the health system *did*, not what the patient *needed*. Access to care, insurance status, distance to a provider, time off work, and trust in the institution all shape spend independently of illness. Where those factors are distributed unequally, an equally sick group accumulates less recorded spend — and therefore lower predicted spend, lower rank, and fewer enrolments at the same level of illness. Measured by cost the model is unbiased. Measured by illness it systematically under-serves the group with less access. This pattern was documented in a real, widely reported health-system deployment, and it is the canonical example of **label bias**. ## Why no feature audit catches it Suppose you did everything the proxy-feature playbook asks. You removed the protected attribute, you checked that the remaining features cannot reconstruct it well, and you documented every input. None of that touches the problem, because the substitution happened before any feature was chosen. Even a model with a perfectly clean feature set will reproduce the gap, because the gap is in what the model was asked to predict. The same shape appears far outside healthcare: - **Clicks as a proxy for satisfaction.** Groups differ in how they engage, so ranking by clicks does not rank by usefulness. - **Past promotions as a proxy for potential.** If promotion decisions were themselves uneven, the label encodes that history as ground truth. - **Customer-service ticket counts as a proxy for problems.** People who do not complain look like people with no problems. Each one is a legitimate-sounding operationalisation that fails when the recording process differs by group. ## How you detect it The test is not a model test; it is a **construct-validity** test: 1. Get a direct measure of the construct for at least a sample — for health, the number of active chronic conditions or uncontrolled lab values; for satisfaction, a survey. 2. Hold the model's score fixed and compare the construct across groups **within score bands**. If patients at the same predicted risk have systematically more active conditions in one group, the label is displacing need by group. 3. Check the label's own generating process for group-dependent recording: does access, reporting propensity, or measurement frequency differ? Step 2 is the one people miss. Comparing the construct across groups overall confounds real prevalence differences with label bias; comparing it *at equal score* is what isolates the substitution. ## How you fix it The fix is a label change, not a model change. Options in rough order of preference: predict the construct directly where it is measurable; predict a composite that is closer to the construct (an illness-burden index rather than cost); or predict the proxy but recalibrate against the construct within groups. Every one of these costs something — data collection, a noisier target, worse apparent accuracy — and each is a decision someone has to authorise. ## The principal-level content This is a governance question more than a modelling one, which is why it reads as a leadership interview item. - **Who chooses the target?** In practice a target is picked early, often by whoever knows the data warehouse, on the grounds that the column exists and is clean. Availability is the strongest and worst reason to choose a label. - **Where is the construct written down?** A model card or design document should state the construct in words, the proxy used, and the known ways they diverge. If nobody can articulate the construct, nobody can notice the substitution. - **What does the review ask?** Add one question to model review: *what is the label, what do we actually want, and how could those two come apart by group?* That single question would have caught the healthcare case at design time. - **Who pays for the better label?** Measuring the construct usually requires collection nobody has budgeted. Making that an explicit, priced tradeoff is the leadership act. - **What is the standing check?** Because divergence between proxy and construct can grow as access or behaviour changes, the construct-validity comparison belongs on a periodic review, not only at launch. A strong answer names the mechanism, distinguishes label bias from feature-proxy bias explicitly, gives the equal-score comparison as the detection method, and lands on the organisational point: the most consequential fairness decision in most projects is made in the first week, when somebody picks the column to predict.
- How would you detect label bias when you have no direct measure of the construct?Buy a sample of it. Collect the construct for a few hundred cases through survey, chart review, or expert adjudication, then compare it across groups within score bands. Absent that, reason about the recording process — who gets measured, how often, and by whom — and document the suspected divergence as a known limitation rather than asserting the proxy is fine.
- Would predicting the construct directly always be better than predicting the proxy?Not always. Directly measured constructs are often scarcer, noisier, and available only for a biased subsample of cases — which can import a fresh selection problem. The judgment is whether the construct label's noise costs less than the proxy label's group-dependent distortion, and that comparison should be made and written down explicitly.
- What single question would you add to model review to catch this class of problem?"What is the label, what do we actually want to know, and how could the two diverge by group?" Asked at design time it is cheap; asked after deployment it is a rebuild. Pair it with a requirement that the construct and the chosen proxy both appear in the model's documentation.
saying these in an interview costs you the question
- Says a proxy label is fine because the model predicts it accurately
- Looks only at features when auditing for bias
- Compares the construct across groups without holding the score fixed
- Picks the target purely because the column is available and clean
- Treats the fix as a model change rather than a label change