skip to content

A clinic no-show model weights 'previous no-shows' negatively - how do you investigate before launch?

level: seniorimportance: should knowfreq 44%

answer

  1. check the plumbing before the theory
  2. does one flipped label explain all of them
  3. how is the feature actually computed
  4. another predictor may be absorbing it
  5. refit and see if the sign holds

basics

~20 s

A wrong-sign weight tilts the decision boundary against domain sense, so treat it as a suspected defect: check how the label and the feature are coded and computed, whether a correlated predictor absorbs the effect, and whether the sign is stable.

solid answer

~50 s

The sign says which way that feature pushes a patient across the boundary, so a negative weight on prior no-shows means the model predicts that a history of missing appointments makes attendance *more* likely. Before theorising, check the plumbing: is the label coded 1 = no-show or 1 = attended, is the column actually 'appointments kept', and is the count computed only from history rather than including the row's own outcome. Then check the model: refit with that feature alone and see whether the direction flips, since a correlated predictor such as 'reminder calls made' — which the clinic triggers precisely because of past no-shows — can absorb the signal. Finally check stability: refit across folds and resamples and see whether the sign holds. Only after those do you decide between fixing a bug, constraining the direction, or documenting a real conditional pattern.

go deeper

for a junior

Know that each weight's sign says which way that feature pushes the prediction, and that a direction contradicting domain sense is something to raise rather than quietly accept.

for a middle

Explain the concrete checks in order: label encoding, how the feature is computed and over what time window, then a single-feature refit to see whether the direction flips only in the joint model.

for a senior

Demonstrate the diagnosis end to end, including a policy-driven feature absorbing the signal, stability across folds, and a decision that does not involve hand-editing the model.

for a principal

Own the process question: which sanity checks are launch gates, who signs them off, and when a defensible direction is worth paying accuracy for in a decision system affecting patients.

## Why the sign is worth a launch gate In a linear model each weight has one unambiguous operational meaning: it is the direction and rate at which that feature pushes the score, and therefore pushes a case across the decision boundary. Positive weight, feature up, score up, closer to being predicted positive. That makes signs the cheapest sanity check available on a model that is about to make decisions about people — you can hand the list of features and signs to a clinic manager, and they will spot in seconds that 'patients who missed three appointments last quarter are predicted *more* likely to show up' is wrong. All of what follows is debugging, not interpretation. The claim under test is 'the model's prediction direction is defensible', not 'this coefficient measures the effect of missing appointments'. ## Step 1: suspect the plumbing first Most wrong signs are bugs, and bugs are cheap to rule out. - **Label orientation.** Is the target 1 = no-show, or 1 = attended? A flipped target flips every sign in the model at once, which is the tell: if *all* the domain-obvious features point the wrong way, it is one bug, not many. - **Feature definition.** Is the column really 'previous no-shows'? Names drift. A column populated as 'previous appointments kept' or 'previous appointments booked' under a stale name produces exactly this sign, correctly. - **How it is computed.** Does the count include the current appointment's outcome, or appointments after the prediction time? A count computed over the full history including the target row is leakage; if a no-show also increments the counter, the arithmetic can invert the apparent direction. - **Missing-value coding.** If missingness is filled with zero or with a sentinel like -1 and missingness itself correlates with new patients, the weight is fitting the sentinel and not the count. ## Step 2: separate 'bug' from 'conditional structure' If the plumbing is clean, ask whether the sign is a joint-fit artefact. Fit the feature on its own against the label. If it is positive alone and negative in the full model, then something else in the model is carrying the signal and this weight is describing whatever is left over once that other predictor is held constant. The common mechanism in this scenario is a downstream feature: the clinic calls patients with a bad attendance record, so 'reminder calls made' is *caused by* prior no-shows. Once the model knows how many reminder calls a patient received, the residual role of the no-show count can genuinely flip — among patients who got the same intensive reminder treatment, the ones who earned it may attend at similar or better rates. That is a real pattern in the data and not a coding error, but it is fragile: it depends on the calling policy staying exactly as it was during the training window. ## Step 3: ask whether the sign is even real A sign is a claim about the data, and claims have uncertainty. Refit across cross-validation folds, bootstrap resamples, or a couple of random seeds and look at the distribution of that weight. If the sign flips from fold to fold, the direction is simply not determined — usually because the feature shares most of its information with another predictor, or because the informative rows are rare. A weight that is not distinguishable from zero should not be argued about; it should be reported as undetermined. Magnitude deserves the same glance. A weight that dwarfs every other in the model, on a feature that is nonzero for a handful of rows, means those rows are being decided almost entirely by that one column. That pattern is a classic leakage signature and it is also an operational risk: one bad upstream join and the score swings wildly. ## Step 4: decide, and write it down The options, roughly in order of preference: 1. **Fix the pipeline.** If it was a coding, timing or naming bug, correct it and refit. Add a regression test that asserts the feature's definition, so it cannot silently drift back. 2. **Remove the mediator.** If a policy-driven feature such as reminder calls is absorbing the effect, consider dropping it. The model gets a sane-direction weight on history, and you stop depending on a calling policy that operations may change next quarter. 3. **Constrain the direction.** Some deployments require monotonic behaviour in a named feature for reasons of policy or trust; imposing it costs a little accuracy and buys a rule you can defend. 4. **Keep it and document it.** If the pattern is real, understood, and the model is held out and validated, ship it with the explanation attached and a monitor on that weight after every refit. What is *not* acceptable is silently flipping the sign by hand, or dropping the feature purely because its direction was embarrassing. Both change the model's behaviour without changing your understanding of why the data looked that way.

  • The sign stays negative only when 'reminder calls made' is in the model. Do you ship it?
    Usually not as-is. That feature is caused by the very history you are modelling, so the model is conditioning on a consequence and its behaviour depends on the calling policy holding constant. I would fit without the mediator, compare held-out performance, and prefer the version whose directions survive an operations change. If the mediator earns its keep, ship it with the dependency documented and monitored.
  • How do you check whether the sign is stable rather than a fluke of one split?
    Refit the model across the cross-validation folds and across bootstrap resamples of the training set, and look at the distribution of that one weight. A sign that flips between fits tells you the direction is not identified — typically collinearity with another predictor or too few informative rows — and should be reported as undetermined rather than argued about either way.
  • Every domain-obvious feature in the model has the wrong sign. What does that tell you?
    That it is one bug, not many. The most likely cause by far is a flipped target encoding, where 1 means attended rather than no-show, which negates the whole weight vector at once. Check the label mapping and the class ordering used at training time before touching a single feature.

saying these in an interview costs you the question

  • Flips the sign by hand to make the model look sensible
  • Drops the feature immediately without diagnosing why
  • Assumes a wrong sign always means multicollinearity
  • Never checks the label encoding or feature definition
  • Treats a weight that is indistinguishable from zero as a finding

context