skip to content

Your naive Bayes prior comes from a 90/10 label count but deployment runs 50/50 — what do you change?

level: seniorimportance: nice to knowfreq 32%

answer

  1. one additive term, not the whole model
  2. only the class mix moved
  3. log of the ratio of the two mixes
  4. identical to shifting the threshold
  5. invalid once the features drift too

basics

~20 s

Swap the prior, do not retrain. The class prior enters the score as one additive log term, so replacing the training log-prior with the deployment class mix corrects the model, provided the per-class feature distributions themselves have not changed.

solid answer

~50 s

Naive Bayes scores a class as `log P(class) + sum of log-likelihood terms`, and the prior is that single leading term, estimated by dividing each class's label count by the total. If only the class mix changed — 90/10 in training, 50/50 in production — you are in a pure prior-shift situation: the likelihoods `P(x | class)` are still valid, so refitting them would be wasted work. Add `log(new_prior / old_prior)` to each class's score, or equivalently rebuild the prior term from the deployment mix. Two caveats. First, this is only sound if the per-class feature distributions really are unchanged; if the features shifted too, the likelihoods are wrong and a new prior papers over it. Second, in a two-class model the correction is a constant shift of the log-odds, so it is mathematically the same move as changing the decision threshold.

go deeper

for a junior

Know that the class prior is just the label proportions in the training data and that it enters the score as one term. That is enough to see why a different production class mix makes the trained prior wrong.

for a middle

Be ready to write the adjustment: add the log ratio of new to old prior per class, leaving the likelihoods alone. Explain why the generative factorisation into prior times likelihood is what makes this surgical.

for a senior

Show that you check which distribution actually moved before acting, that you know the correction is equivalent to a threshold shift, and that you have a way to estimate a deployment mix you cannot directly observe.

for a principal

Own the operating policy: who monitors the class mix, how often the prior is re-estimated, and whether the correction lives in the model or in a downstream threshold — one place, so the same adjustment is never applied twice.

## Where the prior lives in the model The class prior in naive Bayes is estimated by counting labels: `P(class) = rows_in_class / total_rows`. It enters the score exactly once, as an additive term in log space: ``` log_score(class) = log P(class) + sum over features of log P(x_j | class) ``` That structural fact is what makes the fix cheap. The prior term and the likelihood terms are estimated from the training data separately and combined by addition, so one can be replaced without touching the other. ## What prior shift is Training data was 90% class A and 10% class B; production traffic is 50/50. If the *kind* of item in each class is the same as before — a class-B item still looks like a class-B item, only more of them arrive — then `P(x | class)` is unchanged and only `P(class)` moved. This is usually called prior shift or label shift, and it is the one distribution change that a generative model like naive Bayes handles gracefully, because the model factors `P(x, y)` into exactly those two pieces. Contrast it with the change that is *not* handled: if the features themselves have drifted — new vocabulary, a re-scaled sensor, a different user population — then `P(x | class)` is stale, and substituting a fresh prior would leave a wrong model with a confident new prior on top. Diagnose which one you have before reaching for the fix. ## The correction Either rebuild the prior term directly from the deployment mix, or add a per-class offset to the existing score: ``` adjusted_log_score(class) = log_score(class) + log( pi_new(class) / pi_old(class) ) ``` For the 90/10 to 50/50 case with A the majority class, class A picks up `log(0.5 / 0.9)` which is about `-0.588`, and class B picks up `log(0.5 / 0.1)` which is about `+1.609`. The net swing toward B is about 2.2 in log-odds. No counts are recomputed, no likelihoods are re-estimated, and the change is a couple of numbers in a config. ## It is the same move as a threshold change With two classes the decision depends only on the difference of the two log-scores, so adding a constant to one side is identical to moving the decision threshold by that constant. That equivalence is worth stating: a team that has already tuned an operating threshold on production data has, without naming it, absorbed part of the prior correction. Applying both a re-estimated prior *and* a threshold tuned on the shifted data double-counts the same adjustment. ## How much does it actually matter The prior is one term against a sum of many likelihood terms. On a four-thousand-token document, a swing of 2.2 in the log-odds is usually swamped and the predicted label rarely changes. On a short item with a handful of features, or where the likelihood evidence is weak and ambiguous, the prior can decide the outcome outright. So the correction matters most exactly where the model has least evidence — which is also where the errors concentrate, so it is not a negligible fix. ## When you do not know the deployment mix Common in practice. Options, roughly in order of reliability: - Label a small random sample of production traffic and estimate the mix directly. Even a few hundred items pins it down well enough for a log ratio. - Estimate it from the classifier's own predicted label distribution, corrected for the model's known error rates rather than taken at face value — raw predicted proportions are biased by the model's own asymmetric errors. - Treat the offset as a hyperparameter and tune it against the business cost of each error type on whatever labelled production data you have. And if the mix keeps moving — seasonal fraud rates, a campaign that changes traffic composition — build the re-estimation into a scheduled job rather than baking a number in once. ## The related trap A very different response to a 90/10 training set is to rebalance the training data by resampling so the classes are even. That changes the estimated prior implicitly, and it also changes which rows the likelihoods are estimated from — which is unnecessary here, since the likelihoods were fine. Adjusting an explicit prior term is the more surgical and more auditable move, and it leaves you able to change your mind when the deployment mix moves again.

  • How do you apply the new class mix without refitting anything?
    Add `log(new_prior / old_prior)` to each class's log-score, or rebuild the leading prior term from the deployment proportions. Going from 90/10 to 50/50 shifts the majority class by about -0.59 and the minority class by about +1.61, a net swing of roughly 2.2 in log-odds. The likelihood terms are untouched.
  • When is substituting a new prior the wrong fix?
    When the feature distributions moved as well. The substitution assumes P(x | class) is still valid and only P(class) changed. If the vocabulary, the sensor scaling or the user population has drifted, the likelihoods are stale and a confident new prior on top of stale likelihoods is worse than leaving it alone — that case needs refitting.
  • What if you cannot measure the deployment class mix?
    Label a small random sample of production traffic and estimate it directly; a few hundred items is usually enough for a log ratio. Failing that, infer it from the model's predicted label distribution corrected for its known error rates, since raw predicted proportions are biased by the model's own asymmetric errors, or tune the offset against the cost of each error type.
  • Why might the correction change almost no predictions on long documents?
    The prior is a single additive term competing with thousands of log-likelihood terms, so on a long document a swing of about 2.2 in log-odds is easily swamped. It bites on short items and on ambiguous ones where the likelihood evidence is nearly balanced, which is precisely where the model's errors concentrate.

The likelihoods describe what each kind of item looks like; the prior describes how often each kind walks through the door. If only the door traffic changed, you repaint the sign, not the whole showroom.

saying these in an interview costs you the question

  • Retrains the whole model when only the class mix moved
  • Says priors never matter because likelihoods dominate
  • Substitutes a new prior after the features also drifted
  • Applies a re-estimated prior and a re-tuned threshold together
  • Resamples the training rows instead of adjusting the prior term

context