skip to content

When a scoring service cannot log every audit-selection decision in full, how should the decision log's sampling policy be chosen?

level: seniorimportance: should knowfreq 46%

answer

  1. consequence first, volume second
  2. acted-on decisions are never sampled
  3. sample hardest just under the cut
  4. hash the decision id, not a coin flip
  5. store the rate on the row

basics

~20 s

Split by consequence before rate: decisions that actually selected a return are kept whole, the rest sampled by score band with the near-threshold band kept heavily, keep-or-drop decided by hashing the decision identifier, and the rate written on the row.

solid answer

~50 s

A single uniform rate is the wrong shape. First split by consequence: decisions that actioned something — every return selected for audit — are kept at full rate, because those are the ones that get challenged. Sampling applies to the large majority that were scored and not selected. Second, stratify by score band: the band just under the cut is sampled far more heavily than the bulk, because that is where a challenge lands and where a threshold move would change the outcome. Third, decide keep-or-drop by hashing the `decisionId` rather than drawing a fresh random number, so the sample is stable and an outcome arriving weeks later lands on a decision whose inputs were kept. Fourth, store the rate and the policy version on the row, or every count taken from the log is wrong by an unknown factor.

code

pseudocode · 11 lines
pseudocode
function shouldLog(decision, threshold):
    if decision.action == SELECTED_FOR_REVIEW:
        return { keep: true, rate: 1.0 }         // consequence: never sampled out

    if decision.score >= threshold - 0.05:
        rate = 0.25                              // sparse band just under the cut
    else:
        rate = 0.01                              // bulk of the distribution

    draw = hash64(decision.decisionId) / 2^64    // stable per decision, not per call
    return { keep: draw < rate, rate: rate }

go deeper

for a junior

Know that a high-volume scorer cannot afford to keep every decision in full, so the log retains a subset chosen by a stated rule rather than everything it sees.

for a middle

Explain stratification: the sparse band just under the decision threshold is sampled far more heavily than the bulk, because that is where a challenge and a cut change both land.

for a senior

Give both classes — actioned decisions kept whole with a durable write and reconciliation, the remainder sampled deterministically and best-effort off the critical path.

for a principal

Treat the rates as a purchased level of answerability: say which questions the chosen rates can still answer and which the organisation is choosing not to be able to answer.

## Sampling is a policy, not a rate A scorer running over every filed return produces far more decisions than you want to keep in full, and the row is not small — a raw input snapshot plus a feature vector plus versions. So something is dropped. The engineering question is *which* something, and the answer has four parts, none of which is a number on its own. Be precise about which sampling this is: the sampling rate of the **decision log**, not the sampling of training examples. They are unrelated policies with unrelated goals, and conflating them in an interview is a tell. ## Split by consequence first The population divides into decisions that did something and decisions that did not. A return selected for audit sets off a process that touches a person; a return scored 0.11 and left alone does not. The first class is the one that gets challenged, so it is kept at rate 1.0 — no sampling at all. This is usually a small fraction of the volume, which makes it affordable, but affordability is a convenience rather than the reason. Completeness here is required whatever the volume turns out to be. Sampling then applies only to the non-actioned remainder, which is where the volume actually lives. ## Stratify by score band Within the non-actioned decisions, a uniform rate spends the budget where it is least useful. The bulk of the distribution is far from the cut, highly redundant, and adequately described by 1%. The band just under the threshold is the opposite: it is sparse, it is where a filer's "why me and not them" lands, and it is the band a threshold change would move across the line. Sample it at something like 25%. | band | example rate | what the rows answer | |---|---|---| | actioned (selected) | 1.00 | the challenge to a specific selection | | within 0.05 below the cut | 0.25 | was the cut applied consistently; what a cut change would have moved | | the rest | 0.01 | the shape of the score distribution over the year | A flat 1% over everything still estimates the bulk distribution well — that is not what it costs you. What it costs is the near-threshold question, where 1% of a sparse band leaves too few rows to say anything. ## Make the sample deterministic Decide keep-or-drop as a function of the decision identifier — take a hash of `decisionId`, normalise it to the unit interval, keep it if it falls under the rate. Two properties follow, and both matter. 1. **Outcomes join.** An appeal, or a completed audit, arrives weeks after the decision. With a per-decision coin flip you get outcomes attached to decisions whose inputs were thrown away, and the join is full of holes you cannot characterise. With a hash, whether a decision is in the sample was fixed the moment it got its identifier. 2. **Re-evaluation agrees.** Any service that revisits the same decision — a retry, a shadow path, a reconciliation job — makes the same keep-or-drop choice, so you do not end up with half a row. Where you need a bounded-size sample per window rather than a fixed rate — an oversight review asking for a fixed number of cases across a year — reservoir sampling gives a uniform sample of known size without knowing the population size in advance. ## Record the rate, or the log lies quietly A sampled row that does not say what it stands for makes every aggregate wrong. If the near-threshold band was kept at 25% and the bulk at 1%, a naive count of rows over the year is biased 25-fold toward the band near the cut — and it looks like a finding. Store `samplingRate` and the sampling-policy version on each row so counts can be scaled back honestly and a policy change is visible as a change rather than as a trend. ## What the write path can afford The log write should not sit in the scoring request's critical path for the bulk of decisions: serialise a bounded payload, hand it to an append-only stream asynchronously, and on buffer pressure drop with a counter rather than block the decision. But apply different durability to the two classes. For actioned decisions the loss of a row is the failure the whole log exists to prevent, so those get an at-least-once write with retry and a reconciliation count of decisions made against rows landed; for the sampled remainder, best-effort with a dropped-rows counter is the right trade. Stating both halves is what separates a considered answer from "log it async".

  • Why hash the decision identifier instead of drawing a random number per decision?
    Because the sample has to be stable. An outcome that arrives weeks later — an appeal, a closed audit — must land on a decision whose inputs were kept, and any component that re-evaluates the same decision must make the same keep-or-drop choice. A fresh draw each time produces outcomes with no matching row and half-written records, and neither hole is characterisable afterwards.
  • What goes wrong when the sampling rate is not recorded on the row?
    Every aggregate computed from the log is wrong by an unknown factor. A reviewer counting near-threshold decisions cannot know one band was kept at 25% and another at 1%, so a scaled count is either impossible or silently biased toward the heavily sampled band — and a later policy change shows up as a trend rather than as a change in instrumentation.

saying these in an interview costs you the question

  • Applies one uniform rate across the whole score distribution
  • Samples out decisions that actually selected a return
  • Draws a fresh random number per decision instead of hashing
  • Writes every full row synchronously inside the scoring request
  • Omits the sampling rate, then counts rows as if complete
  • Confuses decision-log sampling with sampling the training data