skip to content

What must a decision-log row record when a risk model selects a tax return for audit?

level: juniorimportance: must knowfreq 64%

answer

  1. written at decision time, not rebuilt
  2. what was seen, what was decided
  3. inputs, versions, score, threshold, action
  4. stamp the model version in the row
  5. re-scoring today answers a different question

basics

~10 s

A decision-log row records the decision identifier and timestamp, the snapshot of inputs the model was given, the model version and feature-definition version, the score returned, the threshold in force, and the action taken.

solid answer

~40 s

Write the row on the decision path, because almost nothing in it can be reconstructed afterwards. It carries a stable `decisionId` and the decision timestamp, a snapshot of the inputs the model actually saw, the model version and the feature-definition version that produced them, the raw score, the threshold applied at that moment, and the outcome — selected for review, or not. The model version is the field candidates forget and the one that makes the rest interpretable: without it a reviewer cannot tell which scorer produced the number, and re-scoring the same return today gives a different answer because the scorer has been retrained since. The row says what was decided and what it was decided from. Why the score came out that way is a separate record and a separate question.

code

json · 20 lines
json
{
  "decisionId": "d-2026-03-14-00193f2a",
  "decidedAt": "2026-03-14T09:21:07Z",
  "subjectRef": "tok_9f31c0",
  "modelVersion": "risk-scorer:7.2.0",
  "featureDefVersion": "returns-features:2026-01-19",
  "thresholdPolicyVersion": "selection-cut:2026-02-28",
  "rawInput": {
    "grossIncome": 148200,
    "deductionsClaimed": 61450,
    "filingType": "self-employed",
    "lookedUp": { "priorAuditCount": 0, "sectorMedianDeductionRatio": 0.19 }
  },
  "features": { "deductionRatio": 0.414, "incomeBand": 5, "priorAuditCount": 0 },
  "score": 0.873,
  "threshold": 0.800,
  "decidedBy": "MODEL",
  "action": "SELECTED_FOR_REVIEW",
  "samplingRate": 1.0
}

go deeper

for a junior

Be able to name the fields out loud: decision id and time, the inputs seen, the model version, the score, the threshold, the action. The model version is the one candidates leave out.

for a middle

Explain why each field is unrecoverable later: the artifact has been retrained, the feature definition has moved, the threshold has been retuned. The row is the surviving witness to all three.

for a senior

Show that the row is written on the decision path itself, including for overrides and fallbacks, so a year of selections contains no gaps a reviewer has to guess about.

for a principal

Frame it as what the organisation must be able to answer under challenge, and price that obligation against the write path, the storage and the retention window it commits to.

## Why the row exists A risk model scores filed tax returns and selects some for audit. Months later two things can happen. One filer challenges their selection and asks on what basis it was made. Or an oversight review samples a year of selections and asks the same question of each, in bulk. Both are questions about a **past** decision, and the live system is a poor witness to it: the scorer has been retrained, the feature definitions have been revised, the upstream tables those definitions read have been refreshed, and the threshold has been retuned — often more than once. The record written at the time of the decision is what answers the question. Re-running the system today answers a different one, namely *what would we decide now*. That is the whole design constraint. Every field in the row is there because it is a value that existed at 09:21 on a Tuesday in March and no longer exists anywhere else. ## The fields, and what each one pins | field | what it pins | what breaks without it | |---|---|---| | `decisionId`, `decidedAt` | the identity and moment of one decision | an outcome that arrives later — an appeal, a closed audit — cannot be joined back to the decision that caused it | | subject reference (tokenised) | which filer the decision concerned | a challenge from one person cannot be matched to their record | | input snapshot | what the model was actually given | the decision can only be re-derived from data that has since moved | | model version | which scorer produced the score | the number is uninterpretable, and a discovered bug cannot be scoped to the decisions it affected | | feature-definition version | how inputs were turned into features | stored feature values are read under the wrong definition | | `score` | what the model returned | there is nothing to compare against the cut | | `threshold` and its policy version | the cut actually applied | a near-miss and a clear pass become indistinguishable under today's retuned cut | | action and deciding path | what the system did, and which path decided it | overrides and fallbacks are invisible and look like lost rows | ## The model version is the one people leave out The score without the version is a bare number. Suppose a reviewer sees `0.873` on a return selected last spring. Is that high? Only relative to the distribution that scorer produced, which is not the distribution the current scorer produces. Worse, when a defect is later found in one trained artifact — a feature read from a table that was empty for three days, say — the version stamp is what turns "we have a bug" into "these 4,100 decisions were made by the affected version and must be reviewed". Without it, the blast radius is the whole year. The same argument applies, less obviously, to the threshold. Thresholds are retuned far more often than models are retrained, and usually without producing a new artifact at all. A reviewer who reads the threshold from current configuration will silently misexplain last spring's decisions, showing a score of `0.84` as a comfortable pass when the cut in force that day was `0.86` and the return was selected by an override. ## Written on the decision path, not assembled afterwards The row is produced by the code that made the decision, at the moment it made it. A record assembled later by joining a request log, a feature table and a configuration snapshot is a reconstruction: it inherits whatever those sources look like now, and it has no way to represent the case where the model was consulted and then ignored. That last case matters more than it sounds. Decisions taken by a manual override, by a rules layer, or by a fallback path when the scorer was unavailable still get rows — with the deciding path named, and with the model's score included if one was produced and set aside. Otherwise a reviewer sampling a year of selections sees gaps and cannot tell a decision the model never made from a record that was lost. ## What belongs in a different record - **Why the score was what it was** — attributions, per-decision explanations, reason codes — is a separate body of work with its own record; the decision row holds inputs and outputs, not causes. - **The serving tier's access log** answers who called what and how fast. It is not the decision log and does not carry the score, the version or the cut. - **The training-run record** answers how the artifact was built. The decision row references the artifact's version; it does not restate its provenance. A good answer names the fields, explains why each is unrecoverable later, and does not try to make one record do all four jobs.

  • Why is re-scoring the return today not a substitute for having logged the score?
    Because the scorer has changed. A retrained artifact, a revised feature definition or a refreshed upstream table each give a different number today, and the reviewer is asking about a decision made under the old ones. Re-scoring answers "what would we decide now", which is a different question from "what did we decide then". Only the row written at decision time answers the second.
  • What goes in the row when the model was not the deciding path — an override or a fallback?
    The same row, with the deciding path named. Record that the decision came from the override or the fallback, which rule or default fired, and the model's score if one was produced and then ignored. Otherwise a reviewer sampling a year of selections sees gaps and cannot distinguish a decision the model never made from a record that was lost.
  • Should the threshold be looked up at review time instead of stamped on the row?
    No. Thresholds are retuned more often than models are retrained, usually without a new artifact. A cut read from current configuration will misexplain older decisions — showing a score as a near-miss when it was comfortably over the value in force that day. Stamp the threshold that was actually applied, and the policy version it came from.

A dispensing record names the prescription, the batch the tablets came from and who signed for it. Nobody doubts today's shelf; the batch number is there because next year somebody may ask about this one tablet.

saying these in an interview costs you the question

  • Says the score alone is enough, without the model version
  • Assumes re-running the model later reproduces the original decision
  • Treats the serving tier's access log as the decision log
  • Logs a row only when the model selected the return
  • Reads the threshold from current configuration at review time