skip to content

questions

20

What must a decision-log row record when a risk model selects a tax return for audit?

level: juniorimportance: must knowfreq 64%

answer

  1. written at decision time, not rebuilt
  2. what was seen, what was decided
  3. inputs, versions, score, threshold, action
  4. stamp the model version in the row
  5. re-scoring today answers a different question

basics

~10 s

A decision-log row records the decision identifier and timestamp, the snapshot of inputs the model was given, the model version and feature-definition version, the score returned, the threshold in force, and the action taken.

solid answer

~40 s

Write the row on the decision path, because almost nothing in it can be reconstructed afterwards. It carries a stable `decisionId` and the decision timestamp, a snapshot of the inputs the model actually saw, the model version and the feature-definition version that produced them, the raw score, the threshold applied at that moment, and the outcome — selected for review, or not. The model version is the field candidates forget and the one that makes the rest interpretable: without it a reviewer cannot tell which scorer produced the number, and re-scoring the same return today gives a different answer because the scorer has been retrained since. The row says what was decided and what it was decided from. Why the score came out that way is a separate record and a separate question.

code

json · 20 lines
json
{
  "decisionId": "d-2026-03-14-00193f2a",
  "decidedAt": "2026-03-14T09:21:07Z",
  "subjectRef": "tok_9f31c0",
  "modelVersion": "risk-scorer:7.2.0",
  "featureDefVersion": "returns-features:2026-01-19",
  "thresholdPolicyVersion": "selection-cut:2026-02-28",
  "rawInput": {
    "grossIncome": 148200,
    "deductionsClaimed": 61450,
    "filingType": "self-employed",
    "lookedUp": { "priorAuditCount": 0, "sectorMedianDeductionRatio": 0.19 }
  },
  "features": { "deductionRatio": 0.414, "incomeBand": 5, "priorAuditCount": 0 },
  "score": 0.873,
  "threshold": 0.800,
  "decidedBy": "MODEL",
  "action": "SELECTED_FOR_REVIEW",
  "samplingRate": 1.0
}

go deeper

for a junior

Be able to name the fields out loud: decision id and time, the inputs seen, the model version, the score, the threshold, the action. The model version is the one candidates leave out.

for a middle

Explain why each field is unrecoverable later: the artifact has been retrained, the feature definition has moved, the threshold has been retuned. The row is the surviving witness to all three.

for a senior

Show that the row is written on the decision path itself, including for overrides and fallbacks, so a year of selections contains no gaps a reviewer has to guess about.

for a principal

Frame it as what the organisation must be able to answer under challenge, and price that obligation against the write path, the storage and the retention window it commits to.

## Why the row exists A risk model scores filed tax returns and selects some for audit. Months later two things can happen. One filer challenges their selection and asks on what basis it was made. Or an oversight review samples a year of selections and asks the same question of each, in bulk. Both are questions about a **past** decision, and the live system is a poor witness to it: the scorer has been retrained, the feature definitions have been revised, the upstream tables those definitions read have been refreshed, and the threshold has been retuned — often more than once. The record written at the time of the decision is what answers the question. Re-running the system today answers a different one, namely *what would we decide now*. That is the whole design constraint. Every field in the row is there because it is a value that existed at 09:21 on a Tuesday in March and no longer exists anywhere else. ## The fields, and what each one pins | field | what it pins | what breaks without it | |---|---|---| | `decisionId`, `decidedAt` | the identity and moment of one decision | an outcome that arrives later — an appeal, a closed audit — cannot be joined back to the decision that caused it | | subject reference (tokenised) | which filer the decision concerned | a challenge from one person cannot be matched to their record | | input snapshot | what the model was actually given | the decision can only be re-derived from data that has since moved | | model version | which scorer produced the score | the number is uninterpretable, and a discovered bug cannot be scoped to the decisions it affected | | feature-definition version | how inputs were turned into features | stored feature values are read under the wrong definition | | `score` | what the model returned | there is nothing to compare against the cut | | `threshold` and its policy version | the cut actually applied | a near-miss and a clear pass become indistinguishable under today's retuned cut | | action and deciding path | what the system did, and which path decided it | overrides and fallbacks are invisible and look like lost rows | ## The model version is the one people leave out The score without the version is a bare number. Suppose a reviewer sees `0.873` on a return selected last spring. Is that high? Only relative to the distribution that scorer produced, which is not the distribution the current scorer produces. Worse, when a defect is later found in one trained artifact — a feature read from a table that was empty for three days, say — the version stamp is what turns "we have a bug" into "these 4,100 decisions were made by the affected version and must be reviewed". Without it, the blast radius is the whole year. The same argument applies, less obviously, to the threshold. Thresholds are retuned far more often than models are retrained, and usually without producing a new artifact at all. A reviewer who reads the threshold from current configuration will silently misexplain last spring's decisions, showing a score of `0.84` as a comfortable pass when the cut in force that day was `0.86` and the return was selected by an override. ## Written on the decision path, not assembled afterwards The row is produced by the code that made the decision, at the moment it made it. A record assembled later by joining a request log, a feature table and a configuration snapshot is a reconstruction: it inherits whatever those sources look like now, and it has no way to represent the case where the model was consulted and then ignored. That last case matters more than it sounds. Decisions taken by a manual override, by a rules layer, or by a fallback path when the scorer was unavailable still get rows — with the deciding path named, and with the model's score included if one was produced and set aside. Otherwise a reviewer sampling a year of selections sees gaps and cannot tell a decision the model never made from a record that was lost. ## What belongs in a different record - **Why the score was what it was** — attributions, per-decision explanations, reason codes — is a separate body of work with its own record; the decision row holds inputs and outputs, not causes. - **The serving tier's access log** answers who called what and how fast. It is not the decision log and does not carry the score, the version or the cut. - **The training-run record** answers how the artifact was built. The decision row references the artifact's version; it does not restate its provenance. A good answer names the fields, explains why each is unrecoverable later, and does not try to make one record do all four jobs.

  • Why is re-scoring the return today not a substitute for having logged the score?
    Because the scorer has changed. A retrained artifact, a revised feature definition or a refreshed upstream table each give a different number today, and the reviewer is asking about a decision made under the old ones. Re-scoring answers "what would we decide now", which is a different question from "what did we decide then". Only the row written at decision time answers the second.
  • What goes in the row when the model was not the deciding path — an override or a fallback?
    The same row, with the deciding path named. Record that the decision came from the override or the fallback, which rule or default fired, and the model's score if one was produced and then ignored. Otherwise a reviewer sampling a year of selections sees gaps and cannot distinguish a decision the model never made from a record that was lost.
  • Should the threshold be looked up at review time instead of stamped on the row?
    No. Thresholds are retuned more often than models are retrained, usually without a new artifact. A cut read from current configuration will misexplain older decisions — showing a score as a near-miss when it was comfortably over the value in force that day. Stamp the threshold that was actually applied, and the policy version it came from.

A dispensing record names the prescription, the batch the tablets came from and who signed for it. Nobody doubts today's shelf; the batch number is there because next year somebody may ask about this one tablet.

saying these in an interview costs you the question

  • Says the score alone is enough, without the model version
  • Assumes re-running the model later reproduces the original decision
  • Treats the serving tier's access log as the decision log
  • Logs a row only when the model selected the return
  • Reads the threshold from current configuration at review time
open as a page

A mortgage underwriting scorer is retrained from an unchanged code revision a month later and comes out different - why?

level: juniorimportance: must knowfreq 52%

basics

~20 s

A code revision pins the training instructions, not the values they read: the applications the training query returns, the resolved library versions, the run-time configuration and the seed all moved. Rebuilding a model version means pinning those inputs too.

open as a page

Many advertiser tenants share one ad-screening inference tier; one tenant triples its traffic and all tenants slow down - why?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Every tenant draws on the same finite pool of workers, accelerator time and queue slots. The extra requests wait in the same line, so queueing delay rises for everyone as occupancy climbs. Nothing has to fail for this to happen.

open as a page

Which must an audit-selection decision log keep to survive a feature-definition change: the raw return figures or the computed features?

level: middleimportance: must knowfreq 57%

basics

~20 s

The raw submitted figures survive, because they do not depend on your code. Keep them, keep the vector the model was actually fed as proof of what it saw, and stamp the feature-definition version linking the two.

open as a page

Rolling the dropout early-warning model back to its previous version mid-incident: what must that rollback cover beyond the trained artifact?

level: middleimportance: must knowfreq 68%

basics

~20 s

A model version is a pin set, not a file: the artifact plus the preprocessing and feature-definition versions it reads, its decision threshold and its output contract. Roll back all of them together, by flipping one pointer to a retained, still-loadable version.

open as a page

What must a rebuild record pin, besides the code revision, to reconstruct the underwriting model version that declined an application?

level: middleimportance: must knowfreq 68%

basics

~20 s

Five pins: the code revision, the training-data snapshot identity as a content digest, the fully resolved hyperparameters and seed, an environment digest covering base image and dependency versions, and the artifact digest the run produced, all tied to one training-run identifier.

open as a page

On a shared ad-screening inference tier, what does a per-tenant concurrency quota protect that a per-tenant request-rate quota does not?

level: middleimportance: must knowfreq 56%

basics

~20 s

Occupancy. A concurrency quota caps how much of the tier a tenant holds at once, so it still binds when requests get slower. A request-rate quota counts arrivals, and arrivals stay flat while in-flight work multiplies.

open as a page

A student dropout early-warning model starts flagging the wrong cohort while error rates and latency stay flat - how do you size this incident's severity?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Severity is the count of wrong decisions already delivered and acted on, not a failure rate. Bound the start at the last known-good change, count the flags published since, and escalate on how irreversible the adviser action was.

open as a page

When a scoring service cannot log every audit-selection decision in full, how should the decision log's sampling policy be chosen?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Split by consequence before rate: decisions that actually selected a return are kept whole, the rest sampled by score band with the near-threshold band kept heavily, keep-or-drop decided by hashing the decision identifier, and the rate written on the row.

open as a page

A filer challenges an audit selection from last year — what makes the decision log credible rather than a mutable table?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Insert-only write grants with corrections appended rather than applied, each row's hash chained to the previous so an edit is detectable, the chain head anchored beyond the operators' reach, and a sequence check proving no row is missing.

open as a page

The dropout early-warning model is switched off mid-term because every version reads the same broken term calendar - what serves advisers instead?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A rules-based path built from inputs the broken one does not touch: a small deterministic score over recent attendance and failing grades, capped at the same weekly caseload and published in the same output contract, with the quality loss stated in advance.

open as a page

Given only the model version stamped on an eleven-month-old underwriting decline, how does the platform reach its training run?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Through a two-hop chain: the stamped version resolves to an immutable artifact digest, and a lineage entry written when the run finished maps that digest to its training-run identifier and pin set. A moving alias breaks the first hop.

open as a page

Why can a rebuilt underwriting model version match the original's metrics yet differ bit-for-bit from the stored artifact?

level: seniorimportance: should knowfreq 58%

basics

~10 s

Floating-point addition is not associative, so thread count, data order and accelerator reduction order shift the last bits even with every input pinned. Reproducible-in-metric is the achievable claim; reproducible-in-bits costs deliberate determinism work.

open as a page

A screening tier serves 400 tenant policy models but keeps only 48 resident in accelerator memory; why is a rarely-called tenant's end-to-end p99 measured in seconds?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Resident slots are a cache, and the rare tenant is the miss. Called once an hour, its model is evicted between calls, so nearly every request first pays a cold load - fetch, deserialise and warm a multi-gigabyte artifact - before inference starts.

open as a page

In a multi-tenant screening tier where each tenant pins one policy-model version, how should a request reach a replica that already holds it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Resolve the pinned version once at the router from a versioned routing table, then balance only across replicas that already hold that version, keeping the routing key sticky. Cold-loading elsewhere is the fallback, never a different version.

open as a page

Who should be able to pull the dropout early-warning model's kill switch, given that a switch one team alone can reach is not a switch?

level: principalimportance: should knowfreq 44%

basics

~20 s

At least one role reachable at any hour who does not need the owning team's consent - typically the serving tier's on-call - plus the accountable business owner. Pre-authorise the criteria, make the off direction cheap and the on direction gated.

open as a page

How can a year of audit-selection decisions stay reviewable without the log retaining the filers' identities?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Split the record: decision rows carry a per-subject token and get the long retention, while the token-to-identity mapping lives in a separate store with tighter access and its own clock. Most review questions never need identity at all.

open as a page

A dropout early-warning model's kill switch has not been pulled in nine months - what has probably rotted in the path behind it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Everything the switch depends on but nothing exercises: retained versions aged out, fallback thresholds calibrated on a cohort that has moved, fallback inputs whose schema changed, an output contract that advanced past it, and permissions nobody on call still holds.

open as a page

In an underwriting platform, what makes a retired model version un-rebuildable while its artifact still sits in object storage?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

Rebuildability depends on the pins outliving the artifact. Most often the training-data snapshot was deleted under a shorter retention policy; a withdrawn base image, an unavailable dependency version or a rewritten feature definition do the same.

open as a page

One tenant's policy model costs ten times more accelerator time per request; why does an equal-requests-per-tenant fair-share rule still starve the others?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Because the scarce resource is accelerator time, not requests. Counting requests charges a 400 ms inference the same as a 20 ms one, so the heavy tenant can hold most of the tier while staying comfortably inside its equal share.

open as a page