skip to content

Which must an audit-selection decision log keep to survive a feature-definition change: the raw return figures or the computed features?

level: middleimportance: must knowfreq 57%

answer

  1. one side depends on your code
  2. field names outlive their definitions
  3. raw figures survive a redefinition
  4. the vector proves what the model saw
  5. stamp the feature-definition version too

basics

~20 s

The raw submitted figures survive, because they do not depend on your code. Keep them, keep the vector the model was actually fed as proof of what it saw, and stamp the feature-definition version linking the two.

solid answer

~40 s

A stored feature value is the output of one definition, and definitions get revised. When `deductionRatio` is redefined from deductions-over-gross-income to deductions-over-adjusted-income, last year's stored `0.41` keeps its field name and is now read under the new meaning — a silent misreading nobody's schema check catches. The raw submitted figures do not have that problem: they are what the filer sent, so either definition can be recomputed from them. But raw figures alone do not prove what the model consumed, so a bug in the feature pipeline stays invisible. Keep all three: the input snapshot as the durable evidence, the computed vector as the fidelity record, and the `featureDefVersion` stamp that says which definition turned one into the other.

code

json · 10 lines
json
{
  "decisionId": "d-2025-04-02-0071bc",
  "featureDefVersion": "returns-features:2025-03-01",
  "features": { "deductionRatio": 0.41 },
  "rawInput": {
    "grossIncome": 148200,
    "deductionsClaimed": 61450,
    "adjustedIncome": 96300
  }
}

go deeper

for a junior

Know that the figures a filer submitted and the numbers your code derives from them are different things, and that a decision log has a reason to keep both.

for a middle

Explain the failure concretely: a redefined ratio keeps its field name, so a stored value from last year is read under this year's meaning and no schema check notices.

for a senior

Describe the full snapshot — submitted figures, values looked up at decision time, the vector actually fed to the model, and the definition version that links them.

for a principal

Weigh the storage and privacy cost of retaining raw figures for years against the cost of a log that cannot be defended when a selection is challenged.

## The two candidates A scorer for tax-audit selection is handed a filed return, derives features from it, and produces a score. When you write the decision row, you can store what came in, what the model was fed, or both. The question sounds like a storage-cost trade-off. It is really a question about which of the two is defined by something outside your control. - **The raw submitted figures** — gross income, deductions claimed, filing type — are what the filer sent. Their meaning is fixed by the return form, not by your repository. They mean the same thing in five years as they did on the day. - **The computed feature vector** — `deductionRatio`, `incomeBand` — is the output of a function you wrote and will rewrite. Its meaning lives in the definition, not in the field name. ## What a redefinition does to a stored vector Take one concrete revision. Under the definition in force in early 2025, `deductionRatio` divided deductions by **gross** income. In 2026 it was revised to divide by **adjusted** income, because the sector team argued that was the comparison that mattered. | the same return | gross-income definition | adjusted-income definition | |---|---|---| | deductions 61,450 · gross 148,200 · adjusted 96,300 | `deductionRatio` = 0.41 | `deductionRatio` = 0.64 | Now consider a row from 2025 that stored only `{"deductionRatio": 0.41}`. A reviewer in 2026 reads it under the definition they know, sees 0.41, and concludes the filer's deductions were modest. They were not. Nothing signals the mismatch: the field name is identical, the type is identical, every schema check passes. And the value cannot be repaired, because 0.41 cannot be inverted back to the two numbers the other definition needs. The row that also kept `{"grossIncome": 148200, "deductionsClaimed": 61450, "adjustedIncome": 96300}` can produce either figure on demand. That is the sense in which the raw snapshot *survives* the change. ## Why the computed vector is still worth storing If raw figures are the durable evidence, why keep the vector at all? Because the raw figures tell you what the model *could* have been given, not what it *was* given. The feature pipeline sits between them, and the interesting failures are in that gap: a join that silently returned null and was filled with a default, a unit that was parsed as thousands, a lookup that timed out and fell back to a global average. Recomputing from the raw figures reproduces the intended features, not the ones that were actually scored — so the bug becomes invisible exactly when you are investigating it. So the vector is the fidelity record and the raw snapshot is the durable one. They answer different questions and neither substitutes for the other. ## The third category nobody stores Some inputs are neither submitted nor computed: they are **looked up** at decision time. A prior-audit count read from an online store, a sector median read from a reference table. These are part of what the model saw, and the store has moved on since — the prior-audit count for that filer is now 1, not 0. Recording only the request payload therefore produces a row that looks complete and is not. Put looked-up values in the same input snapshot, marked as looked up rather than submitted, so the boundary between "the filer told us this" and "we fetched this" is visible to whoever reads the row. ## What to store, in one list 1. The **raw submitted figures**, verbatim, subject to the privacy limits on what may be retained. 2. Any **values looked up at decision time**, in the same snapshot, marked as such. 3. The **computed feature vector** exactly as passed to the model. 4. The **feature-definition version** that links 1 and 2 to 3. 5. The **model version**, so the score is interpretable against the right scorer. ## The costs, stated honestly Keeping raw figures is the expensive choice on two axes. It is larger — a full return snapshot is many times the size of a short feature vector — and it is the privacy-heavy part of the row, because submitted figures are exactly the fields a retention policy is about. Teams that store only the vector usually got there by optimising one of those two, not by deciding the vector was sufficient. The answer that shows judgment names that trade-off rather than pretending it does not exist: the raw snapshot is what makes the log defensible under challenge, so it is retained under a stated window with direct identifiers tokenised, not dropped to save space.

  • What has to be true for the raw snapshot to be recomputable into features at all?
    The feature-definition version must be on the row and that definition must still be executable, and any reference data it joins against — a sector median, a prior-audit count — has to be snapshotted alongside the raw figures or pinned by version. A definition that reads a mutable upstream table recomputes to a different number every time you run it, which defeats the point of keeping the raw figures.
  • Values the scorer fetched at decision time are neither submitted nor computed — where do they go?
    Into the same input snapshot, marked as looked up rather than submitted. A prior-audit count read at 09:21 is part of what the model saw, and the store has moved on since. Recording only the request payload leaves the row incomplete in the particular way that looks complete, so nobody goes looking for the missing half.

saying these in an interview costs you the question

  • Stores only the computed vector because that is what the model saw
  • Assumes a stored feature name means the same thing every year
  • Plans to recompute old features using the current definition
  • Treats the submitted request payload as the full input snapshot
  • Keeps raw figures but never records which definition version ran