How do you decide whether a defect's origin is code, data, configuration or a third party?
answer
- Ask what would have to change
- Smallest artefact whose correction removes it
- Fix location is not origin
- Same build correct elsewhere means value
- Single-valued field, written tiebreaks
basics
~20 sAsk which single artefact would have to have been different for the defect never to exist: the logic, a data set, a value applied to one environment, or a component the team does not author. That artefact names the origin.
solid answer
~50 sUse one test, applied consistently: **what is the smallest artefact whose correction makes this defect not exist?** If the logic is wrong for inputs it was meant to handle, origin is code. If the shipped artefact honours its stated contract and a data set violated an assumption nobody wrote down, origin is data. If the same build behaves correctly elsewhere and only a value applied to this environment is wrong, origin is configuration. If a component nobody here authors departed from its documented behaviour, origin is third party. Two disciplines make the answers comparable: the origin field is single-valued, and the boundary cases have a written tiebreak the whole team applies the same way. Do not let the fix location decide - the repair very often lands in the consumer while the origin sits upstream, and that mismatch is exactly what the escape field is for.
code
pseudocode · 12 linesclassify_origin(defect):
if logic wrong for inputs it was meant to handle -> CODE
if contract honoured but a record broke an unwritten -> DATA
assumption and the fix is to correct records
if same build correct elsewhere and a value applied -> CONFIGURATION
to this environment is wrong
if nothing we author is wrong, only the platform -> ENVIRONMENT
if an external component departed from its documented -> THIRD_PARTY
behaviour
if nobody ever decided the behaviour -> REQUIREMENT
# the file the fix touches never decides the answergo deeper
Be able to name the categories and give one clean example of each. Remember the one question that decides them: what is the smallest artefact whose correction would have prevented this defect?
Expect to be handed a boundary case - unhandled input, a wrong setting, an upstream change - and asked to justify a category. Say which artefact had to change and why the fix location did not decide it.
Show that you value consistency over precision: closed list, single value, written tiebreaks, periodic calibration. Be able to say what decision each aggregate would actually drive, and to reject a classification that would drive none.
Own the scheme's cost and its blast radius across teams. Decide how many categories are worth the friction, keep the field blameless so the data stays honest, and be ready to retire a distinction nobody has ever acted on.
## One question, applied every time The origin categories look obvious until you classify twenty real defects, at which point most of them sit on a boundary. What makes a taxonomy usable is not a longer list, it is **one deciding question applied identically by everyone**: > What is the smallest artefact whose correction would mean this defect never existed? That artefact names the origin. Everything below is that question applied to the categories that get confused. ## Code The logic is wrong for inputs it was supposed to handle. The specification was clear, the design supported it, and the implementation does not match. This is the default for a mis-ordered condition, an off-by-one boundary, a wrong operator, a missed case in a branch. The trap is the reverse direction: a defect fixed *in the code* is not automatically code origin. If the code correctly implements a decision that was itself wrong, the origin is requirement or design and the code change is only where the repair was cheapest to apply. ## Data The executable honours its stated contract, but a data set violated an assumption the contract never made explicit: a field arriving empty, a duplicate key, a value outside a range nobody documented, a reference pointing at a removed row. Most "bad data" defects are really two defects, and separating them is the point of the exercise. There is the specific bad record, and there is the fact that nothing rejected it. Record origin as data when the artefact that has to change is the data set; but if the same shape of input will keep arriving forever, the honest origin is usually requirement or design - nobody ever decided what the system should do with it - and the missing guard is what the escape field points at. A useful discriminator: would you fix this by editing records, or by deciding a rule? Editing records is a data origin. Deciding a rule is a requirement or design origin with a data trigger. ## Configuration The shipped artefact is correct and identical everywhere; a value applied to one environment is wrong - a threshold, a limit, a feature flag, an endpoint, a credential scope, a retention window. The tell is that another environment runs the same build correctly. Two traps here. First, a value that no code path validates is *also* a robustness weakness; keep origin as configuration and record the missing guard as the containment gap, rather than arguing the origin. Second, if the value was correct and the code read it wrongly, that is code origin - the configuration was never the thing that had to change. ## Environment Distinct from configuration: nothing the team authored is wrong at all. A platform version, a capacity limit, a clock skew, a locale or an infrastructure difference makes correct software behave incorrectly in one place. If the fix is "make the environments match" rather than "change a value we own", origin is environment. ## Third party A component nobody on the team authors behaves differently from its documented contract, or changes behaviour between versions. This is the category teams most often argue about, because the fix is nearly always local: you add a boundary check, pin a version, or adapt. Record origin as third party when their behaviour is the thing that changed, and then be honest in the escape field - "we consumed it with no boundary check" is a containment gap you own, and it is where the prevention work actually is. Note the version-change case specifically. If the component behaved as documented all along and the team's assumption about it was wrong, origin is design, not third party; nothing outside changed. ## Why the discipline matters more than the metaphysics Aggregates are the reason these fields exist. A quarter that shows twelve configuration-origin defects says invest in validating configuration at start-up; a quarter that shows twelve data-origin defects says invest in input contracts at the boundary. Both signals disappear if two people would classify the same defect differently. So: keep the field **single-valued** so counts add up; keep a **fixed closed list** rather than letting free text creep in; write down the **tiebreaks** for the recurring boundaries - fix location does not decide, missing validation stays with the artefact that was wrong, an unspecified behaviour is requirement not code; and **calibrate periodically** by having several people classify the same handful of defects independently and reconciling the differences. Consistency beats accuracy here, because a scheme applied two ways aggregates into noise that looks like data. Finally, resist the pull to record the *person* or the *team*. Every origin category is reached by a human decision. Adding who made it buys nothing actionable and reliably makes the whole field less honest.
- An external feed sent a field in an unexpected form and your ingest stored zero. Is that third party or code?Origin is third party if the feed changed from its documented behaviour, and requirement or design if the form was never agreed. Either way the escape is yours: the ingest accepted an input it could not represent and said nothing. Classifying origin outward and escape inward is what stops the record becoming a blame exchange with the provider.
- Why insist that the origin field holds exactly one value?Because its only purpose is aggregation. Multi-valued origins cannot be counted without double counting, and every defect starts to accumulate three plausible tags, which flattens the distribution and hides the signal. Keep one origin, one escape, and put nuance in the notes.
- How do you keep two people from classifying the same defect differently?Calibrate. Take a handful of recently closed defects, have several people classify them independently, compare, and turn every disagreement into a written tiebreak rule. Repeat occasionally as the categories drift. Consistency is worth more than any individual judgement being philosophically right.
It is the same question a mechanic asks about a stalling engine: is the part wrong, is the fuel wrong, is the setting wrong, or is the part someone else supplied not what the sheet said?
saying these in an interview costs you the question
- Classifying by which file the fix touched
- Calling every unhandled input a data defect
- Tagging a defect with three origins at once
- Recording the developer or team as the origin
- Blaming a third party without recording your missing boundary check
- Inventing new categories per defect instead of a closed list