skip to content

Defect Management

Where a defect goes after anyone finds it, in a test run or in the field: writing it up, ranking its urgency, moving it through triage, and reading a quarter of them as process data.

on this pageshow

questions

page 1 of 2

What must a defect report contain for someone else to reproduce and fix it?

level: juniorimportance: must knowfreq 74%

answer

  1. Someone else must see it again
  2. Two jobs: reproduce, and decide
  3. Which build, which environment, which data
  4. Two statements, not one sentence
  5. Evidence: logs, traces, identifiers, timestamps

basics

~10 s

An unambiguous one-line title, the exact build identifier and environment, numbered preconditions and steps, the expected result and the actual result written separately, and evidence such as logs, screenshots or traces.

solid answer

~40 s

Every field serves one of two jobs: letting a stranger reproduce the failure, or letting a triager decide what to do without asking you. For reproduction you need the exact build string, not "latest"; the environment and data state, including whether the tier is shared; numbered preconditions and steps with one action each; and the expected and actual results as two separate statements, with the expected pointing at the requirement or rule it comes from. For the decision you need evidence - the log window around the failure, a stack trace, a correlation identifier, a timestamp with time zone, a cropped screenshot - plus enough about impact and frequency to weigh it. Before filing, search existing reports, including closed ones, for the same symptom. File the observation, not a guess at the cause.

go deeper

for a junior

Be ready to list the fields from memory - title, build, environment, preconditions, numbered steps, expected, actual, evidence - and to walk through filing one small real defect using all of them.

for a middle

Explain why each field exists: which ones serve reproduction and which serve the triage decision, and what specifically goes wrong downstream when the build identifier or the expected result is missing.

for a senior

Show judgement about evidence: what to attach and what to redact, which identifiers make a failure findable in aggregated logs, and how you stop a shared tier from becoming an unstated precondition.

for a principal

Own the cost argument - round trips, unreproducible reports and duplicate filings measured against a fixed release cadence - and be able to say which reporting habits you would hold teams to and which you would leave to judgement.

## What the report is actually for A defect report has two jobs, and every field earns its place against one of them. The first is **reproduction**: someone who did not see the failure must be able to see it again. The second is **decision**: whoever triages it must be able to judge impact and route the work without opening a conversation with you. A field that serves neither is ceremony; a missing field that serves either one is a round trip. On a 3-week release train, one "cannot reproduce, please clarify" round trip can push a fix past the branch cut and into the next train, which is why an unreproducible report often costs more than no report at all. ## The one-line title The title is the only part most readers ever see - in a triage list, in a search result, in a scan of what changed. An unambiguous title names three things: the artefact, the trigger, and the observed symptom. "Export broken" names none of them. "Points ledger export truncates at 8,192 rows when a member has more than one tier history" names all three and is findable by the next person who hits the same symptom. Prefer the system's own vocabulary - the screen name, the job name, the literal error text - over your paraphrase, because that is what someone will type into the search box. ## Build and environment "Latest" is not a build identifier. Record the exact version string of the component under test, plus anything else in the picture whose version can move: the client, the platform, the data set, the configuration flags. Record which deployed tier you were on and whether it is shared - on a shared tier, another person's actions are silently part of your preconditions, and that alone explains a great many "works on mine" arguments. Locale, time zone and clock matter whenever formatting, scheduling or expiry is involved: a points balance that expires at midnight behaves differently depending on which midnight. ## Preconditions and numbered steps Preconditions are the state the system must be in before step 1: which account, which data, which flags, which prior activity. Steps are numbered, one action each, and phrased so that two people executing them do the same thing. "Log in as a user with some points" is not a step; "Open the ledger for a member holding 4,180 points with 3 pending accruals" is. Numbering is not decoration - it lets the developer and the confirming tester say "it diverges at step 6" instead of re-describing the whole path. ## Expected versus actual Write them as two separate statements, never as one sentence joined by "but". The expected result is your oracle - the thing that says the behaviour is wrong - and it should point at its source: a requirement, an acceptance criterion, a documented rule, or a consistent behaviour elsewhere in the product. The actual result is what the system did, quoted verbatim where possible, including error text and codes. Separating them separates two different arguments: a reader can dispute your expected result, which is a requirements conversation, without disputing your observation, which is a facts conversation. Reports that merge the two tend to lose both. ## Evidence Evidence shortens diagnosis and proves you saw what you say you saw. Useful evidence is specific: the log window around the failure rather than the whole file; a stack trace or error identifier; a correlation identifier that lets the developer find the same request in aggregated logs; a timestamp with its time zone; a screenshot cropped to the wrong value but with enough surroundings to locate it; a short recording when the failure is about sequence or timing. Redact secrets and personal data as you attach. ## A worked example On build 7.4.219, the nightly accrual job for a loyalty-points ledger failed part-way through. The weak report is "accrual job crashes". The actionable one names the build, states that the job ran against the 12,400-member fixture set on the shared integration tier, gives the numbered steps that trigger a run, states expected ("all 12,400 members accrued, job exits successfully") and actual ("job aborts at member 1,873; log shows the connection pool at its ceiling of 24 with none released"), and attaches the log window plus the pool metric for the ten minutes before the abort. That report carries a diagnosis for free: it is a resource exhaustion, and the reader can see the resource. ## What makes a report worthless A report nobody can reproduce consumes triage time, is closed unresolved, and if the defect is real it comes back later with less context. Two habits prevent most of that. Search for an existing report before filing, using the symptom's own words and the literal error text, and include closed reports. And file the observation, not your guess at the cause: a suspected root cause in the title anchors everyone on the wrong component, and it is the easiest way to make a well-evidenced report useless.

  • Before filing, how do you check whether the same defect is already reported?
    Search with the symptom's own vocabulary rather than your title - the literal error text, the screen or job name, an identifier from the trace - and include closed and rejected reports, not just open ones. Widen to the affected component if the exact wording finds nothing. When you find a match, add your build, environment and evidence to it rather than opening a second report; a closed "could not reproduce" that suddenly acquires a reliable reproduction is far more useful than a fresh duplicate.
  • What do you do when the expected result is not written down anywhere?
    Say so explicitly in the report. State what you observed, what you believe the behaviour should be, and the basis for that belief - a consistent behaviour elsewhere in the product, an established convention, a plain user expectation. Filing it as a question about intended behaviour is legitimate. Filing it as a defect while hiding that the oracle is your own judgement is not, because triage will treat your assumption as a requirement and argue about the wrong thing.
  • Why is putting your suspected root cause in the title risky?
    Because the title is what everyone reads and searches, so a wrong guess anchors the whole thread on the wrong component and the report gets routed to a team that cannot act on it. It also makes the report unfindable by anyone who hits the symptom, since they will search for what they saw, not for your theory. Put the observable symptom in the title and your hypothesis in the body, clearly labelled as a hypothesis with the evidence behind it.

A defect report is a recipe, not a restaurant review: the reader has to be able to cook the same dish and get the same burnt result.

saying these in an interview costs you the question

  • Files a title like "it does not work"
  • Says the build is "latest" instead of a version
  • Writes the steps as prose with no numbering or preconditions
  • Gives only the actual result and leaves expected implied
  • Attaches no evidence because the failure looked obvious
  • Puts a guessed root cause in the title

context

open as a page

What is defect density, and why does the denominator you choose change the answer?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Defect density is a count of confirmed defects divided by a size unit — a thousand lines of code, a functional size unit, a module, a feature. Change the unit and the same product scores differently.

open as a page

What is the difference between a defect's origin phase and its escape phase?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Origin phase is where the defect was introduced - the requirement, design, code, data or configuration work that created it. Escape phase is the last check that should have caught it and did not. Every defect has both.

open as a page

In a tracked defect's explanation, how do the symptom, the condition that produced it, and the weakness that let it reach a user differ?

level: juniorimportance: must knowfreq 64%

basics

~20 s

The symptom is what was observed. The condition that produced it is the specific state or input that made the code fail that time. The weakness is the missing guard or assumption that allowed the whole family of failures.

open as a page

What is the difference between a defect's severity and its priority?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Severity measures a defect's technical impact — how badly the product's behaviour breaks. Priority measures business urgency — how soon it should be fixed relative to other work. Different people set them, and the two values move independently.

open as a page

Walk through the states a tracked product defect passes through, and who moves it between them?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A tracked defect moves through new, triaged, assigned, fixed, verified and closed, with side exits to rejected, deferred and reopened. The reporter files it, triage accepts and assigns it, a developer fixes it, and someone other than the fixer verifies it before close.

open as a page

How do you choose between a why-chain, a fishbone diagram and a fault tree for one tracked product defect?

level: middleimportance: must knowfreq 58%

basics

~20 s

Match the technique to what you already know. A why-chain fits a single causal line you can walk back, a fishbone diagram fits an unknown cause spread across categories, and a fault tree fits a failure several routes can each produce.

open as a page

Which fixed defects earn a root-cause investigation, and why do blanket always-or-never rules fail?

level: middleimportance: must knowfreq 58%

basics

~20 s

Escaped, repeating, costly and undetectable defects earn an investigation; a typo caught in review does not. Blanket rules fail alike: investigating everything turns the work into a form nobody reads, investigating nothing pays for the same weakness again.

open as a page

Which facts make a low-severity defect a higher priority than a crash?

level: middleimportance: must knowfreq 71%

basics

~20 s

Reach, frequency and the absence of a workaround. A cosmetic defect on a path every user crosses, hit on every attempt, with nothing the user can do about it, outranks a crash reachable only through a configuration almost nobody runs.

open as a page

When a defect report is rejected, how do not a bug, duplicate, cannot reproduce and works as designed differ?

level: middleimportance: must knowfreq 61%

basics

~20 s

They assert different things: not a bug means the product behaved correctly and the report was mistaken; duplicate means another open record already covers it; cannot reproduce means nobody could make it happen again; works as designed means the behaviour matches the specification, so changing it needs a specification decision.

open as a page

A schema-drift mismatch in a fleet telematics ingest reached the field. How do you judge which phase should have contained it?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Walk the chain of checks the change passed through, find the earliest one that could cheaply have caught it, and decide whether that check was absent, wrong, or skipped. That answer, not the phase that reported it, is the escape phase.

open as a page

How do you cut a defect down to a minimal reproduction before filing it?

level: middleimportance: should knowfreq 55%

basics

~20 s

Remove or simplify one factor at a time - steps, data, flags, concurrency - resetting and re-running after each change. Keep removals that still fail, restore the ones that stop it, and finish when nothing else can go.

open as a page

How do you decide whether a defect's origin is code, data, configuration or a third party?

level: middleimportance: should knowfreq 44%

basics

~20 s

Ask which single artefact would have to have been different for the defect never to exist: the logic, a data set, a value applied to one environment, or a component the team does not author. That artefact names the origin.

open as a page

A defect is closed by rejecting the one input value that produced it. Why does the rest of the defect family stay open?

level: middleimportance: should knowfreq 52%

basics

~20 s

Rejecting one value removes one occurrence, not the reason that value was dangerous. The weakness stays: the boundary nobody checked, the assumption two components did not share. So a neighbouring value, or the same value by another route, still fails.

open as a page

Which people must attend a root-cause session for a tracked defect, and what does an analysis based only on the defect record produce?

level: middleimportance: should knowfreq 38%

basics

~20 s

Invite whoever can show the material that settles a link: the person who made the change, whoever can reproduce the failure, whoever owns the data or configuration involved. Without them the session produces the most plausible story, not an explanation.

open as a page

How do you write up an intermittent defect that reproduces only some of the time?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Report failures over a stated number of attempts on a named build, the conditions that raise or lower that rate, evidence captured from a failing run, and what you ruled out. "Fails sometimes" is not a report.

open as a page

A defect arrival curve flattens two weeks before release. How do you tell stability from exhausted testing?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Normalise arrivals by test effort before believing the curve, then probe with something the suite has never exercised. A flat curve over falling effort, or over a suite that has already found everything it can, is not evidence of stability.

open as a page

A defect can be fixed at the failing value, at a missing guard, or in the design behind both. How do you choose the layer?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Choose by reach, cost and how soon each pays back. Refusing the failing value is cheap but covers one occurrence; a guard covers the family by the next release; a design change removes the class over months.

open as a page

Two independent routes can each produce the same defect on their own. How does decomposing it as alternatives change what counts as fixed?

level: seniorimportance: should knowfreq 31%

basics

~20 s

Each route reaches the same failure on its own, so closing one leaves the defect reachable. Decomposing as alternatives makes the other routes explicit, forces verification along each, and exposes any shared step where one change closes them all.

open as a page

How do you spot a class of related defects, and what does investigating the class once give you?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Group defects by cause shape (same area, same missing check, same escape route), never by shared symptom. One pass forces an explanation that must account for every member, a far harder test than explaining one.

open as a page

Should a defect's root-cause investigation run before the fix, after it, or after release, and who owns it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Perishable evidence decides. Investigate before the fix only when the failing state is the evidence; otherwise preserve it, fix, then explain, because the change the fix required is itself evidence. Ownership follows the evidence: the fixer plus the context holder.

open as a page

How do you argue a priority escalation for a defect product ranked low?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Bring evidence, not adjectives: measured reach, failure frequency, who performs the workaround and its recurring cost. Ask to rank above one named competing item, and if the answer stands, record the accepted risk and the trigger that reopens it.

open as a page

A fix lands for a tax-filing wizard defect that only reproduces on the 6-hour nightly run. How do you confirm it before closing?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Agree a confirmation criterion before the fix lands, replay the reported steps on a named build containing the fix, and get enough evidence that a pass means something — a forced-condition harness or a stated number of clean cycles — verified by someone other than the fixer.

open as a page

Leadership wants a per-team defect density target. How do you set defect measures without distorting behaviour?

level: principalimportance: should knowfreq 40%

basics

~10 s

Refuse the single per-team target and offer paired measures instead: every efficiency number gets a counter-measure, definitions are published and frozen, trends beat thresholds, and nothing is wired to ranking or pay.

open as a page

When a defect class keeps recurring, how do you choose between a regression case, a static rule, and a design change?

level: principalimportance: should knowfreq 41%

basics

~20 s

Pick the highest rung on the containment ladder the class can afford: make it unrepresentable by design, else catch it mechanically with a static rule if it has a machine-checkable signature, else pin one behaviour with a regression case. Accepting and monitoring is a legitimate fourth answer.

open as a page

Your team writes a cause paragraph on every defect and nothing changes. What policy do you set instead?

level: principalimportance: should knowfreq 33%

basics

~20 s

The paragraph is the wrong output. Either narrow the funnel to few investigations, each owned, timeboxed and ending in one committed change, or keep breadth and cut depth with a two-field note feeding a monthly review of defect classes.

open as a page

How do you decide which open defects ship as known issues rather than block a release?

level: principalimportance: should knowfreq 44%

basics

~20 s

Deferral is an explicit, owned decision, not a queue that fills by default. Ship a defect as a known issue only when its harm is bounded and reversible, a workaround or accepted-risk statement exists, an owner and revisit point are named, and support and users are told.

open as a page

How is defect removal efficiency calculated, and what does a low value for one phase tell you?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Defect removal efficiency is the share of the defects present at a stage that the stage actually caught: defects found there, divided by that count plus the ones that escaped it and were found later, as a percentage.

open as a page

What does Orthogonal Defect Classification add over a free-text cause field on a defect record?

level: middleimportance: nice to knowfreq 19%

basics

~20 s

It replaces prose with a few independent attributes drawn from fixed short lists - what kind of defect, whether something was missing or incorrect, what surfaced it, what it affected - so defects can be counted and compared instead of only read one at a time.

open as a page

showing 1–30 of 31