skip to content

How do you probe a mild symptom for the worse failure hiding behind it?

level: seniorimportance: should knowfreq 46%

answer

  1. The symptom is one sample
  2. Push the dials around the defect
  3. Displayed, or actually stored?
  4. Follow the value into exports and totals
  5. Ask whether anything would ever notice

basics

~20 s

Treat the visible symptom as one sample of a defect, not its full extent. Vary data, timing, role and configuration around it, then follow the bad value downstream - into storage, exports and totals - to see whether something worse happens unseen.

solid answer

~50 s

The symptom you noticed is whichever consequence happened to be visible on the screen you were on, so before reporting I ask two questions. First, does the failure get worse under nearby conditions? I push the same defect with larger data, boundary and extreme values, a different role, repeated execution, concurrent execution, and a different regional or unit configuration. Second, does the wrong result escape the screen? I check whether it is written to storage, copied to another service, included in an export or rolled into an aggregate - because a wrong number that is only displayed is a nuisance, and the same wrong number persisted where nothing re-reads it is silent corruption. This work - follow-up testing - is timeboxed, and I report both the mild symptom and the worse behaviour I could demonstrate, saying which is which and what I did not check.

code

pseudocode · 12 lines
pseudocode
observed = read_from_screen(trip)          # 12.40, expected 12.47

checks = {
  "stored"     : read_record(trip).fare,             # 12.40  -> persisted
  "other_view" : read_from_supervisor_view(trip),    # 12.40  -> not a render bug
  "export"     : read_weekly_export(trip),           # 12.40  -> leaves the product
  "aggregate"  : settlement_total(week).contains(trip),
  "detected"   : any_alert_or_reconciliation_for(trip)   # none -> silent
}

blast_radius = count_records_matching(promotion_applied_after_start, last_week)
# 1,873 trips, average difference 0.07 each

go deeper

for a junior

Recall the core idea: what you saw is one visible consequence, not the whole defect. Before reporting, at minimum re-read the value from another screen to find out whether it was merely displayed wrongly or actually saved wrongly.

for a middle

Explain both probing axes - varying volume, boundaries, timing, role and configuration around the defect, and following the bad value into storage, copies, exports and totals - and be able to say why a persisted wrong value is a different defect from a rendered one.

for a senior

Demonstrate the judgement: which probe to run first, how detectability and repairability change what the defect actually costs, how you timebox open-ended probing, and how you report the mild symptom and the worse finding without blurring which evidence supports which.

for a principal

Own the systemic question this raises: which classes of defect your product cannot detect on its own, and whether the answer is more probing by testers or reconciliation and alerting so silent corruption stops depending on somebody noticing a rounding oddity.

### The symptom is a sample, not the population When a defect surfaces, what you see is not the defect - it is whichever consequence happened to be rendered on the screen you happened to be looking at, under the conditions you happened to be using. A cosmetic-looking symptom under one set of conditions can be a data-integrity failure under another, and the report that describes only the first one is very often the report that never gets fixed, because it is honestly not worth fixing as described. Follow-up testing - sometimes called corollary testing - is the deliberate probing that happens after a defect is found and before it is reported. It has two independent axes. ### Axis one: vary the conditions around the defect Stay on the same defect and push the dials near it: - **Volume and repetition** - once versus two hundred times, one record versus a large history. Failures that are harmless once are sometimes cumulative. - **Boundaries and extremes** - zero, empty, maximum length, negative, the first and last item, values just over a limit. - **Timing and concurrency** - two operations at once, an action during a background job, a slow network, a retried request. - **Identity** - a different role or permission level. A wrong value shown to its owner is different from one shown to another customer. - **Configuration** - regional and unit settings, feature switches, a different deployment target. - **Interruption and recovery** - close the page mid-operation, cancel, disconnect, then look at what state remains. ### Axis two: follow the bad value downstream This is the axis most testers skip, and it is where the nastier failure usually lives. For a wrong value, ask in order: 1. **Is it only displayed, or is it stored?** Re-read the record from a different screen and, if you can, from storage directly. 2. **Does anything copy it?** Another service's cached copy, a search index, an event another team consumes. 3. **Does it leave the product?** Exports, statements, notifications, files handed to a third party. 4. **Is it aggregated?** A wrong value that lands in a total or a payout is far more expensive than one that is visible and obviously wrong. 5. **Would anyone ever notice?** A failure the product itself detects is bounded. One that nothing detects can accumulate for months. 6. **Can it be repaired afterwards?** If the original input is gone, the recovery cost belongs in the report. ### Honest reporting of what you found Report both layers. State the symptom you actually observed, then the worse behaviour you demonstrated and the conditions it needed, and keep them clearly separated so the reader can see which claims rest on what. Say what you did not check - 'I did not test whether the export carries the same value' - because an unchecked area presented as a checked one is exactly the habit that makes the rest of your report untrustworthy. Follow-up testing is also open-ended, so it needs a timebox: probe for an agreed period, file with what you have, and leave a note about the remaining unexplored directions. ### A worked example On a ride-hailing dispatcher, a completed-trip card shows a fare of 12.40 where 12.47 was expected. As a display defect this is trivial and would sit in a backlog forever. Twenty-five minutes of follow-up work changes the picture. Re-reading the trip from the supervisor view shows 12.40 as well, so the value is stored, not merely rendered. The stored value is the one the weekly settlement sums, and nothing in the product ever compares it against the fare the rider was charged, so no screen, alert or report would ever flag the difference - it is silent. Varying the conditions shows it needs a promotion applied after the trip starts, which affected 1,873 trips in the last full week, at an average of 0.07 each. The defect is now demonstrably a settlement-accuracy problem with a measured population and no detection path, and the report has a completely different reception - not because anything was exaggerated, but because the probing found what the mild symptom was concealing.

  • How long do you spend on follow-up probing before you file, and how do you decide?
    I timebox it - twenty to forty minutes is typical - and spend the first slice on the downstream axis rather than on varying conditions, because persistence and detectability change the report's meaning most. If the value turns out to be display-only and self-correcting, I stop early and file the mild version. If it is stored and nothing reconciles it, that alone justifies extending the probe, because the cost of an undetected wrong value grows with every day it goes unfixed.
  • You demonstrate a worse failure than the one you first saw. Do you file one report or two?
    One, if the evidence points at a single underlying cause, with the mild symptom described as how it was noticed and the worse behaviour as what it actually does. Splitting the same cause into two reports invites one of them to be fixed and the other closed as a duplicate. If the probing reveals a genuinely separate cause - a second, unrelated place that also loses precision - that is a second report.
  • Why does 'nothing in the product would ever notice this' change how a mild defect should be read?
    Because detectability bounds the damage. A wrong value that a later check catches costs one reconciliation; a wrong value nothing compares against accumulates silently and is often unrepairable once the original inputs age out. Undetectability is a factual, checkable property of the system - so stating it is raising impact honestly rather than inflating it, and it is usually the single most persuasive line in the report.

A damp patch on a ceiling is the symptom; the follow-up work is going into the roof space to find out whether it is a dripping pipe or a beam that has been quietly rotting all winter.

saying these in an interview costs you the question

  • Files the first symptom without probing further
  • Assumes a wrong number is only a display issue
  • Never checks whether the bad value is stored
  • Ignores exports, copies and aggregated totals
  • Probes indefinitely and files nothing
  • Presents unchecked directions as if they were verified

context