skip to content

How do you report a defect that reproduces on only one platform in your support matrix?

level: middleimportance: should knowfreq 53%

answer

  1. Prove the platform is the variable
  2. Record the cells that passed too
  3. Move one axis at a time
  4. Attempts and failures, never "sometimes"
  5. Name the failing cell's tier

basics

~20 s

State the exact platform coordinates where it fails and the cells where the same steps passed, isolate one axis at a time so the report names the differing variable, give a reproduction rate, and name the failing cell's tier.

solid answer

~50 s

The report has to do one job an ordinary defect report does not: prove the platform is the variable. So it carries full coordinates for the failing cell — platform family and exact build, engine or runtime version, device class, form factor, locale, network — plus the cells where the same steps passed. Then it shows the isolation: same platform version on a different device class, same device with a different version, so the reader can see which single axis moves the outcome. Add a reproduction rate rather than "sometimes" — nine of ten attempts reads differently from two of thirty — attach evidence captured on the failing platform, and note what you ruled out: a dirty shared device, a stale build, a proxy. Finally, state the failing cell's tier and keep technical impact separate from business urgency: a total failure in a best-effort cell and a cosmetic one in a flagship cell schedule very differently.

go deeper

for a junior

Be ready to list the coordinates a platform-specific report needs — platform and build, engine or runtime, device class, orientation, locale — and to say why the cells that passed matter as much as the one that failed.

for a middle

An interviewer expects the isolation method: hold everything constant, move one axis, and report which axis changes the outcome, with attempts-and-failures counts rather than the word intermittent.

for a senior

Show how the report becomes a decision: tier of the failing cell, impact separated from urgency, evidence quality good enough for someone without that platform, and verification that spans the tier after the fix.

for a principal

Own the pattern rather than the report. Be ready to talk about what a rising count of one-platform defects says about the matrix, device pool hygiene, and where instrumentation should live so field-only reports become reproducible.

### The extra burden this report carries Any defect report needs steps, expected behaviour and actual behaviour. A platform-specific one needs a further claim: **the platform is the differing variable**. Until that is shown, the most likely reading of "only fails on that one" is that the person had a dirty device, an old build, a different account, or a different network. Reports that skip the isolation get bounced, and the round trip is expensive because the reader usually cannot reproduce it at all. ### What to capture **Coordinates of the failing cell, in full.** Operating system family and the exact build identifier, not just the marketing version. Engine or runtime version. Device class and model family. Form factor, viewport size, pixel density and orientation. Locale and timezone. Network class. Build identifier of the product under test. Account or tenant type, and any feature flag states. **The negative evidence.** List the cells where the same steps were run and passed. This is the part people omit and the part that makes the report credible. "Fails on 2 of the 14 fully supported cells; passed on the other 12, list attached" is a different artefact from "broken on tablets". **The isolation, one axis at a time.** Hold everything constant and move one thing: - same platform version, different device class → separates hardware from software - same device, different platform version → identifies a version boundary - same device and version, different engine or runtime → points at the client stack - same everything, different locale or display density → catches layout and formatting causes The axis that changes the outcome is the finding. Reporting it saves the assignee the whole first day. **A reproduction rate.** State attempts and failures, and the conditions of each attempt. "Nine failures in ten cold starts, zero in ten warm starts" is a diagnosis in itself; "intermittent" is not. If it is nondeterministic on unchanged code, say so explicitly and give the observed rate rather than letting the reader assume the platform is the cause. **Evidence captured on the failing platform.** Screen recording with timestamps, platform logs from the same window, network capture if relevant, and the state of the stored data. Evidence from the passing platform is worth attaching too, for the diff. **The cell's tier.** Say whether the failing cell is fully supported, best-effort, deprecated or outside the matrix. That is what turns the report into a scheduling decision. Keep the two judgements apart: technical impact (does the feature work at all?) is an attribute of the defect; business urgency (when must it be fixed?) is a decision informed by the tier and the affected customer segment. ### A worked example A hotel booking channel manager's front-desk console showed the previous night's rate on one browser-engine generation only, on the availability screen, after the back office pushed a change. The report read, in essence: fails on engine generation N-2 at desktop and tablet form factors, 9 of 10 attempts on a cold load, 0 of 10 after a hard reload; passes on generations N and N-1 across all six console cells, and passes on N-2 when the client's local cache is cleared first. Same account, same build, same tenant, evidence attached from both a failing and a passing run. The isolation pointed straight at a **stale-cache read** on a client-side revalidation path that behaves differently on that older engine generation, and the assignee reproduced it inside an hour. Note what made the report cheap to act on: the negative cells, the two reproduction rates that separated cold from warm load, and the cache-clear observation that named the mechanism without asserting a cause. ### Common mistakes - **Naming a device instead of a cell.** "Broken on the small tablet" leaves out version, engine, orientation and locale — all of which the reader must now guess. - **Asserting a cause.** "The platform's engine is buggy" is a hypothesis; report the observation and let the isolation carry the argument. - **Filing before eliminating the environment.** Shared device pools keep state between runs; a previous session's data explains a surprising number of "one platform only" reports. - **Conflating technical impact with urgency.** A crash in a deprecated cell may well be scheduled behind a layout defect in the flagship one, and the report should let that decision be made rather than pre-empting it. - **Not re-testing the fix on the passing cells.** A platform-specific fix has a real chance of regressing the platforms that worked, so the verification runs across the tier, not just on the cell that failed.

  • The defect reproduces on one device in a shared pool but not on the identical model at your desk. What now?
    Suspect the environment before the product. Shared pool devices keep state: leftover accounts, stored data, changed system settings, a different platform patch level, a proxy or certificate installed by an earlier run. Reset the pooled device to a known state and repeat; if it then passes, the finding is about pool hygiene, and it is worth reporting as such because it will otherwise consume triage time repeatedly.
  • How does the failing cell's tier change what you write in the report?
    It changes the scheduling ask, not the facts. In a fully supported cell you state it as a candidate release blocker and say what the release-facing impact is. In a best-effort cell you state the impact and explicitly note the tier, so the decision to defer is made knowingly. In an unsupported cell you record it and close it as out of scope, keeping it searchable in case that platform is ever promoted.
  • What do you do when the defect reproduces on a platform you cannot access?
    Get an oracle you trust before filing anything as confirmed: a customer-supplied recording with timestamps, platform logs, and the exact coordinates; or a hosted device pool session. Failing that, add targeted instrumentation and ask the reporter for one specific extra observation rather than a general retest. Mark the report clearly as unreproduced in-house so nobody reads a customer claim as verified.

saying these in an interview costs you the question

  • Names a device but no version, engine or orientation
  • Omits the cells where the same steps passed
  • Writes "intermittent" instead of a reproduction rate
  • Asserts a cause in the platform instead of reporting observations
  • Files without ruling out a dirty shared device or stale build
  • Treats technical impact and business urgency as the same field

context