Your handheld target list has not changed in two years. How do you decide what it should run on now?
answer
- The list is a decision, not furniture
- Weight the field by real usage
- Which rows failed differently?
- Re-cutting resets comparability
- Fixed cadence or stated threshold
basics
~20 sRe-cut it from field evidence rather than habit: usage-weighted share across version lines, capability tiers and screen classes, and which rows have ever caught a failure nothing else caught. Every re-cut costs comparability, so change it on a decided cadence.
solid answer
~50 sStart by separating two decisions that get confused. What the product commits to supporting is a product decision made elsewhere; the target list is which of those combinations the suite actually runs on, and that one is yours. Then bring evidence to it. Pull the field distribution across all three axes — version line, capability tier, screen class — weighted by real usage rather than device count, because a row with few devices and heavy use matters more than a long tail of idle ones. Cross that with which rows have ever produced a failure no other row produced. A popular row that never fails differently costs run time and teaches nothing; a small-share row that repeatedly finds unique failures earns its slot. The trade-off is **stability against relevance**: every re-cut resets the history you compare runs against, so re-cut on a stated cadence or threshold, not opportunistically.
code
yaml · 13 linestarget_list:
revision: 7
cut_on: "2026-02-14"
evidence_window: "90 days of field usage"
core_rows: # every change; kept stable so trends stay comparable
- { line: v11, tier: mid, screen: narrow, usage_share: 0.31, unique_finds_12mo: 4 }
- { line: v12, tier: high, screen: wide, usage_share: 0.22, unique_finds_12mo: 1 }
wide_pass_rows: # periodic, not per change
- { line: v9, tier: low, screen: narrow, usage_share: 0.07, unique_finds_12mo: 6 }
moved_out_of_core:
- { line: v8, reason: "usage share under 1% across two windows; product support unchanged" }
review_policy:
trigger: "any row crosses 5% usage share in either direction, or every third release"go deeper
Know that a target list is a decision someone made from evidence at a point in time, not a fixed property of the product, and that the field it was drawn from keeps moving after it is written.
Be able to say what evidence would justify a change: the share of real usage on each version line, capability tier and screen class, and whether a row has ever caught something no other row caught.
Show how you would run the re-cut in practice — pulling a field distribution over a stated window, weighting by usage rather than device count, and confirming a trend across windows before moving anything.
Own the trade-off. A stable list keeps results comparable and cheap to reason about; a list that tracks the field stays relevant but resets your baselines. Defend the cadence or threshold you chose, and keep a fixed core so trends survive the change.
A target list that has not moved in two years is not stable, it is unowned. The field underneath it has moved on all three axes — version lines, capability tiers, screen classes — and the question is what evidence should move the list, and how often you are willing to pay for moving it. ## Two decisions, not one - **What the product supports** is a commitment to users about which devices and versions will work. That is decided outside the suite, and it changes rarely and deliberately. - **What the suite runs on** is a subset of that commitment chosen for information value and cost. It can and should change more often. Confusing the two produces both failure modes at once: teams add every supported combination to the run list because "we support it", and teams refuse to touch the run list because they think changing it changes the promise. Keep them separate and each becomes tractable. ## The evidence that should move a row | Signal from the field | What it argues for | The caveat | | --- | --- | --- | | A version line's usage share has fallen across two windows | move it from every run into a periodic wide pass | share can be seasonal; look at a trend, not one week | | A new version line is appearing and climbing | add it before it is the majority, not after | early adopters are unrepresentative of who follows | | A capability tier is over-represented in real failure reports | strengthen that row, or split it | a spike can be one bad build rather than the tier | | A row has produced no unique failure in a year | it is a candidate for lower frequency | check whether it ran the cases that could have failed | | Displays of a new geometry are shipping in volume | add a screen row that crosses the layout switch | geometry only needs a row if an arrangement actually changes at it | Two things are worth naming explicitly. First, weight by **real usage, not device count** — sessions, or whatever unit reflects exposure — because a row that is 2% of devices and 15% of activity is not a 2% risk. Second, **popularity and information are different arguments.** A hugely popular row that behaves identically to another row in the list adds cost and no signal; you keep it for the reassurance, not for the coverage, and you should know which of those you are buying. ## Cadence or trigger There are two defensible policies and the conditions pick between them: 1. **Fixed cadence** — re-cut every two or three releases regardless. Cheap to run, predictable, and it forces someone to look. It suits a product whose field moves smoothly and whose team would otherwise never revisit the list. 2. **Stated threshold** — re-cut when a signal crosses a line you wrote down in advance (a row drops below some usage share for two consecutive windows; a new line passes some share). It suits a field that moves in jumps, such as a market where one device family dominates, but it needs the thresholds agreed before the data arrives, or the argument becomes about the threshold instead of the row. What is not defensible is re-cutting opportunistically after each embarrassing escape, because that ratchets the list upward forever: rows are added by incident and never removed by evidence. ## What a re-cut costs Changing the list breaks comparability. Pass rate, duration and flakiness trends are only meaningful across the same targets, so a revision resets the baseline you were reading. The mitigation is structural rather than heroic: - Keep a **small fixed core** that changes rarely, and let the remaining rows rotate through a periodic wide pass. Trends live on the core. - **Version the list** with a revision number and the date and window of the evidence behind it, and record which revision produced each set of results, so a change in pass rate can be attributed to the list rather than blamed on the product. - Record **why each row exists** — the axis it covers and the failure class it is there to catch. A row whose reason nobody can state is the cheapest thing in the list to remove. ## The axis that quietly stops being true Version lines announce themselves and displays are visible, but capability drifts silently. A row labelled "low tier" that was chosen three years ago may now be an ordinary mid-range device: it still runs, it still passes, and it no longer produces the memory-pressure and thermal failures it was added for. Re-cutting means re-checking what each label denotes, not just re-checking the shares next to it.
- A row has a large share of real usage but has never produced a failure no other row produced. Do you keep it?Usually keep it, but in a periodic pass rather than every run — and check first whether it has never failed because it is genuinely equivalent to another row, or because few cases ever ran there. Popularity justifies exposure; information justifies the per-change slot. Conflating the two is how lists grow without getting better.
- How do you stop a re-cut from destroying your ability to compare runs over time?Keep a small core of rows fixed across revisions and rotate the rest, version the list with a revision number and date, and record which revision produced each run's results. Then a drop in pass rate can be attributed to the list change rather than misread as a product regression, and the core keeps a continuous trend line through the revision.
- Which axis is most often left un-recut, and why does that matter?Capability. Version lines announce themselves and screen geometries are visible, but hardware tiers drift quietly: the low row chosen three years ago is now mid-range. It still runs and still passes, while no longer producing the memory-pressure and thermal failures it was added to catch — so the list looks complete while a whole failure family has stopped being covered.
saying these in an interview costs you the question
- Keeps rows because they have always been there
- Counts devices instead of weighting by real usage
- Adds every new version line and removes nothing
- Treats the run list as the support commitment
- Re-cuts opportunistically after every escaped defect
- Assumes a low-tier row stays low forever