How do tags on scenarios keep feedback fast on a large scenario suite?
answer
- Labels the runner can select on
- One suite, several triggers, different subsets
- Few orthogonal axes beat many ad-hoc labels
- A selection matching nothing still reports green
- Selection and parallel running are complements
basics
~20 sTags are labels the runner selects on, so each pipeline stage runs only what it needs: a fast subset per change, the full pack nightly. The value comes from a small, enforced taxonomy, not the mechanism.
solid answer
~40 sThe runner takes a selection expression over the labels attached to scenarios, so one suite can serve several triggers: a commit stage runs the fast tier while excluding anything needing an external dependency, a change-scoped stage adds the capability area that was touched, and the nightly run takes everything. The mechanism is trivial; the discipline is the taxonomy. Keep two or three orthogonal axes — tier, dependency, capability — and refuse labels that rot, such as team, sprint or work-item identifiers. Enforce exactly one tier tag per scenario in the build, or new scenarios silently fall out of every fast subset. Guard against the worst failure by asserting a minimum matched count per stage: a renamed or misspelled tag otherwise runs zero scenarios and reports green.
code
pseudocode · 10 linescommit_stage:
select: tier:fast and not needs:external
require_min_matched: 34
budget_seconds: 300
change_scoped_stage:
select: (tier:fast or capability:earn) and not needs:external
nightly:
select: allgo deeper
Know that scenarios carry labels and that a run can be filtered to a subset by those labels, so not every change has to wait for the whole pack. Be able to say why a fast subset exists.
Explain the selection expression including negation, propose a small taxonomy with orthogonal axes, and name the enforcement that keeps it honest: one tier tag per scenario, checked in the build.
Bring the failure modes. Describe a stage that went green on zero matched scenarios after a rename, and the minimum-matched-count guard; explain why parallel execution needs per-scenario data before it buys anything.
Own the budget policy: what wall-clock target each trigger has, who approves an addition to the fast tier, and how you keep selection from becoming a place where inconvenient scenarios quietly go to die.
## What a tag is, at run time A tag is a label attached to a scenario, or to a group of scenarios, that the runner can select on. At run time you give the runner a selection expression over those labels — one tag, a union, an intersection, or a negation — and it executes only the matching scenarios. That is the whole mechanism. Everything interesting about tags is *policy*: which labels exist, who applies them, and which selection each pipeline stage runs. Selection matters because a plain-language suite has one wall-clock number and several audiences. A developer waiting on a commit build wants an answer in minutes. A release decision wants the whole pack. Without selection you get one setting for both, and it is always the wrong one: either the commit build is too slow to be run on every change, or the release gate is too thin to trust. ## A taxonomy that survives contact The common mistake is to tag by whatever seemed salient the day the scenario was written — the team, the sprint, the person, a work-item identifier. Those labels rot within a quarter and no stage can be built on them. A taxonomy worth keeping has a small number of orthogonal axes, each answering a question a pipeline stage actually asks: - **Tier** — is this in the fast subset or only in the full pack? One tier tag per scenario, always present, is the axis that stages are built on. - **Dependency** — does this scenario need something the commit stage does not have: an external sandbox, a message broker, a scheduled job, a seeded data set? Stages select with a negation on this axis rather than by guessing. - **Capability or risk area** — earn rules, burn rules, expiry, reversal. This axis is what lets you run "everything touching the ledger" after a change to the ledger. - **Lifecycle** — a temporary label for scenarios that are being migrated or that need a data reset, with an expiry date agreed when it is applied. Two or three axes is plenty. Every extra axis multiplies the ways a scenario can be mislabelled. ## Mapping the axes onto stages For the eleven-person team owning a loyalty-points ledger, the full pack of 340 scenarios ran in about 47 minutes, which is far too slow for a commit build. Their commit stage selects the tier tag for the fast subset and negates the external-dependency tag, which yields 38 scenarios in 4 minutes 12 seconds. The change-scoped stage adds the capability axis, so a change to the earn rules runs those 38 plus everything labelled with the earn capability. The nightly run takes the whole pack with no selection at all. The scenarios did not change; only which of them each trigger runs did. ## The failure modes, and the guards for them **A selection that matches nothing and reports success.** This is the dangerous one. A tag is renamed, or misspelled in the pipeline configuration, and the stage executes zero scenarios and goes green — a gate that no longer gates. The guard is mechanical: assert a minimum matched count for each stage and fail the run when the count falls below it, or when it drops by more than a small margin from the previous run. **Tag rot on new scenarios.** A scenario written without a tier tag is in no fast subset, so it runs only nightly, and nobody notices for months. The guard is a build check that every scenario carries exactly one tier tag, run alongside the suite itself. **The fast subset growing quietly.** Every author believes their scenario belongs in the smoke tier. Give the stage an explicit time budget, report per-tag duration after each run, and treat adding to the fast subset as a change that must displace something. **Tags used as a substitute for fixing something.** A label that means "skip this everywhere" is not selection, it is an exclusion policy, and it needs an owner and a review date rather than a tag that silently accumulates members. ## Selection and parallelism are complements, not alternatives Tag selection reduces *how many* scenarios a trigger runs. Running the selected set across several workers reduces the *wall clock* for whatever set you chose. Teams reach for the second before the first and are disappointed, because a suite whose scenarios share mutable state cannot be spread safely: two scenarios that both mutate the same ledger account produce interference that looks like a product defect. The precondition for parallel running is that each scenario creates and disposes of its own data, which is a design property of the suite rather than a runner setting. Once that holds, selection plus parallel execution together are what keep a large plain-language pack inside a feedback budget the team will actually respect.
- Why can a scenario pack sometimes run no faster, or slower, spread across several workers?Three reasons. Scenarios that share mutable data contend or interfere, so time goes into retries and investigation. Each worker pays its own start-up and fixture cost, which can dominate a short selection. And an uneven duration distribution leaves one long scenario as a tail that no amount of extra workers shortens. Fix the data isolation and split the longest scenarios before adding workers.
- How do you stop the fast tier growing until it is no longer fast?Give the stage an explicit time budget and enforce it in the build, then report per-tag durations after every run so the growth is visible. Make adding a scenario to the fast tier a displacement decision rather than an addition: something else comes out. Without a budget, every author reasonably believes their own scenario is critical, and the tier drifts back toward the full pack.
- What tag axes would you refuse to introduce?Anything tied to a moment rather than a property: sprint numbers, work-item identifiers, the author, the team that happened to write it. They are correct on the day and meaningless a quarter later, and no stage can be built on them because nobody maintains them. If a scenario's origin matters, that belongs in version history, not in a label the runner selects on.
Tags are the labels on a ledger's index cards: useful only if everyone files under the same small set of headings, and worthless once people start inventing their own.
saying these in an interview costs you the question
- Invents a new tag per feature or per sprint
- Trusts a green stage that executed zero scenarios
- Thinks parallel workers remove the need for selection
- Leaves new scenarios untagged and assumes they run
- Uses a skip label with no owner or review date