skip to content

A Cucumber-JVM pack outgrew the pipeline budget — how do you choose between tag tiering, parallelism and sharding?

level: principalimportance: should knowfreq 44%

answer

  1. Ask what each stage must gate
  2. Measure per scenario, not in total
  3. Every lever trades something different
  4. Coverage, isolation, infrastructure or time
  5. Watch escaped regressions, not just minutes

basics

~20 s

Measure per-scenario durations first, then decide what each pipeline stage must gate. Tag tiering buys speed by covering less, in-JVM parallelism buys it by demanding isolated glue, and sharding across agents buys it with infrastructure and result aggregation.

solid answer

~50 s

Start from two numbers, not from a lever: where the time actually goes per scenario, and what each stage is meant to gate. Then the levers trade differently. A **tagged fast tier** — say a 7-minute smoke tag on every push with the full pack on merge — is the cheapest and costs coverage: a regression outside the tag reaches the main branch. **In-JVM parallelism** costs no coverage but demands parallel-safe glue and can simply move the bottleneck onto the system under test. **Sharding across agents** scales furthest and adds a partition rule, an aggregation step to turn several partial results into one verdict, and a real bill. **Moving assertions down the pyramid** is slowest to deliver and the only lever that makes the pack permanently smaller. Most real answers combine the first three and fund the fourth continuously.

code

properties · 2 lines
properties
# supplied per pipeline stage rather than compiled into the suite class
cucumber.filter.tags=@smoke and not @flaky

go deeper

for a junior

Know that a large scenario pack is usually split into a fast tagged tier and a fuller one that runs later, and that the split trades feedback speed against coverage.

for a middle

Explain each lever concretely: what a tag filter changes, what parallelism requires of the glue, and what sharding adds in partitioning and aggregation work.

for a senior

Show a rollout you would defend — measure first, fix the slow tail, keep the gating tier off the most unstable layer, and prove it worked with flake and escaped-defect numbers rather than minutes.

for a principal

Own the risk decision behind the split: what each stage gates, who accepts what reaches the main branch, the infrastructure budget sharding creates, and the review convention that keeps the fast tier honest over years.

## Decide what you are protecting before you pick a lever A pack of 214 scenarios that takes 31 minutes against a 12-minute pull-request budget is not one problem. It is three questions that people collapse into one: 1. what must be true **before a change merges**; 2. what may be discovered **an hour later**, on the main branch; 3. what may wait **until tomorrow**, in a nightly run. Every lever below is really an answer to that question, and choosing a lever without answering it first is how teams end up with a fast pipeline that catches nothing. So begin with measurement. Per-scenario duration, not total wall clock — the total tells you the pack is slow, the distribution tells you whether eleven scenarios own half the time or whether all 214 are uniformly mediocre. Those two shapes have different answers, and the second one is much worse news. ## The four levers and what each one costs | Lever | Buys | Costs | Wrong when | |---|---|---|---| | Tag tiering | immediate feedback on a subset | coverage between stages | the untagged remainder is where regressions actually live | | In-JVM parallelism | the same coverage in less time | parallel-safe glue, load on the system under test | the suite has shared state you have not fixed | | Sharding across agents | scale beyond one machine | partitioning, aggregation, infrastructure spend | the pack is slow because of a few scenarios, not its size | | Moving assertions down | a permanently smaller pack | time, and a team willing to delete scenarios | the behaviour genuinely only exists end-to-end | **Tag tiering** is where most teams start, because it is a text edit. A 7-minute tagged tier on every push and the full pack on merge is a defensible split — as long as you say out loud what you have accepted: a regression outside the tag reaches the main branch and is found by a later stage. That is a risk decision, not a technical one, and it belongs to whoever owns the release. The failure mode is tag rot: nobody adds new scenarios to the fast tier, it thins out over two quarters, and the gate becomes ceremonial. **Parallelism** keeps all the coverage, which is why it is the most attractive lever, and it demands the most from the glue: no shared static state, no scenario depending on another's leftovers, per-scenario test data. It also relocates the bottleneck. Six threads against one test database or one browser grid slot produces timeouts that read like flakiness, and the temptation is then to add retries, which buys back the runtime you just saved. **Sharding** is the lever that scales furthest and the one people underestimate. Splitting a pack across six agents means owning a partition rule — by feature file, by tag, or, best, by recorded duration so the shards finish together — plus an aggregation step that turns six partial results into one verdict and one report. It also changes triage: a failure now lives in one shard's log, and reproducing it locally means knowing which shard it ran in. **Moving assertions down the pyramid** is the only lever that makes the number smaller instead of hiding it. A scenario that asserts a pricing rule the domain layer can assert directly does not belong in an end-to-end pack. This is slow, unglamorous work, and it is the difference between a suite that stays viable for three years and one that gets rewritten. ## A worked decision On the vinyl-record marketplace pack: measurement shows 214 scenarios, 31 minutes, and a long tail — 38 browser-driven scenarios account for 19 of those minutes. A browser upgrade lands mid-sprint and eight of those 38 turn intermittent for a week. That shape argues for a specific combination. Keep the gating tier off the browser entirely: a 7-minute tagged tier of service-level scenarios that a browser upgrade cannot destabilise, so the gate stays trustworthy exactly when the environment is moving. Run the full pack on merge with modest fixed parallelism, sized against the test database rather than the agent. Shard only if merge-stage time still misses its budget after the tail is addressed — because sharding a pack whose slowness is concentrated in 38 scenarios buys much less than fixing those 38. ## How you know it worked Speed is the easy metric and the misleading one. Track alongside it: - **escaped regressions per stage** — how often something the fast tier could have caught was caught later instead; - **retry and flake rate**, because a suite that halved its runtime and doubled its retries is slower in practice; - **shard skew**, the gap between the fastest and slowest shard, which is the efficiency of your partition rule; - **agent minutes per merge**, the bill the sharding decision creates; - **the fast tier's share of the pack**, watched over quarters, so tag rot is visible before it is fatal. And make it a standing rule rather than a project: every new scenario declares which tier it belongs to at review time, and the reviewer is entitled to ask whether it needs to be a scenario at all. That convention is what keeps the budget from being renegotiated every six months.

  • How do you keep a tagged fast tier from rotting over a year?
    Make tier membership a review decision rather than an afterthought: every new scenario states which tier it belongs to, and the reviewer can push back. Then track the tier's share of the pack and its escaped-regression count as standing metrics. A tier nobody adds to is a gate that quietly stopped gating, and only a measured share makes that visible in time.
  • What has to exist before you shard a Cucumber-JVM pack across agents?
    A partition rule that balances by measured duration rather than by scenario count, an aggregation step that combines partial results into one verdict and one report, and a way to tell which shard a failure came from. Without the last two, sharding trades a slow pipeline for an unreadable one and triage cost quietly replaces the time you saved.
  • The pipeline is fast now and defects still reach production. What do you re-examine?
    The gating decision, not the speed. Look at what the fast tier omitted for each escaped defect: if the omissions cluster, the tier is drawn in the wrong place; if they are scattered, the pack is testing the wrong layer and the answer is moving assertions down rather than moving scenarios between stages.

saying these in an interview costs you the question

  • Reaches for parallelism before measuring per-scenario time
  • Adds a smoke tag without naming the coverage accepted
  • Shards a pack whose slowness is a handful of scenarios
  • Raises thread counts and retries to hide flakiness
  • Treats pipeline minutes as the only success metric