How do you design an automated case over a slow, lossy link so its result is repeatable?
answer
- The link is an input, not weather
- Name the profile before writing the case
- Fixed delay, loss rate, bandwidth cap
- Assert behaviour bands, not exact durations
- Record the profile in the failure artefact
basics
~20 sDeclare the link as an input: a named profile with fixed added delay, loss rate and bandwidth cap, applied at one controlled point in the request path. Then assert product behaviour under that profile, not raw elapsed time.
solid answer
~50 sTreat the degraded link the way you treat the seeded account or the test data row - a **declared input**, not ambient weather. Give the case a named profile: fixed added round-trip delay, a fixed loss rate, a bandwidth cap, and whether variation in delay is applied. Apply it at a single controlled point the run owns - a hop in the request path, or the device's own network configuration set by the run - so two engineers get the same behaviour. Restore the profile afterwards so it does not leak into the next case. Assert **behaviour bands** rather than exact durations: a progress surface appears while the call is outstanding, repeated taps produce one effect, the app's own deadline is honoured. Put the profile name in the failure artefact, or a red result tells you nothing about the conditions that produced it.
code
pseudocode · 21 linesprofile LOSSY_MOBILE:
added_delay_ms = 300
delay_variation_ms = 80
loss_percent = 2
downlink_kbps = 400
uplink_kbps = 150
case "purchase completes on a weak link":
setup:
impairment.apply(LOSSY_MOBILE, at = request_path_hop)
artefacts.record("profile", LOSSY_MOBILE.name)
steps:
open(product_screen)
tap(buy)
expect visible(progress_surface)
expect disabled(buy)
await settled(purchase_call, within = app_declared_deadline)
expect visible(confirmation) or visible(failure_notice)
expect server.orders(for = user).count == 1
teardown:
impairment.clear(at = request_path_hop)go deeper
Be ready to say why a slow-link case needs its conditions written down. Recall the three basic knobs - added delay, loss rate, bandwidth cap - and that they are set by the run, not by luck.
Explain the mechanics: which defect each knob exposes, where the impairment is applied so the case is portable, and why the profile is set and reverted inside the case rather than left on the machine.
Show the production judgment: assert behaviour bands rather than durations, keep the impaired set small and in its own lane, and make every red result carry the profile that produced it so failures are attributable.
Own the tradeoff between fidelity and control - a hop the run owns is repeatable but models no radio, while the client's own configuration is faithful and hard to guarantee. Decide which flows are worth the runtime at all.
A case that runs "on whatever connection the machine happened to have" is not a network-variability case. It is an ordinary case with unexplained variance, and its failures are unattributable: nobody can tell whether the product regressed or the building did. ## The link is an input, not the weather The design move that makes everything else possible is to stop treating the connection as ambient and start treating it as a **declared input**, exactly like the account the case signs in as or the data row it seeds. Concretely, the case names a *profile* and the run applies it deliberately: - a fixed **added round-trip delay** (the same number on every machine); - a fixed **loss rate**, expressed as a share of packets dropped; - a **bandwidth cap** in each direction; - whether **variation in delay** is applied on top of the fixed delay, and within what range. Two engineers running the same case with the same profile should observe the same product behaviour. When they do not, the product changed - which is the only thing a red result is allowed to mean. ## What each knob actually exposes The knobs are not interchangeable. Choosing them by what defect you are hunting is what separates a designed case from a vague "slow network" case. | Knob | What it models | What it exposes | | --- | --- | --- | | Added delay | distance, congestion, a busy cell | serial chains of round trips, code that assumes a fast first response | | Loss | a weak or fading radio | stalled transfers, half-finished uploads, retransmission pile-ups | | Bandwidth cap | a crowded shared link | oversized payloads, unbounded image or media loading | | Delay variation | a moving user | ordering assumptions between two concurrent calls | A screen that issues six dependent calls in sequence looks fine at one millisecond of delay and unusable at three hundred; a screen that downloads full-resolution media looks fine until the cap is applied. The profile is how you choose which of those two defects the case is for. ## Where to apply the impairment There are three usual placements, and they trade fidelity against control: 1. **At a hop in the request path that the run owns** - a gateway or stand-in the suite configures before the case and resets after. Highest control, works identically on every machine, and it is the placement that makes a case portable to a shared pipeline. It does not model the radio itself. 2. **At the device or emulated environment's own network configuration**, set programmatically by the run. Closer to the real client stack, but it needs the run to have that level of access to every target, and it is easy to leave the setting behind after a failure. 3. **Manually, by an engineer degrading a shared connection** - never acceptable for an automated case. It is unrepeatable, it degrades everyone else on that connection, and it cannot run unattended. Whichever you pick, two rules hold. The profile must be **set and reverted inside the case's own setup and teardown**, so a crashed case cannot poison the next one. And the profile must be **recorded in the failure artefact** alongside the screenshot and the log, because a red result whose conditions are unknown is a re-run, not a diagnosis. ## What to assert under the profile The temptation is to assert a duration, and it is the single most common way these cases become flaky. Wall-clock time under loss is a distribution, not a value; asserting that a call finished in under a specific number of milliseconds converts every unlucky retransmission into a false red. Assert **behaviour bands** instead: - a progress surface is present while the call is outstanding, and gone once it settles; - the control that started the call cannot be triggered a second time while it is outstanding; - the app's own deadline is honoured - after it passes, a definite outcome is shown rather than an indefinite wait; - exactly one effect exists on the server side for one user action; - entered data survives whatever the app decides to do. Each of those is a property of the product under the declared profile, and each stays true across the whole distribution of arrival times the profile produces. ## Keeping the case affordable Impaired-link cases are slow by construction, and a suite that runs every flow under a profile mostly buys itself a long, unreliable pipeline stage. Keep the set deliberately small: - pick the two or three flows where a degraded link is genuinely part of the product promise - the purchase, the upload, the first load after sign-in; - run them in their own lane rather than inside the main regression pack, so their runtime is visible and separable; - pin one profile per case and change it only when you can say which defect the new profile is for; - when a case fails, reproduce it by name - "the case failed under the lossy profile" is a reproduction instruction; "it failed on the network" is not. The end state is a small number of cases whose conditions are written down, whose failures point at the product, and whose runtime is a known cost rather than a surprise.
- Why is a bandwidth cap a different test from added delay, if both just make things slower?They fail different code. Added delay punishes the number of sequential round trips a flow makes, so it exposes chatty screens that could batch or parallelise. A bandwidth cap punishes payload size, so it exposes full-resolution media, unpaginated lists and responses carrying fields the screen never renders. A flow can be clean under one and unusable under the other.
- The same impaired case passes locally and fails in the shared pipeline. What do you check first?Whether the profile is actually the same. Impairment applied at a hop the run owns travels; impairment applied to a developer machine's own configuration does not. Check that the pipeline target got the profile set, that a previous failed case did not leave a different one behind, and that the pipeline's own baseline latency is not stacking on top of the added delay.
- How do you stop these slow cases from dominating the pipeline's wall-clock time?Keep the impaired set to the few flows where a degraded link is part of the product promise, and run them in a lane of their own rather than inside the main pack. Their runtime then shows up as a named cost the team can decide about, instead of quietly inflating every run and pushing people to shorten the profile until it stops testing anything.
It is the difference between saying a car was tested in bad weather and saying it was tested at four degrees on a wet surface. Only the second one can be repeated by someone else.
saying these in an interview costs you the question
- Runs the case on whatever connection the machine has
- Asserts an exact duration under a lossy profile
- Re-runs until green instead of fixing variance
- Leaves the impairment applied after the case ends
- Degrades a shared office connection by hand
- Reports a red result without the profile that produced it