How does a suite decide which evidence to capture on every case and which only on failure?
answer
- Not all evidence costs the same to gather
- Cheap to produce, expensive to keep
- Buffer in memory, write only on failure
- Renderings and dumps cost seconds each
- A store nobody prunes is also a cost
basics
~20 sSplit by cost. Evidence that is cheap to produce and expensive only to keep — the step log, durations, identifiers — is produced always into a bounded buffer and written only on failure. Evidence expensive to produce is gated behind a failure.
solid answer
~50 sCapture has three costs — wall-clock time per case, storage plus the attention of whoever reads it, and the risk of perturbing the run — and they fall differently on different evidence. Cheap items cost microseconds and a few kilobytes: the ordered step log, the durations, the identifiers. Produce those unconditionally but hold them in a bounded in-memory buffer per case and discard it when the case passes, so a green run writes nothing. Expensive items — a rendering of the interface, a process dump, whole request and response bodies — cost seconds and megabytes each, and across a few thousand cases that is a real slice of the run plus a store nobody prunes. Trigger those from the failure path only. The rule of thumb: if it is free to produce and expensive only to keep, buffer it; if it is expensive to produce, gate it behind a failure.
code
pseudocode · 18 lines# cheap evidence is always produced, but only kept when the case fails
buffer = RingBuffer(capacity = 500) # per case, in memory
on_interaction(step):
buffer.push(step.name, step.ok, step.elapsedMs) # microseconds, no disk
on_case_finished(case, outcome):
if outcome.passed:
buffer.discard() # a green case writes nothing
return
write(buffer.drain(), droppedEntries = buffer.dropped)
if policy.expensiveCaptureAllowed(case):
write(renderCurrentView()) # ~1s, hundreds of KB
write(bounded(redacted(case.bodies), maxBytes = 256_000))
if policy.dumpsAllowed(case) and outcome.kind == "crash":
write(processDump()) # seconds, tens of MBgo deeper
Know that gathering evidence is not free: producing a rendering or a dump takes real time and keeping it takes real space. Be able to say why a suite would not simply record everything on every case.
Explain the split — evidence cheap to produce and expensive only to keep, versus evidence expensive to produce at all — and describe holding the first kind in a bounded buffer that is written out only when the case fails.
An interviewer at this level expects consequences with numbers: the per-case seconds a heavy capture adds across a few thousand cases, the store nobody prunes, and the way capture can shift timing enough to mask or create the failure it is recording.
Own the budget. Decide what evidence the estate gets by default, what a team must justify to exceed it, how long any of it is retained, and who answers for the personal data that ends up inside a long-lived store of failure artefacts.
## Capture has three costs, not one "Storage is cheap" is the argument that produces suites which record everything, and it answers only one of three costs. 1. **Time per case.** Rendering the interface, dumping a process or serialising whole bodies takes real wall-clock time. A capture that adds one second across four thousand cases adds an hour to a suite that people are already complaining is slow. 2. **Storage and attention.** Storage is genuinely cheap; the store *nobody prunes* is not, and neither is the reviewer who opens a bundle of ninety files and reads none of them. More evidence past a point produces less diagnosis. 3. **Perturbation.** Capture is work the product would not otherwise do. It changes scheduling, it can force the application to compute something it had not, and on a timing-sensitive case that is enough to mask a race on one run and create one on the next. Evidence that alters what it measures is worse than none. The first and third costs are the reason "capture everything always" is a design error rather than a budget question. ## Cheap to produce, expensive to keep The useful split is not by importance but by **where the cost falls**. Some evidence is nearly free to produce and only costs anything when you persist it: the step log, per-step durations, the identifiers, the resolved configuration the case ran under. Produce that unconditionally, into a bounded in-memory buffer scoped to the case, and simply discard the buffer when the case passes. A green run then writes nothing at all while still having had the evidence available the whole time. The buffer must be bounded. A case that loops a thousand times would otherwise hold a thousand entries, and the last few dozen are the ones that matter; a ring that keeps the most recent N entries and drops the rest is the right shape, with a note in the bundle when entries were dropped so nobody reads a truncated history as a complete one. ## Expensive to produce Other evidence costs real time the moment you ask for it, so it can only be gated behind an actual failure: - A rendering of the interface as it stood. - A dump of the relevant process or its memory. - Full request and response bodies for the whole case rather than sizes and status. - Anything that requires another round trip to the system to collect. ## A capture budget | Evidence | Cost to produce | Cost to keep | Capture when | |---|---|---|---| | Identifiers, target, clock, worker | Free | Bytes | Always | | Step log with outcomes and durations | Microseconds | Kilobytes | Always, buffered; written on failure | | Interface rendering | Around a second | Hundreds of kilobytes | On failure | | Process or memory dump | Seconds | Tens of megabytes | On failure, and not for every failure | | Full request and response bodies | Milliseconds to seconds | Megabytes | On failure, bounded and redacted | The rule of thumb that falls out: **if it is free to produce and expensive only to keep, buffer it; if it is expensive to produce, gate it behind a failure.** ## When always-on is right after all Two cases justify capturing on green runs, and both should be deliberate rather than default. The first is comparison: when a failure is only intelligible next to a passing run, having the passing run's rendering is genuinely valuable — so allow it as a named, opt-in subset or a low sampling rate with short retention, rather than for every case of every run. The second is a case under active investigation, where you turn capture up for a handful of cases and turn it back down when the question is answered. The failure mode is leaving it on; a setting whose cost nobody feels is never reverted. ## What the store owes you back Two obligations follow from keeping failure evidence at all: - **Retention.** Decide how long a bundle lives before anyone writes one. Evidence for a failure from eight months ago has no reader, and the decision not to make one is a decision to keep it forever. - **Exposure.** Renderings and payload bodies carry whatever real-shaped data the run used — names, tokens, contact details. An unbounded, long-lived store of failure artefacts becomes a place that data lives without anybody having decided it should. Bound the bodies, redact the fields you can name, and put the retention window on the store rather than in a habit. The suite that gets this right is not the one that captures the most. It is the one where a failure produces a small, complete bundle in seconds, a pass produces nothing, and nobody is afraid to look at the store.
- What is the risk of running a heavy capture on every case rather than only on failures?It can change the outcome. Reading interface state, dumping a process or serialising large bodies takes time and can force work the product would not otherwise do, which shifts timing enough to hide a race on one run and create one on the next. Evidence that alters the thing it measures is worse than no evidence.
- A team wants renderings kept for passing cases so they can compare against a failing one. How do you handle that?Make it bounded and opt-in rather than the default: a named subset of cases, or a low sampling rate, with a short retention window. The comparison is genuinely useful, but paying for it on every case of every run buys a store that is almost entirely images nobody will ever open.
- What does over-capturing cost beyond time and storage?Reviewability and exposure. A bundle of ninety files means the reader opens none of them, so more evidence yields less diagnosis. And renderings and payload bodies carry whatever real-shaped data the run used, so an unbounded, long-retained store becomes a place personal data lives without anyone having decided it should.
A dashcam that keeps a rolling few minutes and only saves the clip when the impact sensor fires. Recording continuously is cheap; keeping every hour of it is not.
saying these in an interview costs you the question
- Captures everything on every case because storage is cheap
- Assumes capture is free and cannot affect the run's timing
- Keeps failure evidence forever with no retention decision
- Turns off all capture to make the suite finish faster
- Treats more artefacts per failure as strictly better
- Buffers the step log without bounding how much it holds