A suite reports each case outcome to a case repository in its own call as that case finishes, instead of one bulk push at the end of the run. What does the per-case shape buy, and what does it cost?
answer
- N round trips versus one
- what survives a killed run
- bulk endpoints are paced, not free
- chunked flush over a durable file
- overall success can hide row errors
basics
~20 sPer-case calls fill the cycle live and survive a run that dies halfway. They cost a round trip per case, run into bulk rate limits, and make the run depend on the repository staying healthy throughout.
solid answer
~50 sStreaming one call per case buys two real things: a cycle you can watch fill up while a long suite is still running, and partial results when the run is killed mid-flight. It costs a network round trip per case, which on a large suite dominates the suite's own runtime; it multiplies your exposure to the rate limiting that bulk endpoints usually carry; and it couples execution to the repository's availability, so a slow repository slows the tests. A bulk push at the end inverts every one of those. The shape most teams settle on is a middle one: buffer outcomes in memory, write them to a durable file as they are produced, then flush in chunks — with a final flush at the end — so a crash loses at most one chunk and the file remains a replayable source of truth.
go deeper
Know that outcomes can be sent one at a time as cases finish or in one batch at the end, and that the batch is a single call carrying a list of item-and-outcome pairs.
Explain both costs concretely: N round trips and rate-limit exposure on one side, total loss of a killed run on the other. Describe chunked flushing as the practical compromise.
Demonstrate that you persist outcomes to a durable artefact before pushing, flush from a teardown step that still runs on failure, and reconcile accepted counts against produced counts.
Own the availability argument: decide deliberately whether a test run may depend on a third-party repository being healthy for its whole duration, and design the integration so the answer is no.
## The two shapes There are only two ways to get a run's outcomes into a case repository, and every real implementation is a point between them. - **Stream**: as each case finishes, the suite makes a call recording that one outcome against that one item in the cycle. - **Bulk**: the suite accumulates outcomes and makes one call at the end carrying a cycle identifier and a list of item identifiers with their outcomes. Almost every repository offers the bulk shape precisely because the streaming shape is what naive integrations do first, and it is what melts under a real suite. ## What streaming genuinely buys 1. **Live visibility.** A cycle that fills in as the run proceeds lets someone watching a forty-minute suite see the first failures long before the run ends. For a nightly regression pack this is worth real money. 2. **Survivability.** If the runner is evicted, the container is OOM-killed, or the job hits its time budget, a streaming integration has already recorded everything up to that moment. A bulk integration that never reaches its final call records nothing at all — the most painful failure in this whole area, because the expensive part (running the tests) succeeded and only the cheap part was lost. 3. **Simplicity of failure.** One call maps to one outcome, so a failed call has an obvious blast radius. ## What streaming costs - **A round trip per case.** On a suite of any size the accumulated latency becomes a visible fraction of the run. Ten thousand cases with a modest per-call latency is not a rounding error; it is minutes of pipeline time spent talking rather than testing. - **Rate limiting.** Bulk import and result endpoints are rate limited by essentially every hosted product. A per-case integration is the shape most likely to trip that, and the shape least able to recover, because backing off in the middle of execution either stalls the suite or drops results. - **Availability coupling.** The suite now needs the repository to be healthy for the whole duration of the run, not for one moment at the end. A repository maintenance window turns into a test failure, or worse, into a suite that hangs. - **Retry amplification.** Each of N calls can fail independently, so you need per-call retry, and each retry raises the double-record risk that an idempotency key exists to solve. ## Reading a bulk response properly The usual defect in bulk integrations is not the shape but the response handling. A bulk endpoint commonly returns an overall success while reporting per-row problems inside the body — a handle that resolved to nothing, an item not present in this cycle, a malformed outcome. A push step that inspects only the transport status will report a green push over a cycle that took half the rows. Always parse the body, count what was accepted, and compare it with what the harness produced. ## Comparing the two | Property | One call per case | One bulk push at the end | |---|---|---| | Visibility during the run | Live | Nothing until the end | | Run dies mid-flight | Partial results kept | Everything lost | | Network cost | One round trip per case | One, or one per chunk | | Rate-limit exposure | High | Low, and easy to pace | | Repository outage | Breaks the run | Breaks only the push step | | Retry semantics | Per call, N chances to duplicate | One unit, one key to dedupe | ## The shape most teams land on 1. **Write outcomes to a durable local file as they are produced.** This costs nothing, cannot fail for network reasons, and turns the push into a replayable step rather than the only copy of the truth. 2. **Buffer and flush in chunks.** A chunk sized to the endpoint's tolerance gives most of streaming's survivability at a fraction of its call count, and chunk boundaries are natural retry units. 3. **Flush unconditionally at the end**, including on an aborted or failing run, so a red suite reports as fully as a green one. 4. **Pace and back off on transient rejections**, rather than hammering a rate-limited endpoint, and let the durable file cover anything that still does not land. That design answers the interview question honestly: streaming is not wrong, it is a tradeoff with a specific and well-known cost, and chunked flushing over a durable file buys nearly all of its upside without paying that cost. The decision about which test failures should stop a pipeline stage is a separate matter entirely; this is only about how the outcomes travel.
- Your bulk push at the end never runs because the job is killed at its time limit. What changes?Write outcomes to a durable file as they are produced and flush in chunks during the run, so an eviction loses at most the last chunk. Make the final flush run in a teardown step that executes even when the suite step fails or is cancelled, and treat a cycle that received nothing from a job that clearly ran as an alertable condition.
- How do you size a chunk when the endpoint's limits are not documented in a way you can rely on?Start conservative, measure, and let the pacing be adaptive rather than a constant. Treat rejection-for-rate as a signal to back off and shrink, and success as licence to grow slowly. Because the durable file already holds every outcome, a chunk that fails can simply be replayed, which means you can tune aggressively without risking data.
Streaming is posting every receipt as it prints; bulk is mailing the shoebox at the end of the month. One keeps someone informed all along and survives losing the box; the other is far less work but loses everything if the box never gets posted.
saying these in an interview costs you the question
- Says one call per case is always simpler and therefore better
- Assumes bulk endpoints accept unlimited rows at any rate
- Keeps outcomes only in memory until the final push
- Checks the transport status and ignores the response body
- Puts the final flush in a step skipped on failure