You set the stage contract for an upload pipeline used by many teams: when is deep-copying each chunk buffer the honest choice over moving it?
answer
- footprint against concurrency, not depth
- move by default, copy as exception
- retry, lifetime, or a tiny payload
- one failure rule for every stage
- bound anything a stage retains
basics
~20 sMove by default, so footprint tracks in-flight chunks rather than pipeline depth. Allow a copy only where a stage must retry from the original bytes, must keep them for a different lifetime, or where the payload is small enough that the copy is noise.
solid answer
~50 sStart from the arithmetic. With 64 chunks of 8 MiB in flight, a move-only contract keeps the payload footprint near 512 MiB whatever the pipeline depth; if every one of four stages holds its own version, the same workload can reach 2 GiB. So moving is the default and a copy is an exception that has to be argued. The exceptions are real, though: a stage that must retry from the original bytes when a downstream call fails, a stage that keeps the payload on a different lifetime such as an audit or a deferred checksum, and payloads small enough that duplication is below the noise floor. The second half of the decision is organisational: a move-only contract forces every consuming team to restructure retries, and that cost is paid once, in review, rather than continuously, in memory.
go deeper
Take the default away from this: hand the payload on rather than duplicating it, and duplicate only when something genuinely needs a second, separately owned version.
Be able to price it — copy traffic grows with payload size and pipeline depth, while a handoff by move stays flat — and name the cases where a copy is genuinely required.
Bring the operational consequence: with moves, peak footprint is concurrency times chunk size; with copies, it also multiplies by how many stages hold a version at once.
Own the whole contract — default direction, failure rule, retention bound, and a declared exception route — and state the adoption cost you are imposing on the consuming teams.
## Frame it as a footprint decision, not a style one A stage contract is a decision about what the pipeline's memory does under load, so start with the number. Take **64 chunks of 8 MiB in flight** across **four stages**: - **Move-only:** one live block per in-flight chunk, so 64 x 8 MiB = **512 MiB** of payload, independent of how many stages the chunk passes through. - **A copy held at every stage:** up to four live versions per chunk, so 4 x 512 MiB = **2 GiB** for the same workload, plus 3 x 8 MiB = 24 MiB of copy traffic per chunk. The difference is not a micro-optimisation; it is whether the service fits in its box. It also changes the shape of the capacity model: with moves, footprint is a function of **concurrency** alone, which operators can control with one dial. With copies it is a function of concurrency multiplied by **depth**, which changes whenever someone inserts a stage. ## Where a copy genuinely earns it A copy is a purchase, and there are things worth buying: - **Retry from the original bytes.** If a downstream call can fail after consuming the buffer and the source cannot be re-read — a stream, a one-shot reader — someone must hold the original. A copy is one answer; returning ownership on the error path is usually a cheaper one. - **A different lifetime.** An audit record, a deferred checksum, a sampled payload kept for diagnosis: these outlive the pipeline, so they cannot be satisfied by a handle that leaves with the chunk. - **A small payload.** Duplicating a few hundred bytes of header or metadata is below the noise floor, and forcing a move-only discipline there buys nothing but awkward code. - **A representation change anyway.** If the stage must re-encode, compress or reframe the bytes, it is walking them regardless; the copy is the work, not an overhead on it. ## Where the copy is really covering for something else Just as often, a copy is a symptom: 1. **Nobody knew who was allowed to release the buffer**, so everyone kept a version. The fix is an ownership contract, not more memory. 2. **The retry design was never written down**, so each stage defends itself. The fix is a single failure contract across all stages. 3. **A stage wants to report on a chunk after handing it on.** The fix is to capture the small values it needs before the transfer, not to duplicate megabytes. ## What a contract for many teams has to settle Because dozens of teams write stages against this interface, the contract must answer more than move-or-copy: - **The default direction:** a stage takes the owning handle and is expected to consume it. - **The failure rule:** what a stage that took ownership does when it fails — release, or hand ownership back with the error — decided once, so retry code is uniform everywhere. - **The retention rule:** whether a stage may keep a buffer past its call at all, and with what bound, because unbounded retention is how a downstream outage turns into a memory incident. - **The exception route:** how a team that genuinely needs a copy declares it, so the footprint model can account for it instead of being surprised by it. ## The trade you are actually making A move-only default buys a predictable footprint and a single release site per chunk; it costs every consuming team a restructuring of the places where they assumed they could still look at the payload afterwards. A permissive copy policy buys easy adoption; it costs a footprint that grows with pipeline depth and a retry path whose peak memory nobody can state. For a shared platform the first is usually right, because the restructuring is paid once in review while the footprint is paid continuously in production — but the honest version of the answer names the exception route as part of the same decision, rather than pretending no stage ever needs a second version of the bytes. The weak answer to this question is a rule ('always move'). The strong one is a default, the small set of cases that override it, the number that justifies the default, and the migration cost the default imposes on the teams that have to adopt it.
- A team says it needs a copy so it can retry after a downstream failure. What would you offer instead?Have the failing stage hand ownership back with its error. The caller then retries against the original bytes with no duplication, and the retry path's memory is visible in the signature rather than hidden in a defensive copy. A copy is only needed when two parties must hold the payload at once.
- What breaks first if the copy policy is permissive and the pipeline later grows a stage?Peak footprint, silently. Each additional stage that holds its own version multiplies the payload resident at once, so a change that looks local raises the whole service's memory ceiling. Under a move-only contract, adding a stage leaves footprint untouched because the block never duplicates.
- How do you keep the exception route from becoming the norm?Make the exception explicit and accounted: a declared copying stage, a stated payload size and retention bound, and inclusion in the pipeline's footprint model. What is measured stays rare; what is merely discouraged becomes the default the moment a deadline appears.
saying these in an interview costs you the question
- States a blanket rule with no default and no exceptions
- Ignores that copy footprint multiplies by pipeline depth
- Lets stages retain buffers after failure with no bound
- Treats a defensive copy as free because it is simple
- Leaves each team to decide the failure contract separately