skip to content

A worker makes one encrypt call to a central service per record at hundreds of records a second — what does batching those calls buy?

level: middleimportance: should knowfreq 41%

answer

  1. one operation, one network wait
  2. rate times latency equals calls in flight
  3. the service still performs every encryption
  4. one slow call now holds a hundred records
  5. a retry repeats the whole batch

basics

~20 s

Batching removes round trips, not work. A hundred fields in one call is still a hundred encryptions inside the service; what falls is the number of network waits, request authorisations and concurrent calls the worker must keep in flight.

solid answer

~40 s

Each record currently adds one network round trip to the hot path, and the number of calls that must be in flight is the record rate multiplied by the call latency — 400 records a second at 6 ms means about 2.4 calls always outstanding, and far more if the service slows down. Batching a hundred fields into one call cuts round trips and per-request overhead by a hundred, so request-rate limits and connection concurrency stop being the constraint. It changes nothing about the service's cryptographic work: still 400 operations a second. And it costs you back on two axes — one slow call now delays a hundred records, and one retry repeats a hundred operations.

code

pseudocode · 10 lines
pseudocode
// 400 records/s, one encrypt call per record, 6 ms round trip
for each record in stream:
    record.fieldCiphertext = service.encrypt(keyName, record.applicantId)  // one round trip
// 400 calls/s x 0.006 s = ~2.4 calls in flight, continuously

// same load, 100 fields per call
for each chunk of 100 records in stream:
    ciphertexts = service.encryptMany(keyName, fieldsOf(chunk))            // one round trip
    attach(chunk, ciphertexts)
// 4 calls/s x 0.006 s = ~0.024 in flight; the service still performs 400 encryptions/s

go deeper

for a junior

Remember that each operation is now a network call, so encrypting a field is no longer free. A loop that protects one field per record makes one call per record.

for a middle

Explain the arithmetic: calls in flight equal rate times latency, and batching cuts the number of round trips while leaving the service's per-operation work exactly where it was.

for a senior

Talk about what a batch costs back — tail latency spread across every record in it, retries that repeat a hundred operations, and partial failures you now have to unpick entry by entry.

for a principal

Treat throughput as a capacity question for the estate: per-caller request limits, what the shared service is sized for, and whether this workload's shape belongs on a per-operation service at all.

## One operation, one network wait Encrypting a field used to be a function call: microseconds, with no failure mode worth naming. Sending the same operation to a service turns it into a request — serialise, authenticate, cross the network, wait, parse. For a worker consuming a few hundred claim messages a second and protecting one applicant field in each, that is one round trip per record on the hot path, and the round trip dominates the operation. The cryptography is not the cost; the waiting is. ## The arithmetic that actually constrains you The relation to hold in your head: **the number of calls in flight at any moment equals the request rate multiplied by the time each call takes.** 1. **400 records a second, 6 ms per call.** 400 x 0.006 = about **2.4 calls in flight**, continuously. A single-threaded synchronous worker can only sustain 1 / 0.006 = about **167 records a second**, so the queue backs up long before the service is stressed. 2. **Raise the rate to 1,000 a second** at the same latency and you need 6 calls in flight — still small, but now a concurrency setting somebody has to have chosen on purpose. 3. **Let the service's latency degrade to 60 ms** and the same 400 records a second needs **24 calls in flight**. Nothing in your code changed; somebody else's latency became your concurrency requirement. Line three is the one that bites in production. ## What batching changes, and what it does not | Cost, per second | One call per record | One call per hundred | |---|---|---| | Network round trips | 400 | 4 | | Request authorisations | 400 | 4 | | Cryptographic operations at the service | 400 | 400 | | Calls in flight at 6 ms | about 2.4 | about 0.024 | | Records delayed by one slow call | 1 | 100 | | Operations repeated by one retry | 1 | 100 | Two rows fall by a hundred, one row does not move at all, and two rows rise by a hundred. That is the trade in one table: **batching amortises the transport and the per-request overhead; it does not amortise the work.** If the service meters or bills per operation, batching saves nothing there. If it limits per request — and many do, because request rate is what protects a shared service — batching is exactly the right lever. ## What a batch costs back - **Tail latency spreads.** The slowest call now delays every record travelling inside it, so records that would have been fast inherit the worst call's latency. - **Retries multiply.** A failed call repeats every operation it carried, turning one transient failure into a hundred repeated operations. - **Partial failure becomes yours.** A batch response may succeed for some entries and fail for others, and the worker needs a per-entry result path it never needed before. - **Plaintext for a hundred records is assembled at once** rather than one at a time — a bigger thing to hold, and a bigger thing to get wrong on an error path. - **Flushing becomes real.** A batch that never fills has to be flushed on a timer, and that timer is now part of the worker's latency. ## Sizing it 1. Start from the service's stated request limit and its per-operation cost, not from a round number. 2. Pick a size where the batch filling up, not the flush timer, is the normal case at your usual rate. 3. Bound the batch by what you are willing to repeat on a retry, not by what fits in one request. 4. Measure the high percentile **per record**, not per call — the per-call number will look excellent and mean nothing to the record that waited behind ninety-nine others. ## The shape batching is not Batching reduces how often you ask. It does not change the fact that you must ask, and it does not shorten the path for a single urgent record. If the workload cannot tolerate asking at all — a value needed on every request of a high-traffic page, or bulk objects measured in gigabytes — the answer is not a bigger batch but a different arrangement of where the key material sits, which trades away the property that every operation passes through the service. Reach for that deliberately, not because a batch size kept growing.

  • Why does batching not reduce what the service has to do?
    Because the unit of work is the cryptographic operation, not the request. A hundred fields in one call is still a hundred encryptions inside the service. What batching removes is a hundred network waits, a hundred request authorisations, and the connection concurrency those needed.
  • What breaks first if the batch is made very large?
    Retry cost and tail latency. A failure anywhere in the call repeats every operation it carried, and the slowest call delays every record inside it, so the high-percentile latency of records that would have been fast gets much worse. Partial-failure handling also stops being trivial.

saying these in an interview costs you the question

  • Thinks batching reduces the service's cryptographic work
  • Assumes one call per record is fine because a call is milliseconds
  • Sizes the batch without considering what a retry repeats
  • Ignores that one slow call now stalls every record with it
  • Confuses round trips removed with latency removed