Why must a latency gap measured against a hosted embedding endpoint be repeated before it means anything?
answer
- one call is not a measurement
- what else lives inside wall-clock time
- averaging only kills independent scatter
- blocked trials fold load into the difference
- resolution costs calls and money
basics
~20 sA single call's latency is dominated by network jitter, queueing and co-tenant batching, which swamp any model-shape effect. Only a difference that survives averaging many interleaved calls carries information, and confounds that move with load never average away at all.
solid answer
~50 sPer-call latency against a hosted endpoint is mostly not about the model. Network path variation, server-side queueing, batching with other tenants' traffic, cache warmth and autoscaling all contribute, and each is large relative to the difference being sought. So one observation is noise. Averaging helps, but only against variation that is independent between calls: uncertainty in a mean shrinks roughly as one over the square root of the sample count, so resolving a gap half as large costs about four times the calls. The harder limit is that some confounds are not independent noise — load that drifts with time of day, an autoscaler adding a warmer instance — and those bias the estimate rather than scattering it, so more calls sharpen a wrong number. Measurements must therefore be interleaved and paired, not run in blocks, and some differences never resolve at any budget.
go deeper
Know that a single call's timing tells you nothing, because network and queueing variation are larger than any model effect you are looking for.
Explain the mechanics: which terms besides model compute sit inside wall-clock time, why averaging shrinks uncertainty only as the square root of the sample count, and why that makes fine distinctions expensive.
Demonstrate measurement hygiene — interleaved conditions, repeated windows, a control that must show no difference — and be able to say plainly which claimed effects can never be resolved from outside a shared service.
Be ready to say what this exposure is worth against the cost of doing anything about it, given that the resolution an attacker needs to learn something interesting is also the volume that ordinary abuse controls already see.
## Why the naive measurement says nothing An attacker outside a hosted text-embedding service wants to know whether inputs of one kind are processed differently from inputs of another kind, because that difference would say something about the model's shape or its preprocessing. What they can observe is wall-clock time from sending a request to receiving the vector. That number contains the model's compute, and it also contains: DNS and connection setup, the network path and its congestion that second, TLS work, load-balancer and admission queueing, time spent waiting for a batch to fill alongside other customers' requests, cache warmth, and whatever the platform's autoscaler did in the last minute. Most of those terms are larger and more variable than the effect being sought. A single-call comparison is therefore not a weak measurement; it is no measurement. ## What averaging buys, and its price For variation that is independent from call to call, repeated sampling works in the ordinary way: the uncertainty in an estimated mean falls roughly in proportion to one over the square root of the number of samples. The consequence is unpleasant for the attacker. Resolving a gap half as small requires roughly four times as many calls; a gap one tenth as large costs about a hundred times as many. On a metered endpoint, that cost is denominated in money, and in traffic volume that a rate limiter or an anomaly detector may notice. The "free" channel stops being free precisely at the resolution where it becomes interesting. ## The floor that averaging cannot cross The more important limit is structural. Averaging removes zero-mean, independent scatter. It does nothing about a confound that is **correlated with the thing being compared**. If all the calls of type A were sent in one window and all the calls of type B in another, then any drift in load, any autoscaling event, any cache that warmed up in between, is folded directly into the difference — and adding calls makes the biased estimate tighter, not truer. This is the single most common defect in reports of this kind: a beautifully narrow confidence interval around an artefact of scheduling. The discipline that makes such a measurement credible is not exotic: alternate the two conditions so both experience the same load conditions, repeat across separate time windows, and include a control condition that is known to be identical to itself and therefore must show no difference. If the control shows a gap, the apparatus is measuring the platform. And some effects genuinely never resolve — where the model-side difference is small relative to per-call variability that is itself driven by other tenants, no budget of calls separates them, because the noise is not a property the attacker can shrink. ## What survives this limit What survives is coarse and robust: relationships rather than small offsets. Latency that grows roughly with input length, a visible step where behaviour changes at some length, or an order-of-magnitude difference between two endpoints — those clear the noise floor easily and are cheap to establish. Fine distinctions — a few percent difference attributed to a particular architectural detail — do not, and are the ones most often over-claimed. ## Why the limit is the interesting half This is the leaf's actual content: the channel exists, it is free at low resolution, and its resolution is bounded by infrastructure the attacker does not control. An answer that only says "timing leaks information" has said the easy half. The answer that lands is the one that states what it costs to sharpen the measurement, and what stays unresolvable no matter what is spent. ## The boundary worth stating out loud This is not the cryptographic timing question — there is no secret whose comparison must be made constant-time here. And it is not a serving-performance question — nobody is tuning batching or defending a latency objective. The actor is outside the service, paying per call, trying to decide whether a number they measured is about the model or about the platform.
- Roughly how much extra traffic does an attacker need to resolve a gap half as large?About four times as many calls, because the uncertainty in a mean falls roughly with the square root of the sample count. That quadratic cost is what makes fine timing distinctions expensive on a metered endpoint — the volume becomes visible to rate limiting and anomaly detection well before the resolution becomes interesting.
- Why does interleaving the two conditions matter more than collecting more samples?Because interleaving is what converts a confound into shared noise. If both conditions experience the same load, autoscaling and cache warmth, drift affects them equally and cancels in the paired difference. If they are run in blocks, the drift is inside the estimate, and every additional sample tightens a confidence interval around an artefact.
- What would a negative control look like in a timing measurement of this kind?A pair of conditions that are known to be identical to the service — the same input sent under both labels — which must therefore show no difference. If the control separates, the measurement is reading the platform rather than the model, and no result from that run should be believed.
saying these in an interview costs you the question
- Reports a difference from a handful of calls
- Runs all of one condition, then all of the other
- Assumes more samples fix a confounded comparison
- Ignores that co-tenant batching is outside their control
- Treats the free channel as free at any resolution