A timing finding against your paid embedding endpoint reproduces once in five runs — how do you triage it?
answer
- one in five is a confound's signature
- interleaved or blocked collection?
- a control that must show nothing
- compare the effect against the spread
- rate it against what replies already disclose
basics
~20 sA one-in-five reproduction against a shared endpoint is unproven until the measurement rules out the platform: were conditions interleaved, was a null control run, does the effect exceed per-call spread? Then rate severity against what replies already disclose.
solid answer
~50 sStart with the apparatus, not the model. One-in-five reproduction against a shared hosted service is the signature of a confound: conditions collected in separate blocks pick up load drift, autoscaling and cache warmth, and a cold-start window can manufacture a gap that vanishes later. Were the two conditions interleaved and paired, so both experienced the same load? Was a negative control run — the same input under both labels — which must show no difference? Is the claimed effect larger than the per-call spread, or a few milliseconds sitting inside a wide tail? Then re-run across separate time windows before believing anything. Finally, rate it: even a confirmed family-level timing signal usually adds only a coarse prior on top of what the reply shape and error text already disclose, so it is a low-severity note, not an incident.
code
text · 10 linesclaim: "latency separates short-form from long-form inputs beyond tokenization cost"
cond input_len n_calls median_ms p10-p90_ms collection_window
A 32 200 41.2 28 - 96 Tue 02:10-02:25 UTC
B 32 200 47.9 26 - 214 Tue 14:40-14:55 UTC
...
sampling: all of A, then all of B
control: same input under both labels -- not run
repeats: 1 window per conditiongo deeper
Know that an intermittent timing result against a shared service is usually the platform, and that the first question is how the measurement was collected, not what it implies.
Be able to name the confounds inside wall-clock latency and explain why interleaving and a negative control are what make a paired comparison believable.
Show you can run the triage end to end: suspect the confound, name the check that rules it out, demand dispersion alongside medians, and set a severity ceiling based on what the endpoint already discloses anyway.
Own the standard your team applies to findings of this class, so that measurement-quality expectations are set in advance and coarse shape inferences do not consume product usability in response.
## The situation Someone hands you a finding: measurements against your hosted text-embedding endpoint show that inputs of one kind take measurably longer than another, and they conclude something about the model behind it. It reproduced once in five attempts. Your job is to decide whether this is a real signal about the model's shape or an artefact of the platform, and what it is worth either way. ## Step one: assume the platform until shown otherwise Wall-clock latency against a shared service contains far more than model compute: connection setup, network path variation, admission and queue time, waiting for a batch to fill alongside other customers' traffic, cache warmth, and autoscaling events. Any of those can generate a difference of the size typically reported, and several of them drift on a timescale of minutes to hours. Intermittent reproduction is exactly what a load-correlated artefact looks like: it appears in the windows where the drift happened to line up with the labels and disappears elsewhere. ## Step two: interrogate the apparatus Three questions settle most findings of this class. **Were the conditions interleaved?** If all of condition A was collected and then all of condition B, any drift between the two collection windows is inside the reported difference. Alternating the conditions makes both experience the same load, so drift cancels in the paired comparison. Blocked collection is the most common defect here and the easiest to spot in a report. **Was there a negative control?** A pair of conditions known to be identical — the same input submitted under both labels — must show no difference. If the control separates, the apparatus is measuring your platform, and every other number from that run is uninterpretable. A report without a control has not established that its instrument works. **Is the effect larger than the spread?** A gap in medians that is small compared with the per-call variation is not a finding, however many decimals it is quoted to. Ask for the dispersion, not only the central value. One more: **repeat across separate windows**, ideally at different times of day and, if the service is multi-region, from more than one vantage. Real model-shape effects are stable across those; platform artefacts are not. ## Step three: rate it honestly, even if it is real Suppose it survives all of that. What has been established? Almost always a coarse relationship — latency grows with input length, or a step appears at some length — which supports a family-and-scale inference. Compare that with what the endpoint already discloses to any paying customer on the first call: the exact output dimension and numeric type, and, from error text, close to the exact input ceiling. A timing result that merely corroborates the same family-level picture adds very little. That makes it a low-severity note in most reports. It is worth recording, because it is cheap for an attacker and effectively permanent, but it does not justify degrading the product to obscure a shape the product must expose. The severity conversation should also state what a family-level prior is actually worth to someone assembling a copy: it trims their search, and it is not what determines whether their copy is any good. ## What a good triage write-up says Name the confound you suspect, state the specific check that would rule it out, and state the ceiling on severity if the check passes. "Not reproducible under interleaved sampling; control separated; closing" is a complete outcome. So is "reproduces across three windows with a clean control; effect is a coarse scale relationship, already implied by the disclosed output dimension; logging as low severity". Both are better than escalating a flaky measurement or dismissing it without touching the apparatus. ## The boundary This is not a serving-performance investigation — nobody is tuning batching or defending a latency objective — and it is not a cryptographic timing question, because there is no secret whose comparison must be made constant-time. It is an outsider's inference about a model's shape, and it is being triaged by the person who has to decide what the finding is worth.
- The report shows a 6 ms median gap with a p10-p90 spread of nearly 200 ms. What does that tell you?That the effect is far inside the noise, so the central value alone cannot support the claim. With that much dispersion the medians need many more paired samples to separate at all, and the wide tail in one condition suggests queueing or autoscaling rather than model compute. Ask for interleaved collection and dispersion-aware comparison before reading anything into the gap.
- The submitter says re-running would cost too much in API charges. How do you respond?That the cost is the finding's own limit and it belongs in the report. Resolution scales with the square root of the sample count, so a gap this small inside this much spread needs a lot of calls. If that budget is not available, the honest disposition is unproven rather than confirmed, and the severity ceiling is low enough that spending more is usually not warranted.
- It reproduces cleanly under interleaved sampling with a clean control. What now?Believe it, then rate it. Establish what it actually supports — usually a coarse scale relationship — and compare that against what a paying customer already learns from the reply's exact dimension and the error text's input ceiling. If it only corroborates that picture, it is a low-severity note, recorded rather than remediated.
saying these in an interview costs you the question
- Escalates a flaky measurement without checking the apparatus
- Dismisses it without asking how it was collected
- Accepts blocked collection across different time windows
- Reports medians with no dispersion
- Rates a family-level inference as high severity