A paid text-embedding endpoint returns only vectors — what does its reply metadata leak?
answer
- the reply has a shape, not only content
- errors are documentation nobody meant to write
- latency grows with input length
- a family and a shape, not a spec
basics
~20 sA reply's fixed vector length and dtype, latency that grows with input length, and error text for oversized input reveal an output space, a rough scale and an input ceiling — a family and a shape, never parameters.
solid answer
~40 sAn attacker who is already paying for embeddings gets a second channel for free: the shape of each reply and the way the service behaves around it. A fixed-length float array pins the output dimension and the numeric precision exactly. Error strings for oversized, empty or malformed input frequently name a token ceiling and hint at how text was split before encoding. Wall-clock latency that rises with input length, once averaged over enough calls to beat network and queueing jitter, indicates rough compute scale and whether cost grows about linearly with length or faster. All of it is coarse. It eliminates hypotheses about the architecture family and pins the output space; it is not a specification of depth, width, or training hyperparameters, and nothing here recovers a single weight.
go deeper
Be ready to name the three things visible without extra access — the reply's shape, the latency, and the error text — and to say that together they give a family and an output space rather than parameters.
Explain what each signal pins down and how tightly: dimensionality and dtype exactly, an input ceiling nearly so, compute scale only in a statistical sense after averaging many calls.
Show you can rate this exposure honestly. It is cheap for the attacker, impossible to remove entirely without breaking the product, and worth a bounded prior — so it rarely deserves a high-severity finding on its own.
Own the framing that the purchasable replies, not the model's blueprint, are the asset under threat, and resist spending engineering effort on obscuring a shape the product must expose anyway.
## The vantage this question describes The adversary here is an ordinary paying customer of a hosted text-embedding service: text goes in, a fixed-length float vector comes back, billed per thousand inputs. They are not reading the answers more cleverly — that is a different attack. They are reading everything *around* the answers: how long a call took, what shape came back, and what the service says when a request is rejected. This is **incidental behaviour**, and the reason it matters is that the customer is entitled to all of it. It arrives with traffic they were buying anyway. ## The three signals, and what each one actually pins down **Output shape.** The returned array has a length and a numeric type. That is not an estimate — it is exact, and it is disclosed on the first call. It fixes the dimensionality of the embedding space the product is built on, and the precision the service is willing to hand out. Dimensionality is a strong hint at the family and rough scale of the encoder behind it, because published encoder families cluster on a small number of conventional widths. **Error text.** Requests that are too long, empty, wrongly typed, or in an unexpected encoding produce error strings, and error strings are documentation nobody meant to write. A message that names a maximum token count discloses the tokenizer's input ceiling. A message that distinguishes "too many tokens" from "too many characters" discloses that tokenisation happens before the length check. A message that behaves differently across scripts or languages discloses something about how the text is normalised and segmented. None of this is secret in any deliberate sense; it is simply informative. **Timing.** Latency per call varies with input length. Averaged over many calls, the *shape* of that relationship carries information: roughly linear growth in tokens versus something steeper, a step at a particular length, or a flat region up to some point followed by a rise. That tells an attacker something about compute scale and about batching or truncation behaviour. It is the noisiest of the three signals by a wide margin, and how much of it survives measurement noise is a separate question in its own right. ## What the attacker walks away with A **prior**. Concretely: an output dimension known exactly, an input ceiling known nearly exactly, a preprocessing story partly constrained, and an architecture family narrowed from "anything" to "one of a few plausible ones at roughly this scale". Hypotheses have been eliminated. That is genuinely useful to someone assembling a functional copy — it saves them from sizing their own model absurdly or from wasting paid calls on inputs the endpoint will truncate — but it is a head start, not the thing itself. ## What it is not It is not parameter recovery. No number of latency measurements yields a weight. It is not a specification: layer counts, hidden widths, attention configuration, training corpus, objective and hyperparameters are not recoverable this way, and claiming otherwise is the standard overstatement in a report of this kind. It is also not the expensive part of copying an endpoint — the fidelity of a copy is set by which inputs get bought and how well the replies are fitted, not by whether the copy shares the target's blueprint. ## Where the line sits Two neighbouring subjects look identical and are not. Timing as a *cryptographic* concern — a comparison whose duration depends on a secret, and the constant-time discipline that answers it — is application-security material about secrets, not about model shape. And latency, batching and autoscaling as *serving* behaviour to be tuned against a service objective is a capacity subject; here the same numbers are being read by somebody who does not own the service and cannot see inside it. The distinguishing feature of this leaf is the actor: someone outside, inferring the shape of a model they do not own, from replies they paid for. ## How to talk about it in an interview Say what each signal pins down and how precisely, and then say what it does not reach. The honest summary is that this channel is cheap, always present, hard to remove without degrading the product's usability, and worth a bounded amount — a prior on the model's shape, not the model.
- Which of those three signals is exact, and which is only statistical?The returned vector's length and numeric type are exact and free — one call settles them. The input ceiling from error text is close to exact where the message names a number. Timing is purely statistical: any single call's latency is dominated by network and queueing variation, so only a difference that survives averaging over many calls means anything at all.
- Does suppressing detailed error messages close this channel?It narrows one signal and closes nothing. Generic errors remove the stated token ceiling, but the ceiling is still discoverable by the pattern of accepted versus rejected inputs, and the vector's length and dtype are unavoidable — the product exists to return them. It raises the attacker's effort slightly and costs legitimate integrators real debugging time, so it is a trade rather than a fix.
- Why does output dimensionality hint at an architecture family at all?Encoder families are published at a small number of conventional widths, and a hosted embedding product usually exposes the encoder's own output width or a fixed projection of it. So an observed width is compatible with only a few families at a few scales. It is an elimination argument over a small hypothesis set, not a derivation.
Like judging a sealed parcel by its weight, its dimensions and the courier's rejection slip: you learn the size and class of what is inside without ever opening it.
saying these in an interview costs you the question
- Claims timing recovers layer counts or weights
- Treats a single call's latency as a measurement
- Says nothing leaks because only vectors are returned
- Confuses this with a cryptographic timing side channel