skip to content

A seller's traffic pushes 12% of ranking responses onto the default ordering while uptime stays green. Which property broke?

level: seniorimportance: should knowfreq 34%

answer

  1. uptime measured the host, not the model
  2. ask for the written floor first
  3. then ask who the fallback favours
  4. degradation can be the mechanism
  5. flat aggregates prove preservation, not innocence

basics

~20 s

You cannot say from the traffic alone. Compare the fallback share against the operator's written floor for model-served responses, then ask who the fallback ordering favours. If the seller gains under it, the finding is integrity.

solid answer

~50 s

Two questions decide it, and uptime answers neither. First, **against what tolerance**: model availability is the share of responses the model actually produced, so it needs a written floor — if the operator's floor is 97% model-ranked and you are at 88%, availability is broken regardless of a green dashboard. Second, **who benefits**: if the seller's own listings sit near the top of the plain fallback ordering and far down the model's, then the degradation is the *mechanism* and the chosen ordering is the payoff, which makes this an integrity finding that a fallback-rate alarm will never characterise. The two readings go to different owners and carry different fixes, so the report has to commit to one and say why. And a green uptime figure establishes only that the process stayed up — it was never a measurement of the model at all.

code

text · 10 lines
text
window                        last 7 days
service uptime                99.98%
p99 latency                   within SLO, flat
responses served by model     88.1%
responses served by fallback  11.9%
operator's written floor      >= 97% model-served
...
seller S, median listing rank
  under model ordering        #41
  under fallback ordering     #6

go deeper

for a junior

Know that a service can be fully up while the model behind it is not being used, and that 'uptime is green' therefore does not answer a question about the model.

for a middle

Explain why the fallback share is the number that matters and why it means nothing without a written floor. Be able to state model availability as a rate over responses, not a binary.

for a senior

Show the adjudication: tolerance first, beneficiary second, aggregate metrics treated as weak evidence. An interviewer wants you to name the fix that the wrong reading would produce and why it fails.

for a principal

Own the fact that each of the three goals needs its own tolerance, agreed in advance with a named owner, and that an engagement finding a missing tolerance has found something worth more than the incident.

## The trap in the question Everything in the scenario that is easy to measure — uptime, latency, error rate — is measuring the *host*, and none of it is measuring the *model*. The one number that speaks to the model, the share of responses served by the fallback ordering, is uninterpretable until somebody has written down what share is tolerable. So the honest first answer is not "availability" but "against what tolerance, and who benefits". ## Availability, for an ML system, needs a written floor A ranking service with a plain-list fallback is *designed* to degrade rather than fail. That is good engineering and it is exactly why the usual signals go quiet: the process answers every request, inside its latency budget, with a valid response. The property that is actually at stake is "the served response was produced by the model", and it is a rate, not a binary. That makes the security goal a threshold somebody owns: *at least X% of responses are model-ranked, measured over Y*. Without X written down in advance, 88% is a number with no verdict attached, and the conversation collapses into opinion. With it, the adjudication is mechanical — and note that the three goals each get their *own* tolerance. The share of served responses the operator will accept losing is a different number, set by different people, from the share of orderings it will accept being wrong. ## The second reading, which the availability alarm cannot see Degrading a model is not always the point of degrading a model. A fallback ordering is a *different function*, usually a simpler and more predictable one — recency, price, or a plain relevance score. Anyone who can see both orderings can work out which one favours them. If a seller's listings sit at the bottom of the model's ordering and near the top of the fallback, then pushing traffic onto the fallback is not an attack on availability that happens to help them; it is an attack on the *ordering*, with degradation as the delivery mechanism. The two readings differ in almost every way that matters downstream: | | Availability reading | Integrity reading | |---|---|---| | What was bought | the model's output was absent | a specific ordering was served | | Measured as | share of responses not model-ranked | rank of chosen listings, on affected requests | | Stops when | the fallback rate returns under the floor | the beneficiary stops gaining, whatever the rate | | Owner | the serving/platform team | the marketplace-integrity team | | Wrong fix | capacity, so the model absorbs the traffic | rate limits, which leave the incentive intact | The "wrong fix" row is the practical reason this adjudication is not academic. Reading an integrity attack as an availability incident leads a competent platform team to add capacity, which makes the fallback rate go away and leaves the adversary's incentive completely intact. They come back with a cheaper way to trip the same fallback. ## How to decide, in order 1. **Get the floor.** Ask the operator for the written tolerance on model-served responses. If it does not exist, that absence is itself the first finding, and the engagement should produce the number rather than argue about the incident. 2. **Compare, and say so plainly.** 88% against a 97% floor is a breach of a stated goal; 88% against no goal is an observation. 3. **Split the traffic by beneficiary.** Look at where the suspected beneficiary's listings land under each ordering, on the affected requests only. A large, consistent gap in their favour is what turns this into an integrity finding. 4. **Check the aggregate.** If overall ranking quality metrics are flat, do not read that as "nothing happened". An adversary pursuing a chosen ordering has every reason to leave the aggregate alone — it is the number being watched. Aggregate stability proves the adversary preserved it, not that the outputs were correct. 5. **Write one verdict, and name the other reading.** Commit to the property you can evidence, and record explicitly what the other reading would require to confirm, so the second owner can pick it up. ## The claim you must not make "Uptime was 99.98%, so availability held" is the sentence to strike. Process liveness was never a measurement of whether the model served. In the same spirit, "the fallback ordering is a valid response, so nothing is wrong" confuses a designed behaviour with an acceptable rate of that behaviour; the fallback exists precisely because losing the model sometimes is survivable, and the whole question is how much of the time.

  • The operator has never written down a tolerance for the fallback rate. What do you deliver?
    The missing tolerance, as the first finding. Without it there is no goal to falsify, so no rate can be called a breach and the incident is argued on instinct. I would propose a floor derived from what the fallback ordering costs the business per request, get it owned by the team that runs serving, and only then rate the observed 88% against it. A number nobody agreed to in advance is not a security goal.
  • Overall ranking quality metrics are unchanged through the window. Does that clear the integrity reading?
    No — it is weak evidence in the wrong direction. An adversary who wants one ordering has every reason to leave the aggregate flat, because that is the number under watch. Flat aggregates establish that the aggregate was preserved, not that no outputs were manipulated. The test is per-request and per-beneficiary: where do the suspected listings land on the affected requests, compared with unaffected ones.
  • Why does calling this an availability incident risk making it worse?
    Because the obvious remedy is capacity, and capacity removes the symptom while leaving the payoff untouched. If the adversary was buying a favourable ordering, absorbing their traffic simply raises the price of the same win, and they return with a cheaper way to trip the fallback. Fixes aimed at the wrong property tend to be expensive, effective-looking on the dashboard, and irrelevant to the incentive.

saying these in an interview costs you the question

  • Cites uptime as evidence that availability held
  • Rates the fallback share with no written tolerance to compare against
  • Never asks who the fallback ordering favours
  • Reads flat aggregate quality as proof nothing happened
  • Treats a designed fallback as always an acceptable outcome
  • Files one verdict without recording the alternative reading

context