skip to content

Integrity, Uptime, Secrecy

One crafted input is a different finding under each goal, and a model's confidential asset is usually its training data, not its weights. Interviewers ask because the host's reading does not transfer.

on this pageshow

explore

questions

3

Against a deployed ranking model an outsider can only query, which three goals can they choose between?

level: juniorimportance: must knowfreq 76%

answer

  1. success is undefined until you name it
  2. three properties, not three techniques
  3. wrong, expensive, or revealing
  4. the third needs no weight file

basics

~20 s

Three: a wrong output (integrity), a degraded or missing output (availability), or a fact the model should not reveal, usually about its training data (confidentiality). The goal defines success, so it is named before any technique.

solid answer

~50 s

Three, and they are not three flavours of one attack. **Integrity**: make the model produce a wrong output on inputs the adversary cares about — a mediocre listing ordered above a better one. **Availability**: stop the model from serving its real output at all, for example by pushing enough traffic onto the cheap fallback ordering that ranking is effectively gone. **Confidentiality**: learn something the model should not reveal — most often a fact about the data it was trained on, which is reachable through ordinary query access with no copy of the weight file. The goal is named first because it fixes what counts as success and therefore what the engagement measures: a share of responses that fell back, an ordering the adversary chose, or one bit about one record. The adversarial-ML taxonomies (NIST AI 100-2, for instance) group attacker goals the same way.

go deeper

for a junior

Be ready to name the three properties for a model and give one concrete example of each against the same system. Say plainly that the goal is chosen before the technique, because it decides what counts as success.

for a middle

Explain how each property is measured for a model rather than for a host: a share of outputs that are wrong, a share of responses the model did not produce, and an advantage over a base rate. Note that host monitoring is silent on all three.

for a senior

Show that you would refuse a finding written without its goal and its tolerance. An interviewer expects you to say which evidence you would collect for each of the three and who in the organisation acts on it.

for a principal

Own the argument that the three goals carry different tolerances that someone must write down in advance, and that an engagement with finite time tests one or two of them deliberately rather than sampling all three shallowly.

## Why the goal comes before the technique A finding against a machine-learning system is not "the model can be manipulated". It is "an adversary with *this* access, spending *this much*, achieved *this outcome*, measured *this way*". The outcome half of that sentence is the adversary's **goal**, and until it is named, "success" is undefined — the same crafted request can be a serious finding, a curiosity, or nothing at all depending on which property it was aimed at. Take a concrete setting and hold it for the whole discussion: a marketplace search service whose ranking model orders third-party sellers' listings, with a plain, non-personalised fallback ordering the service can drop to when the model is slow or unavailable. The adversary is a competitor with an ordinary paid seller account. They see the rendered ranked list and their own dashboard metrics — never a score, never a probability vector — and they can upload catalogue content of their own. ## The three properties, as a model expresses them **Integrity — a wrong answer.** The service returns an output the operator would not endorse. Here that is an ordering: the adversary's listing surfaced above listings that should outrank it, or a rival's listing pushed down. Success is measured on the *outputs the adversary cares about*, not on the model's aggregate quality. A model whose overall ranking metrics are unchanged can still be fully broken under this goal, and an adversary who cares about integrity has every reason to keep the aggregate steady, because that is what the operator is watching. **Availability — an answer that costs too much, arrives too late, or never arrives as a model output at all.** For an ML system this is subtler than a process falling over. The service can be up, healthy and inside its latency budget while the *model* is effectively absent, because every request is being served by the fallback ordering. Availability of an ML system is a statement about the share of requests that received the model's real output, so it needs a written floor — "at least 97% of responses are model-ranked" — before anyone can say it was broken. **Confidentiality — a fact the system should not have revealed.** This is the one that surprises people, because for a model the confidential asset is usually not a file. It is the **training data**, and secondarily the learned function itself. A model that was fitted on data reveals something about that data through its outputs, so an adversary who can only send queries can pursue this goal without ever obtaining the weights. Success is not "a file was copied"; it is "the adversary now knows something about a record, a person, or a corpus that they did not know before, at better than chance". ## Why the host's reading does not carry over The familiar three properties are usually reasoned about over a *host*: a file was altered, a service went down, a database was read. An ML system re-expresses each of them over the model's behaviour, and the mapping is not one-to-one: | Property | Host reading | Model reading | |---|---|---| | Integrity | stored data or code was altered | a *specific output* is wrong, with the stored artefacts untouched | | Availability | the process is not answering | the process answers, but not with the model's output | | Confidentiality | data at rest was read or copied | the *training data* is inferred from outputs, with nothing copied | The practical consequence: host-level monitoring is silent on all three. Uptime dashboards do not measure model availability, integrity checks on artefacts do not measure output integrity, and egress controls on the weight file do not measure disclosure through the query interface. ## What each goal costs the adversary, roughly The three are not equally priced from the same seat. Integrity against a ranked list is measurable by the adversary directly — they can see where their own listing landed, so they get free feedback on every attempt. Availability requires sustained volume or unusually costly inputs, and it is loud. Confidentiality is the quietest and the slowest: results come back as a statistical advantage over a base rate rather than as an obvious win, so it needs a plan for what the advantage would prove before the first query is sent. ## The common confusion An adversary may of course pursue more than one goal, and one technique can serve several. That does not make the distinction cosmetic — it makes naming it *more* important, because the measurement, the evidence you must collect, and the person who has to act on the finding are different in each case. If a report cannot say which property was broken and against which stated tolerance, it has not established anything.

  • Why does naming the goal change how you report one single finding?
    Because the measurement changes. The same crafted listing is an integrity finding if it moved the ordering the attacker wanted, an availability finding if enough of them push responses onto the fallback ordering, and a confidentiality finding if the model's replies revealed something about the training corpus. Each is measured against a different tolerance and lands on a different owner, so a report that stops at "the model can be manipulated" hands the operator nothing to decide on.
  • Which of the three is hardest to evidence from outside, and why?
    Confidentiality. Integrity and availability produce outcomes the adversary can see directly — where the listing landed, whether the fallback ordering came back. A privacy result is a statistical advantage over a base rate, so it needs a baseline, a sample large enough to separate the advantage from chance, and an argument about what the inferred fact actually means about a person. Without those three, a suggestive number is not a finding.
  • Can an adversary pursue more than one of the three at once?
    Yes, and they often do — degrading a model's output can be the mechanism for a chosen ordering rather than the point of the exercise. The goals stay distinct because they are measured differently: you still have to say which outcome the adversary was buying and which was a side effect, or you cannot tell whether the tolerance that was exceeded is the one that mattered.

Naming the goal is like a burglary report saying whether something was taken, broken, or merely photographed. The same broken window supports all three, and only the outcome tells you what was lost.

saying these in an interview costs you the question

  • Treats the three goals as three attack techniques
  • Says confidentiality means protecting the weight file
  • Calls any wrong output a successful attack, with no target named
  • Measures a model's availability by the service's uptime
  • Infers the goal from the technique instead of the outcome
  • Reports 'the model can be manipulated' with no measurement

context

open as a page

A team says an attacker who never obtains their model's weight file cannot breach its confidentiality. Why is that wrong?

level: middleimportance: should knowfreq 59%

basics

~20 s

Because the confidential asset is not the file. It is the training data the model absorbed, and secondarily the learned function itself, both of which leak through ordinary query access. A privacy attack can succeed against a model the adversary never obtains.

open as a page

A seller's traffic pushes 12% of ranking responses onto the default ordering while uptime stays green. Which property broke?

level: seniorimportance: should knowfreq 34%

basics

~20 s

You cannot say from the traffic alone. Compare the fallback share against the operator's written floor for model-served responses, then ask who the fallback ordering favours. If the seller gains under it, the finding is integrity.

open as a page