skip to content

Your network-flow alert-triage model stopped flagging one traffic class and request logs look normal — what does that rule out?

level: seniorimportance: should knowfreq 41%

answer

  1. wrong telemetry for this class
  2. the change never sent a request
  3. compare the last two training tables
  4. if it was never recorded, say so

basics

~20 s

Almost nothing. Request logs establish which queries arrived, not what the weights encode. A change written into the disposition queue or the training table produces no request at all, so the evidence that bounds it is the upstream write record.

solid answer

~50 s

They rule out the families that had to come through the endpoint, and only for requests that were actually logged. This model is retrained from analyst dispositions and a shared feed, so an adversary who wants that traffic class ignored never sends a request: they write one artefact a later training run reads. So I stop treating gateway telemetry as the evidence and go to the writes. Who held write on the disposition queue and the feed ingestion between the last two training tables; whether anything compares one training table with the previous one before a retrain consumes it; how long a write can sit before anyone would look. The same evidence separates the benign explanation — drift, or a change in how dispositions get collected — from a deliberate one. If none of it was recorded, the honest report is that the question is unanswerable, not that nothing happened.

go deeper

for a junior

Know that a model's behaviour can change because of what it was trained on, and that request logs only describe requests. Recognising which telemetry answers which question is the level-appropriate takeaway.

for a middle

Be able to list what feeds a retraining loop and explain why a change in any of those sources shows up in the weights rather than in the request path.

for a senior

Show the investigation: compare consecutive training tables, enumerate who held write on each contributing path over the window, establish dwell, and separate drift from a deliberate change with the same evidence.

for a principal

Be ready to say what the organisation can honestly claim when the records do not exist, and to fund the record — attributable writes and retained training-table snapshots — rather than hardening the surface the adversary never touched.

## The shape of the finding A model that triages alerts over network flow records is retrained from what analysts decided about previous alerts, plus a shared feed the pipeline ingests. One class of traffic has quietly stopped being surfaced. Nothing in the request logs looks unusual. The first move is to notice that the request logs were never going to look unusual, whatever the cause. ## What clean request telemetry actually rules out Request logs are a record of requests. They can speak to attacks that had to arrive as a query — somebody probing the endpoint, buying score information, or feeding it crafted inputs — and even then only for the subset of traffic that was logged and retained. They say nothing at all about how the weights came to encode what they encode. An adversary who wrote into the disposition queue, into records the ingestion job pulls from the shared feed, or into the assembly step that builds the training table, made no request. There is no row in the gateway log where their entry would be. "The logs are clean" is therefore not evidence of anything about this class, and — this is the part candidates miss — it is not evidence for the benign explanation either. It does not support drift over an attack; it is simply silent. ## The evidence that would bound it Three records, in order of how much they buy. **A diff between training tables.** Two consecutive training tables, compared. Did the labels or dispositions attached to that traffic class move between them, and by how much? This is the record that tells you whether the model's change originated in what it was trained on at all, or in something downstream like a threshold or a deployment. **The writer set over the window.** Which principals held write on each contributing path between those two tables: the disposition queue, including whatever shared account an outsourced triage tier authenticates as; the ingestion job's identity and, upstream of it, whoever can contribute to the shared feed; and whoever can change the assembly step. Frequently the honest answer is a set rather than a name, and the size of that set is itself the finding. **Dwell.** How long a write can sit before anything would look at it. If the answer is "until someone notices the model behaving oddly", then the window is the whole interval since the last time anybody checked, and the investigation's scope has to be that wide. ## Getting the direction of every claim right Several tempting statements point the wrong way, and an interviewer is listening for them. Clean gateway logs prove which requests arrived, not what the model became. A statement that no unauthorized access occurred bounds unauthorized access, and this class needs none — every write in the story is authorized. Aggregate detection metrics holding flat prove that overall performance was preserved, which is what somebody aiming at one traffic class would want; it is not evidence that nothing changed. And a pipeline that ran without errors proves it read what it was configured to read, not that what it read was chosen honestly. ## Benign explanations deserve the same evidence Most of the time this is not an attack. Analyst dispositioning habits change; a tier gets outsourced or retrained; the shared feed alters its coverage; the traffic class genuinely became rarer. That is an argument for the same records, not against them. A training-table diff plus a write record distinguishes all of these, and it lets you close the finding with a reason rather than with an absence. This is also why "reproduces once in five tries" style triage matters here: a behavioural change observed through the deployed system is noisy, and the durable evidence lives in what the training run consumed, not in re-running the observation. ## What to say when the records do not exist Often they do not. There is no retained snapshot of previous training tables, dispositions carry no per-writer attribution, and the ingestion job's identity is shared. The correct output then is a scope statement and a gap: these are the paths that fed the model, this is the set of principals that could have written to each, nothing reviewed the change before the retrain consumed it, and therefore the question cannot be answered from records that exist. Saying "we found no evidence of tampering" when nothing was ever recorded is the failure mode; the absence of evidence is a property of the instrumentation, not of the pipeline. The remediation that follows is about the record, not about the endpoint: attributable writes on each path, retained training-table snapshots with a diff someone owns, and a named owner per write path. Adding controls to the API in response to this finding would be treating a surface the adversary never used.

  • What would you have wanted recorded before this happened?
    Attributable writes on each upstream path, so that a change has a principal against it, and a retained snapshot of every training table with a diff against the previous one. Together they convert "we cannot tell" into a bounded answer: here is the set of principals who could have written, and here is exactly what changed between the two runs.
  • Could this be ordinary drift rather than an adversary?
    Easily, and usually it is. The point is that the same evidence answers both questions. A diff between the two training tables shows whether that class's labels moved, and the write record shows who moved them. Gateway logs answer neither, so their cleanliness is not support for the drift explanation any more than it is for the attack one.
  • Your provider states that no unauthorized access occurred. Does that settle it?
    It bounds unauthorized access, and this class needs none. Every write in the story is authorized: an analyst dispositioning an alert, an ingestion job writing feed records, a scheduled job assembling a table. The statement you actually need is about the authorized writer set and whether any write was reviewed, which is a different claim and usually one nobody has instrumented.

saying these in an interview costs you the question

  • Treats clean gateway logs as proof nothing happened
  • Searches only for a malicious request
  • Assumes authorized writes cannot be the cause
  • Concludes drift without comparing training tables
  • Reports no evidence of tampering when nothing was recorded

context