skip to content

A triage record shows a wrong routing and an output screen verdict of pass — what does that prove?

level: seniorimportance: nice to knowfreq 30%

answer

  1. each log proves one narrow claim
  2. a pass is a score, not a verdict
  3. which values, not who chose them
  4. the digest shows outcomes only
  5. pair the answer with the lifted fields

basics

~20 s

Only that the answer scored below threshold on the screen's harm categories and that the parser lifted certain values from it. It does not show who supplied those values, and it does not show that the answer's content was benign or malicious.

solid answer

~50 s

Read each record for exactly the claim it supports. The screen's pass proves a score below threshold on a fixed category set — nothing about the consumer. The routing record proves *which* values the parser lifted and acted on, not *who chose them*: a parser has no notion of provenance, so a value that originated in a customer's ticket body and one the assistant inferred look identical once lifted. The morning digest proves a human saw an outcome, not the string that produced it. What actually decides the question is a comparison nobody keeps by default: the answer text as generated set against the fields the parser extracted from it. If a lifted value traces back to a span of the ticket body rather than to the assistant's own reasoning, you have the finding; without that pairing, all three records are consistent with an ordinary day.

code

json · 15 lines
json
{
  "answer_id": "a-84213",
  "output_screen": {
    "scored_categories": ["harassment", "hate", "self_harm", "violence"],
    "max_score": 0.02,
    "verdict": "pass"
  },
  "triage_parser": {
    "read_same_answer": true,
    "lifted": { "queue": "billing-escalation", "priority": "P1", "owner": "on-call" },
    "provenance_of_lifted_values": null,
    "source_region": "[span the parser's grammar read as fields - elided]"
  },
  "morning_digest": { "row": "ticket 84213 -> billing-escalation / P1", "includes_answer_text": false }
}

go deeper

for a junior

Recall that a classifier log records a score against categories, and a routing log records a decision. Neither records where the values in the decision came from.

for a middle

Explain why a parser has no provenance for lifted values: it sees one string from one producer and assigns meaning by position, so inferred and supplied values are indistinguishable.

for a senior

Show you can triage from incomplete records — state what each artefact proves, name the pairing that would decide it, and give a reproduction rate rather than a bare claim of success.

for a principal

Be ready to argue for retaining the answer-to-extracted-field pairing as an evidence question, and to concede honestly what a morning digest of outcomes can and cannot be claimed to buy.

## Three records, three narrow claims An incident in an unattended triage workflow usually arrives as a handful of log lines and a complaint that a ticket went to the wrong queue at the wrong priority. The temptation is to read the logs as a story. They are not a story; each one supports a single narrow claim, and the gap between them is where this class lives. **The output screen's verdict.** A pass means the answer scored below threshold on the harm categories that screen has. It is not a safety verdict and it is not evidence about the consumer, which the screen never saw. Equally, had it been a block, that would prove the text scored high on a category — not that anything was attempted. **The routing record.** It shows the values the parser lifted and the decision written from them. A parser assigns meaning by position; it has no provenance field, because as far as it is concerned the answer is one string from one producer. So a value that arrived because a customer wrote it into a ticket body and a value the assistant inferred from the ticket are indistinguishable in this record. *Which* values were used is proved; *who chose them* is not. **The morning digest.** It proves a human looked at outcomes. It shows ticket, queue, priority — a list of decisions that all look like ordinary triage decisions, because that is exactly what a decision made from a planted value looks like. The digest cannot surface this class, not because the reviewer was careless, but because the artefact does not contain the thing that would give it away. ## The pairing that decides it The evidence that resolves the question is a comparison between two artefacts that are usually stored separately or not at all: 1. The answer as generated, before any consumer touched it. 2. The field values the parser extracted from it. Set them side by side and ask where each lifted value came from in the answer's text, then ask where that text came from in the ticket. A lifted value that traces back to a span of the customer-supplied body — rather than to the assistant's own summary of it — is the finding. Without that pairing you are reasoning from three records that are all individually consistent with nothing having happened. This is also why the class is under-reported rather than rare. The artefact that would show it is the one nobody retains, because retaining generated answers alongside extracted fields is storage nobody budgeted for a workflow whose whole selling point was that it runs without supervision. ## Direction-of-claim discipline The habit worth drilling, because interviewers probe it directly: | Record | Proves | Does not prove | |---|---|---| | screen verdict: pass | scored below threshold on its categories | the answer was harmless, or safe for a consumer | | routing record | which values were lifted and acted on | who supplied those values | | morning digest | a human saw the outcomes | a human saw the answer text | | one reproduction | the construction worked once, on this deployment | the finding is reliable, or general | Every row of that table is a sentence a candidate either says precisely or blurs. Blurring the last row is the most costly in practice: a probabilistic producer means a single reproduction is a data point with a success rate attached, and a write-up that reports it as "reproduced" without the rate will be argued down by the owner and deserves to be. ## What to say when asked "so was it screened?" The answer is yes, and the screen returned an accurate score for the question it was asked. That is not a defence of the workflow and not an indictment of the screen. It is a statement that the coverage everyone is pointing at measured a property that this construction does not touch, and that the record which would have measured the relevant one was never written.

  • Why does the parser's record have no provenance for the values it lifted?
    Because there is nothing to record. The parser receives one string from one producer and extracts values by position; the distinction between a value the assistant inferred and one a customer wrote into a ticket body does not exist at that layer. Provenance would have to be carried from upstream, and the workflow was built on the assumption that the answer is the workflow's own output.
  • What would you ask to be retained so this class is decidable next time?
    The generated answer paired with the field values extracted from it, keyed together. That single pairing turns an undecidable set of logs into a check anybody can run: for each lifted value, where in the answer did it come from, and where in the ticket did that come from. Note this is an evidence question, not a fix — it changes what you can conclude, not what the consumer does.
  • The owner says the digest was reviewed every morning and nothing looked odd. How do you respond?
    That is consistent with the finding rather than against it. A routing decision made from a planted value looks exactly like a routing decision made from a real one, because the digest carries the outcome and not the string. The review was not careless; the artefact reviewed does not contain the discriminating information.

saying these in an interview costs you the question

  • Reads a passing screen verdict as evidence the answer was benign
  • Treats a routing record as showing who chose the values
  • Says the digest review would have caught it
  • Reports one successful reproduction as a reliable finding
  • Assumes the parser stores provenance for extracted fields

context