skip to content

An input screen and an output screen both logged allow for a turn in which the assistant obeyed an encoded span — what does that record prove?

level: seniorimportance: should knowfreq 40%

answer

  1. a verdict is about a surface
  2. allow means below a threshold
  3. the log shows which call, not who chose
  4. the clean record is part of the payoff
  5. provenance stops at the tenancy boundary

basics

~20 s

Two allow verdicts prove that two surface forms scored below threshold on the labels those screens carry. They say nothing about what the generator reconstructed, which call it made, or which argument values it chose.

solid answer

~50 s

The record certifies scores, not behaviour. The input verdict says the characters handed to the screening model fell below a threshold on the labels that model carries; the output verdict says the same about the emitted response. Neither screen ever saw the text the generator reconstructed for itself, so the pair is silent on what the turn actually did. To triage it you go to the things that record behaviour: which tool calls ran and with what argument values, whether those values correspond to anything the user asked for, and where each span in the context came from. And note what the clean record is worth to whoever placed the span — review is sampled, and a turn with no label attached to it is a turn nobody has a reason to open, so the absence of an incident is part of the payoff rather than evidence against one.

code

json · 14 lines
json
{
  "turn_id": "...",
  "tenant": "tenant-4b",
  "input_screen":  { "verdict": "allow", "top_label": "none", "score": 0.03 },
  "context_sources": [
    { "kind": "user_turn",   "chars": 74 },
    { "kind": "tool_result", "tool": "record_lookup", "field": "notes",
      "value": "[transport-encoded span elided, 212 chars]",
      "written_by": "host-account-91", "written_at": "..." }
  ],
  "tool_calls":     [ { "name": "record_update", "args": { "id": "...", "status": "..." } } ],
  "output_screen": { "verdict": "allow", "top_label": "none", "score": 0.05 },
  "review_sampled": false
}

go deeper

for a junior

Remember that an allow verdict is a score on a piece of text, so it tells you the characters scored low and nothing about what the assistant went on to do.

for a middle

Explain why both verdicts can be correct and uninformative at once: each screen scored a surface, and neither saw the text the generator reconstructed for itself.

for a senior

Demonstrate the triage: enumerate context sources and their paths, compare tool-call arguments against the user's request, and state clearly what a single reproduction against one deployment does and does not license you to claim.

for a principal

Own the consequence for assurance: if review is spent on flagged turns and the flags are scores on surfaces, then a clean record is an artefact of what was measured, and any claim built on the absence of incidents inherits that limit.

## Read the record for what it is A screening verdict is a score on a string. Written out as an evidentiary claim, an `allow` from an input screen says: *the characters this model was handed scored below the configured threshold on the labels this model carries.* Every word of that is load-bearing. It is not a statement about the request's intent, about the reconstruction the generator later produced, or about anything that happened after the score. The output verdict is the same kind of claim about a different string — the emitted response — and it inherits the same asymmetry: it scores the surface that was produced, while whatever consumes the response resolves it. ## What the pair is silent about | Artefact | Proves | Does not prove | | --- | --- | --- | | input screen: allow | the input surface scored low on that screen's labels | that the input carried no directive | | output screen: allow | the emitted surface scored low | that the turn's effect was benign | | a tool call in the log | which call ran with which argument values | who chose those values | | a field's provenance | which session and account wrote it, and when | that anyone meant it as an instruction | | no sampled review | nothing was flagged for review | that there was nothing to review | The row people most often over-read is the tool call. A call record is strong evidence of *what happened* and no evidence of *why*: the arguments were emitted by the generator, and the generator's reason for choosing them is not in the log. A call whose arguments do not correspond to anything in the user's turn is the anomaly worth chasing precisely because the log cannot explain it. ## Where the span came from, and where you stop being able to tell If the span arrived in a field of a tool's return value, the provenance you can obtain is the host product's, not the assistant's: some account wrote that note at some time, long before this turn, and no one classified it as input then. For an assistant feature a vendor embeds inside other companies' products, that trail usually stops at the tenancy boundary. The vendor holds its own screen verdicts, the prompt it assembled and the calls it made; it does not hold the host product's data flow. "Where did this text come from" becomes a question only the tenant can answer, and answering it requires the tenant to accept that one of its stored fields was an input path — which is exactly the framing the record does not prompt anyone to adopt. ## The clean record as the payoff It is worth naming plainly that the absence of an incident is part of what was obtained. Review capacity is finite and is spent on flagged material. A turn that both screens allowed, that carries no label and that no sampler selected is a turn with nothing pointing at it. The moderation log and the tenant's audit trail will both, quite truthfully, report that no policy label fired — and a reader will hear that as *nothing happened*, which is a far stronger statement than the records support. Whoever placed the span gets the behaviour and the silence together, and the silence is often the more durable half, because it determines whether anyone looks again. ## How to triage the turn anyway The useful moves all bypass the verdicts: - **Start from effects, not scores.** Enumerate the calls the turn made and compare their arguments with what the user asked for. Divergence between the two is the signal the screens cannot give you. - **Enumerate the context sources.** List every span that entered the window and the path it came in on. The one that no stage treats as input is the one to look at first. - **Ask what the surface was.** If a span in the context is in a representation the pipeline never expands, the input verdict was scored on that representation, and the score is uninformative by construction. - **Be honest about reproduction.** Getting the same behaviour once tells you the construction worked once against one deployment. With a probabilistic generator, report attempts alongside successes; a single reproduction is an existence proof and a rate is a different claim. ## What the finding should say The defensible finding is about the asymmetry and the path: which component scored which surface, which component resolved it, and how the span reached the context window. That is the part with a long shelf life. The particular representation used is the part that gets covered first and teaches a reader least, and a finding that leads with it invites a fix that closes one form and leaves the structure untouched.

  • What in the turn actually tells you what the model did?
    The tool calls with their argument values and the emitted response — not the screen verdicts. A call record proves which call ran and with which values, and nothing about who chose them, so the thing to look for is arguments that do not correspond to anything in the user's request. That divergence is the only part of the record that points at a cause.
  • Why is the clean record valuable to whoever placed the span, and not just a side effect?
    Because review is sampled and spent on flagged material. A turn with no label attached has nothing pointing at it, and the audit trail's truthful statement that no policy label fired will be read as nothing happened. The behaviour and the silence are obtained together, and the silence determines whether anyone ever looks at the turn again.
  • How does a multi-tenant embedding limit what you can reconstruct afterwards?
    The vendor holds its own verdicts, the prompt it assembled and the calls it made, but not the host product's data flow — who wrote the field, when, or what the backend did to it on the way out. Provenance stops at the tenancy boundary, so establishing the origin of a span requires the tenant to investigate a stored field that nobody on either side had classified as an input.
  • You reproduced the behaviour once in five attempts. Is that a finding?
    It is a finding about a construction, reported honestly as one success in five against one deployment. With a probabilistic generator a single reproduction is an existence proof, not a rate, and quoting it without the denominator invites both over- and under-reaction. State attempts, deployment and date, and describe the asymmetry rather than the attempt that happened to land.

saying these in an interview costs you the question

  • Reads two allow verdicts as proof nothing happened
  • Treats a tool-call log as evidence of who chose the arguments
  • Assumes the output screen catches whatever the input screen missed
  • Takes a field's provenance as proof somebody intended an instruction
  • Reports one successful reproduction as a reliable rate

context