skip to content

A callback fired once from a colleague's client and never in your harness — is that a finding?

level: seniorimportance: should knowfreq 38%

answer

  1. one claim or two?
  2. the harness has no renderer in it
  3. placement is a rate, dereference is a yes
  4. a fixture answer needs no model
  5. one line, one client, one version, once

basics

~20 s

Yes, but not as one claim. Split it: whether the outsider's field text reaches the answer, measurable server-side over many runs, and whether a given surface resolves the reference while drawing, testable per client with a fixture and no model at all.

solid answer

~50 s

Yes, but not the finding as filed. Split it in two. Placement — does the outsider's field text reach the answer carrying a reference? — is measurable server-side over many runs and gives an honest rate. Dereference — does surface X issue the request while drawing? — needs no model at all: render a fixture answer in each client and watch. The single log line proves that one client, at one version, at one moment, resolved a model-authored reference. It does not prove the client population, the emission rate, or that anyone read the answer. Your harness's silence is evidence about the harness's display path, not about the construction. Many surfaces also cache a resolved preview, so a second attempt legitimately produces no request. File it with the observed surface and version, the measured placement rate, and an explicit list of what was not tested.

code

json · 12 lines
json
{
  "callback_host_log": [
    { "ts": "T+2d 09:41:22Z",
      "path": "/[38 chars not reproduced here]",
      "count": 1,
      "client": "desktop chat client displaying a shared answer" }
  ],
  "harness_replay": { "renders": 412, "requests_observed": 0,
                      "surface": "plain-text console; resolves no references" },
  "placement_check": { "answers_sampled": 200,
                       "answers_quoting_the_notes_field": 173 }
}

go deeper

for a junior

Recall that a test harness which prints answers as plain text cannot observe a fetch performed by a rich client, so silence there is not a negative result.

for a middle

Explain how to separate whether the untrusted text reaches the answer from whether a given surface resolves a reference while drawing, and why the second needs no model.

for a senior

Demonstrate the triage judgment: measure placement as a rate, test dereference per surface with a fixture, and write severity as conditional on a client population you may not be able to enumerate.

for a principal

Own the reporting standard — what a single observation may be claimed to establish, and how a team's credibility erodes when severity outruns evidence in either direction.

## The situation A data-analysis assistant quotes a free-text notes field from a database row — a field an outsider can set through a public form. Somebody's answer, viewed in a desktop chat client, produced a single request to an unfamiliar host. You are asked to reproduce it. Four hundred renders in the test harness later, nothing. Now you have to write up what you actually know. ## The mistake is filing one claim The report as filed is a compound claim: *outsider text causes a fetch from user clients*. That is two independent claims stapled together, with different evidence requirements and different failure modes. **Claim A — placement.** Does the outsider's field text reach a rendered answer, intact enough to carry a reference? This is a property of the pipeline: retrieval, whether the row is selected, and whether the model quotes rather than paraphrases. It is probabilistic. The honest measurement is a rate over many runs of the same question with the same row, observed *server-side on the emitted answer text* — no display surface involved. **Claim B — dereference.** Does a given surface issue a network request while drawing an answer that contains such a reference? This is a property of a client and its version, and it has nothing to do with the model. Render a fixture answer — text you wrote yourself — in each surface and watch for the request. Deterministic, cheap, repeatable, and it isolates the half that the harness cannot speak to. Splitting the claim is the whole senior move. It converts an unreproducible incident into two things you can each state with a number or a yes. ## What the one log line proves, and what it does not It proves that one client, at one version, at one moment, resolved a model-authored reference while displaying an answer, and therefore that the field text did reach that answer at least once. It does not prove: that other surfaces do the same; that the model emits the reference reliably; that any human read the answer; or that the pipeline is the only way that text could have got there. Getting these backwards is how a finding acquires a severity it cannot support — *all users are affected* is not a conclusion one request can carry. ## What the harness's silence proves Evidence about the harness. A test harness that renders answers as plain text in a console resolves nothing while drawing, so it can never observe claim B regardless of how often claim A succeeds. Non-reproduction there is not evidence about the construction; it is a statement that your observation instrument lacks the component under test. Other honest reasons a second attempt sees nothing: - **Caching.** Many surfaces resolve a given reference once and reuse the result, so a repeat produces no request even where the first one did. Repeating with the *same* reference measures the cache, not the client. - **Version drift.** The colleague's client may have updated between the event and your attempt, in either direction. - **Per-user or per-workspace settings** that change whether previews are expanded at all. - **Placement variance.** The model may quote the field in some fraction of runs; a handful of attempts is not a measurement of that fraction. ## How to file it A write-up that survives review states, separately: 1. The observed event: which surface, which version if known, when, and that exactly one request was seen. 2. The measured placement rate, with the sample size, taken server-side. 3. The per-surface dereference results — which clients were tested with a fixture, which resolved, and, explicitly, which were **not** tested. 4. Severity expressed as a function of the client population, with the population stated as unknown if it is unknown. That last point is where most reports go wrong in both directions: either inflating to *everyone* from one line, or dismissing as *not reproducible* when the reproduction attempt used an instrument that could not observe the effect. ## Is it a bug or a design limit? One confirmed dereference establishes that the class exists in this deployment: untrusted text reaches an answer, and at least one display surface treats part of that answer as something to fetch. That is a finding. What is genuinely unknown is its reach, and pretending otherwise is what makes red-team reports get discounted the next time. Say the reach is unknown, say what it would take to know, and let the owner decide what that is worth.

  • How would you test the dereference half without involving the model at all?
    Write the answer text yourself as a fixture, containing a reference to a host you control, and display it in each surface the team can name — the chat client, the mobile app, the emailed digest, the dashboard tile. Watch for the request. That isolates a deterministic property of each client and version, gives a per-surface yes or no, and removes the model's variability from the experiment entirely.
  • Your second attempt in the same client produced no request. What are you actually measuring?
    Most likely the cache. Surfaces that expand previews commonly resolve a given reference once and reuse the result, so repeating with the same reference tests caching rather than client behaviour. Vary the reference between attempts. Other candidates are a client update between attempts, a per-user setting, and simple placement variance — the model may not have quoted the field that time.
  • What severity do you put on it, given one observation?
    Express severity as conditional and say so explicitly: high impact if the surfaces in wide use resolve references, unknown reach until they are tested. State the measured placement rate, the one confirmed surface, and the list of untested surfaces. A report that asserts a population-wide number from a single log line gets discounted, and the next real finding pays for it.

saying these in an interview costs you the question

  • Treats non-reproduction in the harness as disproof
  • Files placement and dereference as one untestable claim
  • Infers population-wide impact from a single request
  • Repeats the same reference and calls the cache a negative result
  • Concludes a human must have read the answer

context