skip to content

A directive found in one call transcript won't reappear on re-transcription — is that a finding?

level: seniorimportance: should knowfreq 47%

answer

  1. two questions, two kinds of evidence
  2. a replay makes a new object
  3. the draft and its recipients decide it
  4. did it happen versus how often
  5. the transcript proves what was handed over

basics

~20 s

Yes, if the exposure is decidable from what was produced. A failed replay produces a new transcript; it does not overwrite the one that actually reached the model. Separate whether it happened from how often it recurs.

solid answer

~50 s

Split the question in two, because they have different evidence. `Did it happen` is a records question: the transcript that was actually used, the draft it produced, and who that draft went to either show the exposure or they do not, and a fresh run has no bearing on it, producing a different object over the same bytes. `Will it recur` is a rate question, and a single non-reproducing case tells you almost nothing about a rate; you would need many samples to say anything. Then be precise about what each artefact proves. The stored audio proves what was uploaded, not what text existed. The transcript proves what the drafting stage was handed, not that a person said those words. The draft proves what the model produced from that transcript. If the transcript that was used was not kept, you are reasoning from the output alone, and you say so.

code

json · 10 lines
json
{
  "upload_id": "rec-4471",
  "audio_sha256": "9f2c...",
  "trail": [
    {"stage": "upload_scan", "inspected": "audio/m4a, 41.2 MB", "text_seen": null, "verdict": "clean"},
    {"stage": "transcribe", "run": "t-1 (consumed by the draft)", "chars": 18422, "directive_span": "present - 34 chars elided"},
    {"stage": "transcribe", "run": "t-2 (replay, same bytes)", "chars": 18391, "directive_span": "absent"},
    {"stage": "draft_recap", "input_run": "t-1", "included_prior_call_material": true, "recipients": "..."}
  ]
}

go deeper

for a junior

Know that recognition can give different text for the same file, so a failed replay is not proof that nothing happened. Look at what the pipeline actually produced.

for a middle

Be able to list the artefacts in the trail and say what each one supports: the file, the transcript that was used, a replayed transcript, and the drafted output.

for a senior

Show the triage discipline of splitting exposure from recurrence, and say out loud what you could not replay. This is the level where an interviewer is scoring whether you would have closed the ticket.

for a principal

Be ready to set the standard your organisation uses for accepting findings from non-deterministic pipelines, so that intermittent results are neither dismissed nor over-claimed.

## Why this comes up at all In a call-recording assistant, non-determinism sits in an unusual place. People expect the *model* to vary between runs. Here the variance is upstream: the recognition stage guesses over an ambiguous signal, so the same stored bytes transcribed twice can yield different text. The consequence is a triage situation that catches people out — a report arrives saying a spoken directive steered a drafted recap, the engineer re-runs the pipeline over the same recording, and nothing happens. ## The two questions inside one **Did it happen?** This is decidable and it is about records, not replays. The transcript the drafting stage actually consumed, the draft that came out, and its recipients are the evidence. In this setting the interesting failure is usually the audience: content from an adjacent recording or from an earlier call in the same account thread showing up in a recap that went to the wrong side of the table. That either did or did not go out. A fresh recognition run cannot un-send it. **Will it recur?** This is a rate, and one observation is a terrible estimator of a rate. Worse, the variance has two independent sources stacked — recognition and generation — so a single failed replay is weak evidence in both directions. If you need a rate, you need many trials and you should say what your denominator is, rather than reporting `reproduces intermittently` and leaving the reader to guess. Conflating the two is the actual error. Reports get closed as `not reproducible` when the question that mattered was already answered by the output records. ## What each artefact proves, exactly - **The stored audio** proves what was uploaded. It does not prove what text existed downstream, because at that point no text existed. - **A re-run transcript** proves what recognition emitted *this* time. It is a new object, not a measurement of the old one, and its disagreement with the first transcript refutes nothing. - **The used transcript** proves what the drafting stage was handed. It does not prove that any person uttered those words — recognition can emit text with no counterpart in the input. - **The draft and its recipients** prove what was produced and where it went. This is usually the strongest artefact you have, and the one that decides the exposure question. - **A pipeline run identifier** proves which components and versions were in the path, which matters because a component swap can change the behaviour without anyone touching the application. Notice how none of these prove intent, and be careful about asserting it. The direction of every claim has to be right: a directive in a transcript proves the string reached the drafting stage, not that somebody planned to put it there. ## Writing the finding State the class, not the string. The class is: the pipeline's inspection point sits before the artefact that carries text, so anything that reaches the model through recognition was never subject to it. Report the observed instance as one instance, name explicitly what you could and could not replay, and give the exposure separately from the reproduction rate. A report that leads with `reproduced 1 of 5 times` invites the wrong conclusion; a report that leads with `a recap containing another participant's material was drafted and sent, and here is the transcript it was drafted from` does not. If the transcript in question was not retained, say that plainly and reason from the draft. It is a weaker finding and it is still a finding. What is not acceptable is presenting a replayed run as if it were the original, or presenting a failed replay as a refutation. ## The judgment an interviewer is listening for They want to hear that you know a probabilistic pipeline does not yield binary reproduction, that you can name which artefact answers which question, and that you will not let `it did not reproduce` close a case where the output records already settled the exposure. They also want to hear the reverse discipline: one success against one deployment on one day is not a reliable finding either, and you should not inflate it into one.

  • The replay produced a different transcript. What has that measured?
    Only what recognition emitted on that occasion. It is a fresh object built from the same bytes, not a re-observation of the first transcript. It cannot contradict a stored transcript, and it is one sample toward a rate you have not established. Treat it as new data, never as a refutation.
  • Would you file this differently if the transcript that was used had not been retained?
    Yes, and I would say so in the report. Without it, the strongest artefact is the draft and its recipients, which still establishes that material from elsewhere reached an outbound message. The causal step is then inferred rather than shown, so the finding is weaker and should be written as weaker.
  • How would you answer if asked for a reproduction rate anyway?
    By giving a denominator or declining. A rate over a probabilistic pipeline needs many trials, and the variance has two stacked sources here, recognition and generation. `One in five` from five attempts is noise. I would state the sample size alongside any figure so nobody reads it as a stable property.

saying these in an interview costs you the question

  • Closes the report as not reproducible without checking outputs
  • Treats a replay transcript as evidence about the original run
  • Reports a rate from a handful of attempts
  • Claims the transcript proves someone spoke those words
  • Assumes the stored audio settles what text existed

context