In a transcription pipeline, how does the recognition step's own error behaviour become part of the attack surface?
answer
- the stage guesses, it does not copy
- output words with no input counterpart
- normalisation rewrites what was spoken
- two runs, two transcripts
- cheap to hide, hard to aim
basics
~20 sA transcript is a recogniser's guess, not a copy. Its mistakes can add text nobody spoke, and its normalisation rewrites what was said, so the span the model reads is shaped by the stage as much as the speaker.
solid answer
~50 sRecognition is a generative guess over an ambiguous signal, so the output is not a faithful copy of anything. Three consequences follow. First, the recogniser can emit words nobody said — confusable sounds, mis-placed word boundaries, filler produced over noise or silence — so directive-shaped text can appear with no matching pattern anywhere in the audio. Second, the stage normalises: casing, punctuation, numbers spelled out or digitised, disfluencies dropped. A construction that depends on exact characters may not survive the rewrite. Third, the same bytes transcribed twice can differ, so the attacker cannot guarantee the string and the investigator cannot guarantee the replay. Net effect: the attacker gets a lever aimed at how sounds are rendered, and pays for it with a loss of control over the exact text. Optical recognition over a scanned page behaves the same way with glyph confusions and reading order.
go deeper
Know that recognition produces a best guess rather than a copy, and that its output can contain words nobody said. That single fact carries most of the reasoning above it.
Explain both directions: which errors and normalisations create the opening, and which of them destroy an attacker's control over the exact string. An answer that only names the upside is half an answer.
Be ready to reason about which constructions cross this boundary and which cannot — in particular why encoding-level tricks die here while intent-level ones survive, and what that implies for triaging a report.
Frame the stage as a trust transition rather than a format conversion, and be able to say plainly that any assurance stated over pre-recognition artefacts does not carry across it.
## The transcript is an inference, not a transcription in the clerical sense It is tempting to picture the recognition stage as a faithful copier that turns sound into the words that were said. It is not. It is a model producing the most likely text for an ambiguous acoustic signal, and the same is true of optical recognition over a scanned page: the output is the most likely character sequence for a set of pixels. Once you hold that picture, the stage stops being a neutral pipe and becomes a component with behaviour an attacker can aim at. ## What the error behaviour hands the attacker **Text that nobody spoke.** Confusable sounds, mis-placed word boundaries, and content produced over noise, silence, or crosstalk can all put words in the transcript that were never said. That matters for this class in a very specific way: any inspection performed *before* recognition — over the stored file — has nothing to match on, because the string that eventually appears never existed as a pattern in the input. The span was not smuggled past a check; it was created after the check. **Deniability.** Because the output is a guess, the presence of a directive-shaped sentence in a transcript is not proof that a person uttered that sentence. In a call-recording assistant, that ambiguity cuts against whoever is investigating: the artefact that carries the span and the artefact that was uploaded disagree, and neither is authoritative about the other. **A second lever.** The speaker is choosing sounds, and the sounds are rendered by a stage with its own regularities — how it segments, how it punctuates, how it handles a spelled-out sequence. Aiming at the rendering rather than only at the wording is a real degree of freedom, and it is the degree of freedom that a text-level view of the pipeline misses entirely. ## What the error behaviour costs the attacker This is where a good answer separates itself, because the same property that helps also hurts, and hurts more than people expect. - **No character-level control.** Normalisation rewrites casing, punctuation, numbers, and often removes disfluencies. Any construction that relies on exact characters, on unusual codepoints, or on invisible characters is simply gone: recognition emits ordinary words from a plain output vocabulary, so the codepoint tricks that work in a text channel do not survive here at all. - **No repeatability.** Two runs over the same stored bytes can produce different transcripts. A construction that landed once may not land again, which means the attacker cannot rely on it and cannot test it cheaply against the live pipeline. - **No preview.** The attacker generally cannot see what the recogniser emitted, so they are firing blind at a stage whose output they never observe. The honest summary is that recognition is a *noisy, lossy, one-way* channel. It is excellent at defeating anything that inspected the input, and terrible at delivering anything precise. ## Why this is not a question about recognition quality An interviewer may probe here, and the boundary is worth stating explicitly. How accurate a recogniser is, how you measure that, which engine or decoding strategy to use, how reading order is recovered from a two-column page — those are questions about building recognition systems, and they have their own answers. The attacker-side question is narrower: given that this stage guesses, and that its guess is the first artefact in the pipeline containing text, what does that do to reasoning about where the span came from and what any pre-recognition inspection could possibly have covered. ## The chain to hold in your head Sound or pixels go in; the stage emits a normalised string that may contain words with no counterpart in the input; that string is the first text-bearing artefact in the pipeline; downstream, it is handled as ordinary input, sitting in the model's context beside the application's own instructions when the recap is drafted. Every link is unremarkable on its own. The class exists in the join between the second and third.
- Do invisible-codepoint tricks carry through a transcription stage?No, and that is a useful thing to say out loud. Zero-width characters, homoglyphs and bidirectional controls are properties of a text encoding. Recognition emits ordinary words from a plain output vocabulary, so nothing at codepoint level survives the crossing. Those constructions belong to channels where the attacker actually writes the bytes.
- Why does the guessing behaviour matter more here than in a channel where the attacker writes the text directly?Because it decides which artefact is authoritative. When an attacker writes text, the carrier and the payload are the same object and any inspection of the carrier is an inspection of the payload. When recognition produces the text, the two are different objects, and the mismatch is not an implementation bug — it is what the stage is for.
saying these in an interview costs you the question
- Treats the transcript as a faithful copy of the audio
- Assumes zero-width or homoglyph tricks survive recognition
- Thinks the attacker controls the exact transcript string
- Says a non-reproducing transcript proves nothing happened
- Confuses this with tuning recogniser accuracy