Why doesn't a scan of an uploaded call recording catch an injection that appears in its transcript?
answer
- what the file actually contains at rest
- audio encodes sound, not characters
- the text is created downstream
- scan and model read different objects
- clean verdict, wrong artefact
basics
~20 sThe scan inspects audio bytes, and at that moment the instruction text does not exist. Transcription manufactures it afterwards. A clean verdict on the stored file says nothing about the transcript the model later reads.
solid answer
~50 sAn upload-time scan works over the stored artefact: file type, malware signatures, and pattern matching over whatever text the file already contains. A recording of somebody speaking contains no text at all — the sentence is acoustic, and no byte in the file spells it. The transcription stage is what brings the string into existence, and everything downstream treats that string as ordinary input; in a call-recording assistant it lands in the model's context next to the application's own instructions when the recap is drafted. So this is indirect injection whose carrier genuinely did not carry the payload: the scan and the model examined two different objects, minutes apart. `We scan every upload` is a true statement about the wrong artefact. The same reasoning applies to a scanned page whose characters are produced by optical recognition rather than read out of the file.
go deeper
Be ready to say what an upload scan actually inspects and why a spoken sentence is not present in the file's bytes. The line to land: the text is created later, by recognition.
Walk the pipeline stage by stage and name which stage first holds a string, then explain why that string sits in the model's context as ordinary input beside the application's own instructions.
Expect to be pushed on what a clean scan verdict can honestly be claimed to cover in a deployed pipeline, and on where the first artefact containing the span appears in it.
Own the framing that scanning coverage is stated per artefact, not per upload. A control inventory that records uploads scanned without naming the object inspected will mis-describe this entire class.
## The two artefacts A call-recording assistant records customer calls, transcribes them, and drafts a recap and follow-up messages afterwards. Somewhere in that pipeline there is an upload-time scan: the file is checked for type, run past malware signatures, and often passed through content matching that looks for sensitive or disallowed strings in whatever text the file carries. Now imagine a participant on the call says something shaped like an instruction to the assistant — a sentence whose grammatical mood is imperative and whose subject is the drafting behaviour rather than the deal. The scan will not see it, and no amount of tuning will make it see it, because **the sentence is not in the file**. Audio encodes sound pressure over time. There is no byte sequence in the recording that spells those words. The string comes into existence at the transcription stage, as that stage's *output*. That is the whole point of this class, and the mistake it corrects is a very common one: `we scan uploads for malicious content` sounds like coverage and is not, because coverage is a property of an artefact, not of an event. The upload was scanned. The object that contained the directive was never scanned, because it did not exist yet. | Stage | Artefact it inspects | Does the directive text exist here? | | --- | --- | --- | | Upload scan | the stored file as bytes | No — nothing spells it | | Recognition (transcription or optical) | audio or image in, text out | It is created here, as output | | Recap drafting | the transcript sitting in the model's context | Yes, as ordinary input | ## Why it counts as indirect injection Direct prompt injection arrives in the user's own turn: the person typing to the assistant writes the instruction themselves. Indirect (second-order) injection arrives in content the application retrieves, fetches, or is handed — the application feeds it to the model as data, and the model has no reliable way to tell it apart from the application's own instructions, because the instruction hierarchy is a trained preference rather than an enforced boundary. A recorded call is squarely in the second bucket, with a twist that makes it distinctive: the person who spoke need not be the customer of the assistant at all. They were on a call. Somebody else's tooling recorded them. The application's user — a salesperson — uploaded the recording believing they were uploading a conversation, which is exactly what they did. ## What it costs the attacker This is not a cheap construction, and an interviewer will want to hear that. - **You must be in the conversation.** There is no way to place the span except by producing the sound, which means being on the call, on the line at the right moment, or in the room being recorded. - **You do not control the exact string.** Recognition is a guess. The words that land in the transcript are the recogniser's rendering of the sounds, not a copy of them, so any construction that depends on precise wording is unreliable here. - **It is spoken in front of witnesses.** Everything you say is heard live by the other participants, so the span has to survive a human hearing it, not merely a scan missing it. - **You may not know what happens downstream.** Whether the transcript re-enters a model's context, and what that model can reach when it does, is invisible from the call. ## Where it stops working If the transcript is only ever read by a person, the span is a strange sentence somebody said and nothing more — the class exists only because the transcript is fed back to a model that then does something on the strength of it. It also stops working when the recogniser renders the sounds differently, when normalisation rewrites what was spoken, or when the drafting step has nothing worth reaching. The payoff in this setting is usually the *audience*: a drafted recap or follow-up is addressed to people, and the interesting failure is content from an adjacent recording or an earlier call in the account thread being pulled into a message that goes to the wrong side of the table. ## What to actually say in an interview Name the two artefacts and the moment the text first exists. Say that the scan verdict is true and irrelevant. Then say what the verdict *does* prove — that the stored file matched no signature and held no matching text at that moment — and stop there, because claiming more is the error the question is testing for.
- What does a clean upload-scan verdict actually prove here?That the stored file matched no signature and contained no matching text at the moment it was scanned. It is a statement about one artefact at one time. It does not extend to text that a later stage produces, and it says nothing about what the drafting model was handed.
- If the transcript were only ever shown to a human, would the same spoken span still matter?Far less. A span is only directive when something reads it as instruction and can act. Shown to a person, it is a sentence somebody said on a call. The class exists because the transcript re-enters a model's context with drafting and sending behaviour attached, so the payoff depends on what the pipeline does after recognition.
- Is this direct or indirect prompt injection, and why does the distinction matter?Indirect. It arrives in content the application processes rather than in the user's own turn, and the speaker need not be the application's user at all. That matters because reasoning that starts from `the user typed it` — consent, attribution, per-user rate limits — does not apply to a span that a third party spoke into a recording.
Screening a sealed envelope for anthrax tells you nothing about what the letter will say once somebody reads it aloud into a notepad.
saying these in an interview costs you the question
- Claims a better-tuned upload scanner would have caught it
- Treats the recording and the transcript as one artefact
- Calls it direct injection because a user uploaded the file
- Assumes a clean scan means nothing untrusted reaches the model
- Says audio files contain the spoken words as text