skip to content

How do you build audio understanding when the signal is non-speech, like machine noise?

level: middleimportance: nice to knowfreq 32%

answer

  1. a transcript-first pipeline returns nothing here
  2. turn sound into a picture of frequencies
  3. open vocabulary versus a fixed taxonomy
  4. you need a threshold, not a paragraph
  5. your microphone matters more than your model

basics

~20 s

Transcription pipelines produce nothing here, because there are no words. Either send the audio to a model that accepts raw sound and can describe or classify events, or train a dedicated classifier on spectrograms when the categories are narrow and the accuracy bar is high.

solid answer

~50 s

Non-speech audio breaks the usual pipeline: a bearing whine or a bird call has no transcript, so anything built on speech-to-text returns empty and the signal is silently lost. Two approaches remain. An audio-capable multimodal model takes the raw sound and can describe or answer open-vocabulary questions about it, which is excellent when the categories are unknown or shifting, and when you want reasoning that combines sound with other context. A purpose-built classifier over spectrogram features, or an audio-text embedding model that supports zero-shot matching against label descriptions, is better when the taxonomy is fixed and small, when you need calibrated scores and thresholds, and when the volume makes per-clip model calls uneconomic. In practice, teams often use the cheap classifier as a continuous detector and reserve the general model for explaining the clips it flags. Whichever you choose, evaluate on your own recordings, since acoustic conditions dominate performance.

go deeper

for a junior

Know that transcription only works on speech, and that classifying a machine noise or an animal call needs a model that consumes the raw audio rather than a transcript.

for a middle

Explain the spectrogram framing, contrast an open-vocabulary audio-capable model with a fixed-taxonomy classifier, and say which fits an always-on detector versus an ad-hoc question.

for a senior

Show the cascade design, a cheap recall-tuned detector feeding a general model for explanation, and insist on evaluation against your own recordings with recall quoted at a fixed false-positive rate.

for a principal

Own the data strategy: how a zero-shot start bootstraps labels for a cheaper supervised model later, what alert volume the responding team will actually sustain, and where sensor placement buys more accuracy than any model change.

## The failure that starts the conversation Audio pipelines are usually built for speech, and the default reflex is to transcribe first and reason over text. Applied to a factory-floor recording or a field microphone, that reflex produces an empty string or, worse, a hallucinated fragment of speech, and the entire signal disappears before any reasoning begins. Recognising that non-speech audio needs a different path is most of the question. ## What the signal actually is Sound is a waveform, and almost every approach starts by turning it into a time-frequency representation, typically a mel-spectrogram, which is effectively an image of how energy is distributed across frequencies over time. A bearing developing a fault shows up as energy at particular harmonics that were not there last month. A bird call is a distinctive shape in that image. Once you see the problem this way, much of the toolkit is the vision toolkit. ## Path one: a general audio-capable model Frontier multimodal models accept audio directly, alongside text and in some cases video, and are billed per second of media rather than per token of transcript. They can describe what they hear in open vocabulary, answer questions about it, and combine it with other context in the same request. This path wins when the categories are not known in advance, when you want a description rather than a label, when the volume is modest, and when the audio needs to be reasoned about together with something else, for example a recording paired with the machine's telemetry or with the video of the same moment. It is weakest for fine-grained discrimination inside a narrow domain, where a general model has seen far less of your specific acoustic world than a targeted model has, and for anything requiring a stable, calibrated score you can threshold. ## Path two: a purpose-built classifier A supervised classifier over spectrogram features, or an audio-text embedding model that lets you match a clip against natural-language label descriptions without training examples, gives you something the general model does not: a number. Fixed taxonomy, calibrated probabilities, a threshold you can tune against a precision and recall target, and a per-clip cost low enough to run continuously on every microphone. This path wins for always-on detection over high volumes, for regulated or safety contexts where a tunable operating point is required, and for narrow domains such as industrial condition monitoring or bioacoustic species identification, where domain-trained models are strong and general models are mediocre. ## The pattern most systems converge on Cascade them. Run the cheap classifier continuously as a detector tuned for recall, and pass only flagged clips to the general model for description, context and triage. That keeps the always-on cost proportional to the cheap model and the reasoning cost proportional to the rare event. It also gives you two independent signals to disagree, which is a useful escalation trigger. ## What actually determines success Acoustic conditions, not model choice. The same bearing recorded through a cheap microphone three metres away in a reverberant hall is a different problem from a contact sensor on the housing. Background machinery, compression artefacts from the recording chain, sample rate and clipping all move accuracy more than the difference between two candidate models. That is why the only meaningful evaluation is on your own recordings, from your own hardware, in your own environment, with a labelled set that includes the near-misses and the confusable sources. Two further practicalities. Class imbalance is severe, since the anomaly you care about is rare by definition, so headline accuracy is meaningless and you should be quoting recall at a fixed false-positive rate. And labels are expensive, because non-speech audio often needs a domain expert to label at all, which frequently makes a zero-shot or description-matching approach attractive at the start, with a supervised model trained later on the data the first system helped you collect. ## Where this connects to video When the sound accompanies footage, native video input can carry both into one context, billed per second of audio on top of the sampled frames. That is often the cheapest way to catch an event that is audible before it is visible, and it is a good argument against stripping the audio track reflexively when the environment is one where sound leads.

  • Why is headline accuracy a poor metric for an industrial anomaly-sound detector?
    Because the anomaly is rare. A detector that always says normal can score above 99 per cent on a realistic recording set while catching nothing. Quote recall at a fixed false-positive rate instead, or precision at the recall you actually need, and state the operating point. That framing also forces the real conversation about how many false alarms the people responding to them will tolerate before ignoring the system entirely.
  • You have almost no labelled examples of the fault sound. Where do you start?
    Start with an approach that needs no training examples: match clips against natural-language descriptions using an audio-text embedding model, or ask a general audio-capable model to describe anomalous clips. Use it deliberately as a data-collection tool, reviewing its flags with a domain expert. Once you have a few hundred confirmed positives, a small supervised classifier trained on your own recordings will usually beat it and cost far less to run.
  • When should the audio track be sent alongside video rather than handled separately?
    When the two signals are jointly informative or when sound leads picture. An impact you hear before anything visible enters the frame, or an alarm that explains why people start moving, is much easier to interpret in one context than by reconciling two pipelines by timestamp. Native video input tokenizes audio per second on top of the frames, so the extra cost is predictable and often cheaper than raising the frame rate.

Treating non-speech audio with a transcription pipeline is like handing an X-ray to an OCR system: it finds no characters and reports an empty page, even though the image is full of information.

saying these in an interview costs you the question

  • Runs speech-to-text over non-speech audio and gets nothing
  • Assumes a general model beats a domain classifier on narrow faults
  • Quotes overall accuracy on a heavily imbalanced detection task
  • Evaluates on public clips instead of the deployment's own recordings
  • Ignores microphone placement and background noise as the dominant factor

context