An attention speech recognizer repeats words and its alignment jumps backwards — what is wrong?
answer
- plot the weights before theorising
- audio and transcript share one order
- a jump back re-consumes spent audio
- track coverage of source positions
- force the window to advance
basics
~20 sThe alignment has stopped moving forward: the decoder re-attends to audio it already consumed, so it emits the same words again. Plot the alignment map, mask padded frames, and constrain attention to a window that can only advance.
solid answer
~50 sSpeech alignment is inherently **monotonic** — audio and transcript run in the same order — so a backward jump is a defect, not a quirk. The decoder is re-attending to frames it already transcribed and re-emitting their content; the mirror failures are stalling on one region and racing ahead so the tail is dropped. Plot the output-by-frame weight matrix first and check whether the bright band advances steadily. Then check the cheap causes: padded frames not driven to a large negative score before the softmax, so trailing silence attracts weight, and a near-uniform blur, which means the scores never learned to discriminate. If the alignment is genuinely non-monotonic, constrain it — score only a window whose centre may not move backwards, or track coverage per frame and penalise attending to spent frames. Unconstrained global attention is the wrong prior for a task whose alignment never reverses.
go deeper
Know that alignment weights can be plotted as a heatmap, that a healthy speech alignment forms a band moving steadily forward through the audio, and that repeated output usually means attention revisited earlier frames.
Explain why speech alignment is monotonic while translation is not, and describe the concrete failure shapes: a backward jump, a vertical stall, an early sprint to the end, and a uniform blur.
Demonstrate a diagnosis order — plot first, then verify masking of padded frames, then check whether failures track input length or repeated content — and name coverage tracking and forward-only windows as the fixes with the tradeoffs each carries.
Own the prior-versus-freedom argument: constraining alignment for a provably monotonic task removes a failure mode at no real cost, while the same constraint silently caps quality on a task that must reorder, and be able to say which side a new task falls on.
## Reading the picture first An attention decoder gives you a free diagnostic: the matrix of alignment weights, one row per output step, one column per source position, each row summing to one. Plotted as a heatmap it shows exactly where each emitted token read from. For speech recognition the healthy picture is a **near-diagonal band that only ever moves right**. Audio time and transcript order are the same order — there is no language in which the end of an utterance is spoken before its beginning. That property has a name, *monotonic alignment*, and it is a strong prior the plain attention mechanism does not know about. Global attention can put weight anywhere in the source, at any output step, and nothing in the loss forbids going backwards. (Translation is the contrast case: there the band is near-diagonal but *legitimately* broken in places. Where two languages order an adjective and its noun differently, the map shows a small swapped block, two output positions reading two source positions in the opposite order. That is the model being right, not wrong. In speech there is no analogous excuse.) ## What the described failure is Repeated words plus a backward jump is the decoder re-consuming audio. It emits a phrase, its state after emitting it resembles its state before, the scores peak on the same frames again, and the loop repeats — sometimes for many steps, since nothing tells the model those frames are spent. Related symptoms from the same root: - **Stalling**: the band goes vertical, attention fixed on one region while the decoder babbles. - **Truncation**: the band sprints to the end early and the tail of the utterance is never transcribed. - **Blur**: no band at all, weights near-uniform. This is not a monotonicity problem; it means the score function has not learned to discriminate. Early in training this is normal; late in training it means the alignment never took. ## Diagnosis order **1. Plot before theorising.** Dump the alignment matrix for the failing utterances. The shape of the failure tells you which of the above you have, and they have different fixes. **2. Check masking.** Batched audio is padded to the longest utterance. If padded frames are not driven to a large negative score *before* the softmax, they receive real weight, and since padding often sits at the end, the effect looks like attention being dragged to the tail. Two things to verify: the mask is applied to the scores, not to the weights afterwards; and it is the same mask the encoder used. **3. Check whether it is only long inputs.** Failures that appear past some duration usually mean the score function's discrimination degrades as the number of competing positions grows — with hundreds of frames, a softmax over weakly separated scores spreads mass widely, and a spread distribution is easy to pull backwards. **4. Check whether it is only certain content.** Repeated or acoustically similar segments (a repeated phrase, a long steady tone) genuinely produce similar encoder states, so a purely content-based score cannot tell the second occurrence from the first. This is the case that argues hardest for a positional constraint rather than a better score. ## Fixes, cheapest first - **Mask correctly.** Set padded positions' scores to a large negative value before the softmax. Prefer a large finite negative to a literal infinity, so that an all-masked row cannot produce a not-a-number result. - **Coverage.** Accumulate how much weight each source position has received so far and feed that running total into the score function, with a penalty for attending to positions already well covered. This was introduced for translation's over- and under-translation and transfers directly: a spent frame becomes progressively less attractive. - **Windowing.** Restrict scoring at output step `t` to a window of source positions around a centre, rather than the whole source. Constrain that centre to be non-decreasing in `t` and a backward jump becomes *impossible*, not merely discouraged. The window may advance a fixed amount per step or be predicted by the model. - **Monotonic formulations.** More strongly, formulate attention as a left-to-right process that consumes frames and never revisits them. You lose the ability to reorder, which speech does not need, and gain a hard guarantee plus streaming compatibility. ## The judgment behind the fix The real lesson is about priors. Unconstrained attention is the right default when you do not know the alignment structure, and translation genuinely needs reordering. Speech is the case where you *do* know the structure, and paying for freedom you can prove is unnecessary buys you a failure mode instead of accuracy. Choosing a constrained alignment for a monotonic task is not a hack; it is encoding a known property of the problem. The same reasoning tells you when *not* to constrain: if your task really can reorder, hard monotonicity will silently cap quality, and you would rather pay for coverage tracking and keep the freedom.
- In translation, one off-diagonal block appears where two languages order adjective and noun differently — is that a bug?No, it is the model being correct. Alignment follows meaning, so a local order flip between the languages shows up as a small swapped block while the rest of the map stays near-diagonal. Only global patterns worry you: a band that drifts backwards over many steps, goes vertical, or dissolves into uniform weight across the source.
- Your alignment weights put nonzero mass on padded source positions — what is the fix, and why not just zero those weights after the softmax?Set the padded positions' scores to a large negative value before the softmax, so they receive essentially zero weight and the surviving weights still sum to one. Zeroing after the softmax is worse: the pads have already consumed part of the normaliser, so the remaining weights sum to less than one and the context vector is silently scaled down.
- How would you make a backwards jump structurally impossible rather than merely unlikely?Score only a window of source positions around a centre, and constrain that centre to be non-decreasing across output steps — positions before it are simply not candidates. You trade away the ability to reorder, which speech never needs, for a hard guarantee, and the same restriction is what makes streaming decoding feasible.
- The alignment map is a uniform blur late in training — is that the same problem?No. A blur means the score function is not discriminating between source positions at all, so every context vector is roughly the average encoder state. Monotonicity constraints will not help. Look at whether the scores have collapsed in magnitude, whether the decoder state carries any useful signal, and whether the model is learning the alignment at all.
saying these in an interview costs you the question
- Blames the output layer without ever plotting the alignment
- Claims attention weights cannot be inspected
- Treats any off-diagonal weight as a defect
- Masks padding by zeroing weights after the softmax
- Calls a near-uniform alignment map healthy smoothing