skip to content

How do sequence-to-label, per-step tagging and sequence-to-sequence framings of a recurrent model differ?

level: juniorimportance: must knowfreq 70%

answer

  1. Ask how many outputs the head emits
  2. Alignment between input and output steps
  3. Final state versus every step
  4. Output length can be a model decision

basics

~20 s

They differ in output shape. Sequence-to-label emits one prediction for the whole input, per-step tagging emits one prediction aligned to each input step, and sequence-to-sequence emits a new sequence whose length need not match the input.

solid answer

~50 s

The recurrent cell is the same in all three; what changes is which hidden states you read out and where the loss attaches. Sequence-to-label reads out once per example, usually from the final hidden state or from a mean or max pool over all step outputs, and carries one loss term per example - classifying the intent of a customer sentence, say. Per-step tagging emits an output at every input step, one-to-one and in order, with the loss averaged over steps - BIO named-entity tagging over a news sentence, where each token gets `B-PER`, `I-PER` or `O`. Sequence-to-sequence emits a sequence whose length is not tied to the input length and is decided during generation, with a stop symbol ending it - translation or summarisation. Pick the framing from where the labels actually live in your data and when the decision is consumed, not from the architecture you already have.

go deeper

for a junior

Be ready to name the three framings and give one example of each: a single label for a whole sentence, one label per token, and an output sequence of a different length from the input.

for a middle

Explain where the read-out and the loss attach in each framing, and why tagging keeps input and output aligned one-to-one while generation decides its own length and must learn to stop.

for a senior

Show that you pick the framing from where the labels actually live and when the decision is consumed downstream, and that you know generation costs more to serve and adds failure modes the aligned framings cannot have.

for a principal

Own the call when a product wants several kinds of output from the same stream: decide which of them justify the annotation cost of per-step labels, and which are cheap sequence-level decisions dressed up as generation.

## The framing is about output shape, not about the cell A recurrent network reads a sequence one step at a time, updating a hidden state from the previous state and the current input, with the same weights reused at every step. That machinery is identical in all three framings below. What differs is **how many outputs you produce, which hidden states you read them from, and where the loss attaches**. Interviewers open here because a candidate who thinks a recurrent model always produces one label at the end can only build sentence classifiers. ### Sequence-to-label (many-to-one) Input: a sequence of T steps. Output: exactly one prediction for the whole sequence. The classification head is normally attached to the hidden state after the final step, on the reasoning that it has seen everything. A common and often stronger alternative is to compute an output at every step and **pool** them - mean or max over time - before the head. Pooling matters because a single final state is recency-biased: in a long input the early steps have been overwritten many times, and the gradient reaching them is weakest. There is exactly one loss term per example. Typical tasks: deciding whether a support chat needs escalation, classifying the intent of a spoken request, scoring a whole user session as risky or not. ### Per-step tagging (many-to-many, aligned) One output for every input step, in order, one-to-one. The alignment is fixed by the framing - output i belongs to input i - and the model never decides how long its output is. The loss is averaged or summed over the T steps, so every step contributes gradient and every step is supervised. The canonical example is BIO tagging for named entities over news sentences: each token is labelled `B-PER` for the beginning of a person name, `I-PER` for a continuation of one, `O` for outside any entity, with parallel tags for organisations and locations. Part-of-speech tagging, per-frame speech activity detection and per-event anomaly flags in a session all share this shape. A useful contrast: intent classification of exactly the same news sentence is sequence-to-label. Same sentence, same tokens, same recurrent cell - a completely different head, loss and label file. That is the whole point of the distinction. ### Sequence-to-sequence (many-to-many, unaligned) The output is a sequence whose length is not determined by the input length. One recurrence reads the input; a second stage produces the output step by step, each emitted step conditioned on the steps already emitted, until a designated stop symbol is produced. There is no correspondence between output position j and input position j, and the model must learn where to stop. Typical tasks: translation, summarisation, transliteration, generating a free-text explanation of an event. This is the most expensive framing to train and to evaluate. Length is now a model decision, so a whole class of failure - truncated or runaway outputs - exists that the other two framings simply cannot exhibit. ### How to choose Two questions settle it almost always. 1. **Where do the labels live?** If your annotators marked spans inside each example, you have per-step labels and a tagging task. If each example carries one label in one column, you have sequence-to-label. If the target is free-form text of a different length from the input, you need generation. 2. **When is the decision consumed?** A per-step framing gives you a decision you can act on as the sequence unfolds. A sequence-to-label framing gives you one decision, and only once the read-out point is reached. A third consideration is cost. Aligned framings are cheap: one forward pass yields every output. Generation is sequential at inference - you cannot produce step j+1 until you have produced step j - so its serving cost scales with output length in a way tagging does not. ### Common mistakes - **Assuming the output length always equals the input length.** True for tagging, false for the other two. - **Solving a tagging task by classifying each token independently.** That discards exactly the sequential context the recurrence exists to provide; a person name is `B-PER` or `I-PER` depending on what came before it. - **Reaching for generation on an aligned task.** If the output is one label per input step, generating it as free text buys length errors and slower serving for nothing. - **Reading a long sequence-to-label input from the final state alone** when pooling over steps is a one-line change with better behaviour. - **Thinking the framing changes the cell.** It changes the head, the loss and the label file.

  • In a sequence-to-label framing, why is the final hidden state often a weak summary of a long input?
    Because the state is overwritten at every step, so it is recency-biased: information from early steps has survived many updates and the gradient reaching those steps is the weakest in the sequence. On long inputs, pooling the per-step outputs - mean or max over time - usually gives a more stable summary, since every step contributes to the read-out directly rather than only through the chain of state updates.
  • In per-step tagging, why average the loss over time steps rather than take it only at the end?
    Because every step carries its own supervision. Averaging gives each step an equal share of the gradient and delivers a training signal directly to every position, rather than routing all of it through the final state. A loss taken only at the end would waste the labels you already paid annotators for, and would leave early steps learning only through a long chain of state updates.
  • Both tagging and sequence-to-sequence produce many outputs - what really separates them?
    Alignment and who decides the length. Tagging is one-to-one and monotone: output i corresponds to input i and there are exactly as many outputs as inputs. Sequence-to-sequence has no positional correspondence, the output length is decided during generation and ended by a stop symbol, and each emitted step is conditioned on the steps already emitted. That makes generation strictly harder to train, evaluate and serve.

saying these in an interview costs you the question

  • Says a recurrent model always emits one label at the end
  • Assumes the output length always equals the input length
  • Treats named-entity tagging as sentence-level classification
  • Thinks the framing changes the recurrent cell itself
  • Classifies each token independently and calls it tagging

context