skip to content

When should an RNN carry its hidden state across mini-batches on a never-ending accelerometer stream?

level: seniorimportance: should knowfreq 38%

answer

  1. who decides where a sequence ends
  2. batch row must continue its own stream
  3. no shuffling once state is carried
  4. stop the gradient at the handoff
  5. reset when the stream truly breaks

basics

~20 s

Carry it only when each batch row truly continues that row's stream from the previous batch: feed batches in order, stop the gradient at each handoff, and reset the state at real breaks such as the device coming off.

solid answer

~50 s

The default is to start every training segment from a zero state, which is fine when segments are long enough to build their own context. On a wrist accelerometer stream that never ends, cutting it into short segments throws away context at every cut, so you carry the final state of one batch into the next. That is valid only if batch row i always continues the stream that row i held before, which means slicing the stream into as many contiguous sub-streams as the batch has rows, feeding them in order, and giving up shuffling. Detach the carried state so the graph does not grow without bound, and set an explicit reset policy: zero it when the watch comes off, when samples are missing, or when you switch subjects. At serving, keep one state per device, keyed by device id.

go deeper

for a junior

Know the default first: each training sequence normally starts from a zero state, and carrying state between batches is a deliberate exception, not standard behaviour.

for a middle

Be able to state the contract — batch row i must continue the same sub-stream, batches must be fed in order, and shuffling is off — and say why the state is detached at each handoff.

for a senior

Show production judgment: define the reset conditions for real discontinuities, keep per-device state at serving, and name the train/serve skew that a mismatched reset policy creates.

for a principal

Own the tradeoff between carried state and overlapping warm-up segments, and be ready to argue when the operational cost of stateful training and a per-stream state store is worth the extra context.

## The default, and why a stream breaks it Normally each training example is a self-contained sequence: you start from `h_0` (zeros), run the recurrence over the example, compute a loss, and throw the state away. Examples can be shuffled freely because none of them depends on another. A continuous sensor stream has no example boundaries. A wrist accelerometer for fall detection samples all day; a conveyor's vibration telemetry runs until the line stops. You still have to train on finite chunks, so you cut the stream into segments — and every cut throws away the context that preceded it. If a segment is 200 steps and the useful context is 2,000, most segments start blind. **Carrying state** (often called running the layer statefully) is the fix: the final hidden state of segment k becomes the initial state of segment k+1, so context flows across the cut even though the loss is computed chunk by chunk. ## The contract you must honour Carrying state is only meaningful if the next batch really is the continuation of the last one, *row by row*. With a batch of B rows, the setup is: 1. Slice the long stream into **B contiguous sub-streams** (or use B different devices/sessions, one per row). 2. Cut each sub-stream into consecutive segments of equal length. 3. Build batch n from the n-th segment of every sub-stream, and feed batches **in order**. 4. After each batch, keep the final state and use it as the next batch's initial state — **for the same row index**. Break any of these and you are initializing a sequence with a state from unrelated data. The most common way to break it is to shuffle: shuffling is the standard defence against correlated gradients, and carrying state forbids it. That cost is real — consecutive batches are highly correlated, so gradient noise is less independent than usual, and you often compensate with a smaller learning rate or more aggressive averaging. ## Stop the gradient at the handoff If you carry the state as a live part of the computation graph, the graph from batch 1 is still attached at batch 500, and memory grows until the job dies. You carry the state's **values** and cut its gradient path at the boundary, so each batch's backward pass stops at the state it was handed. The state then propagates information forward indefinitely while gradients travel only within a batch. ## The reset policy is the design work "When do I zero the state?" has no default answer on a stream, and getting it wrong is a silent data bug. Reset when the stream is genuinely discontinuous: - the device was removed, powered off, or rebooted; - there is a gap between samples larger than the sampling interval allows; - you move to a different subject, device, or physical asset; - a sub-stream reaches its end and a new one is assigned to that row. Do **not** reset merely because an epoch ended, if the stream itself did not end — and conversely, do not carry a state across a six-hour gap just because the two chunks happen to be adjacent in your file. A stale state is often worse than zeros: it asserts a context that no longer holds. ## Serving looks the same, with a state store Online scoring has the same structure. The service keeps the latest hidden state **per stream**, keyed by device id, updates it as each window of samples arrives, and applies the same reset rules. This is a genuine piece of state infrastructure: it must survive restarts (or be rebuilt), it must not be shared across devices, and it must expire — a state loaded after a long silence should be discarded rather than trusted. Train/serve skew here is nasty and invisible: if training resets at chunk boundaries and serving never does, the model sees state distributions it never trained on. ## The cheaper alternative Often you can avoid carrying state entirely by making segments **overlap**. Take segments with a warm-up prefix — the first K steps rebuild context from real data and their outputs are excluded from the loss and from evaluation. You pay redundant compute on the prefix, but you get back shuffling, restartability, and trivial evaluation. Reach for carried state when the required context is far longer than any segment you can afford to materialize; reach for warm-up prefixes when the context is merely longer than the labelled part. ## How to answer Say when it is needed (context outlives any affordable segment on an unbounded stream), state the row-continuity contract and the loss of shuffling, mention detaching the state, and then spend your last sentences on the reset policy and the per-device state at serving — that is the part that separates someone who has run this in production from someone who has read about it.

  • What goes wrong if you shuffle segments while still carrying the state?
    Each segment is then initialized with the state of an unrelated stretch of data, so the model conditions on context that does not belong to it. The damage is silent — training still runs and the loss still falls — but the layer learns to work from misleading state, and serving, where the state really is continuous, sees a different distribution.
  • How do you get long-context conditioning without carrying state at all?
    Give each segment a warm-up prefix: extend it backwards by K steps, run the recurrence over the whole thing, and exclude the prefix's outputs from the loss and the metrics. Context is rebuilt from real data instead of carried, so you keep shuffling and restartability at the cost of recomputing the prefix.
  • What does the reset policy need to cover on a conveyor's vibration stream?
    Any true discontinuity: the line stopping, a sensor swap or reboot, a gap in samples beyond the expected interval, and a switch to a different machine. Epoch boundaries are not discontinuities and should not trigger a reset. A state carried across a long silence asserts a context that no longer holds and is usually worse than zeros.
  • Why must the carried state be detached from the computation graph?
    Otherwise the graph from the very first batch stays attached to every later one, and memory grows without bound until the job fails. Carrying the values while cutting the gradient path lets information flow forward across batches while each backward pass stays confined to its own batch.

saying these in an interview costs you the question

  • Carries the state across shuffled batches
  • Never detaches, then blames memory growth on model size
  • Has no reset policy for gaps or device changes
  • Keeps one global hidden state for all devices at serving
  • Assumes a carried state is always better than zeros

context