skip to content

A node's processor use climbs while its byte and request rates stay flat: what might it be doing to arriving batches?

level: seniorimportance: should knowfreq 45%

answer

  1. processor time up, bytes and requests flat
  2. pass the framed bytes through, or open them
  3. cost scales with records, not bytes
  4. a step change means someone changed something
  5. realign the writer before buying capacity

basics

~20 s

It may have stopped passing batches through untouched and started decoding and re-encoding them. Anything that forces a node to look inside a compressed batch — a conversion, a stamp, a per-record check — turns a free transfer into processor work that scales with records, not bytes.

solid answer

~50 s

The tell is the mismatch: processor time rose but the two things that usually explain it, bytes and requests, did not. The classic cause is that the node is no longer able to accept the batch as the writer framed it and must **decode and re-encode** it. That happens when a client sends a framing the node cannot store as-is, when the node must stamp or validate something inside every record, or when policy requires the stored form to differ from the sent form. In the good case a node accepts a compressed batch and writes and serves those bytes untouched, spending almost nothing on the contents; in the bad case it pays per record, on every batch, forever, and nobody budgeted for it. The diagnosis is to correlate the step change with a client-side change, and the fix is normally to realign what the writer sends rather than to buy processor capacity.

go deeper

for a junior

Recall that a node can sometimes keep a compressed batch exactly as it arrived and spend nothing on its contents, and that anything forcing it to open the batch costs processor time instead.

for a middle

Explain why the cost scales with the number of records rather than the bytes, and name the situations that force a node to open a batch: a framing it cannot keep as-is, a per-record stamp, a per-record check.

for a senior

Show the diagnosis: a step change with bytes and requests flat, correlated with a client-side change, attributed to specific callers, and fixed by realigning what the writer sends rather than by adding capacity.

for a principal

Decide whether this cost is accidental or structural. If a policy requires records to be held in a different form than they arrive in, that is a standing per-record charge on every future writer, and it belongs in the capacity model and the onboarding rules.

## Two very different things a node can do with a batch When a compressed batch arrives, a node has two broad options: 1. **Pass it through.** Verify what it must at the envelope level, then keep the batch as the writer framed it — hold those bytes, and later hand the same bytes to readers. The contents are never decompressed; the node's cost is dominated by moving bytes, and it is close to free per record. 2. **Decode and re-encode.** Decompress the batch, do something per record, then compress it again — possibly with different settings — before it is held. Now the node's cost scales with the **number of records** and with the compressor's cost, on every batch, forever. The entire economics of this leaf changes between those two cases. The writer's compression choice is a gift to the cluster in case one and a bill to the cluster in case two. ## What forces a node to look inside The reasons vary by platform, but they fall into recognisable families: - **Framing conversion.** A client speaks an older or different framing than the form the node keeps, so every batch is translated. This is the one that appears overnight after a client library is upgraded or a new team onboards with different tooling. - **Per-record stamping.** The node must attach or normalise something on each record before it is held, which cannot be done without opening the group. - **Per-record validation or routing.** A rule that needs the individual records — anything the node must check or decide record by record — defeats pass-through by definition. - **A policy difference in the stored form.** If the cluster requires records to be held in a particular way that differs from how they arrived, the re-encode is structural and permanent. ## Why the symptom looks the way it does This cost has an unusual signature, and recognising it is the point of the question: | Observation | What it rules in or out | |---|---| | Processor time up, byte rate flat | Not a volume problem; something per record or per batch changed | | Processor time up, request rate flat | Not a request-overhead problem either | | Step change rather than a ramp | A client-side or configuration change, not organic growth | | Concentrated on some nodes or streams | Points at a specific caller or a specific stream's traffic | Because the extra work is per record, a caller sending many small records inside each batch can multiply the cost enormously without moving the byte rate at all. This is also why the effect can appear when a writer *improves* its batching: more records per batch, same bytes, far more per-record work for a node that has to open them. ## Where designs differ An answer here has to be careful not to describe one platform as the model: - Designs that **store what the writer framed** have the sharpest version of this: pass-through is nearly free, and losing pass-through is a cliff rather than a slope. - Designs that **unpack every batch on arrival by nature**, tracking each message individually, never had pass-through to lose; there the per-record cost is the baseline, and it is the record count, not the batch, that predicts it. - Designs where the broker **never examines payloads at all** cannot exhibit this failure for payload reasons, though envelope-level work still scales with record count. So the general statement is: the more a node is obliged to understand about individual records, the less the writer's framing decisions are free, and the more the cluster's cost tracks record count rather than byte volume. ## Working the incident 1. **Correlate the step with a change.** Client library versions, a new writer onboarded, a compression choice changed, a cluster setting altered in the same window. This is almost always a change, not a drift. 2. **Attribute it.** Find whether the extra work is concentrated on particular streams or particular callers; per-record cost follows the caller sending the most records, which may not be the caller sending the most bytes. 3. **Confirm the mechanism.** Check whether the affected traffic is arriving in a form the node can keep as-is, rather than assuming. 4. **Fix on the writer's side if you can.** Aligning what the writer sends with what the node can hold untouched removes the cost entirely; adding processor capacity only pays for it. 5. **If the re-encode is required by policy, price it.** It is then a permanent per-record cost that belongs in capacity planning, not a defect to be chased again in six months. The habit worth demonstrating is refusing the reflex answer of *add capacity*. The cluster is doing work it did not use to do; find out why, and the cheapest fix is usually a conversation with whoever changed what they send.

  • Why can this cost appear right after a writer improves its batching?
    Because the extra work is per record. If the writer now packs many more records into each batch at the same byte rate, a node that has to open every batch does far more per-record work than before, while the two metrics an operator usually watches — bytes and requests — stay flat or even improve. The improvement moved cost rather than removing it.
  • How would you tell this apart from a node that is simply short of capacity?
    A capacity problem arrives with growth: bytes, requests or connections rise and processor time rises with them. This arrives as a step with the inputs unchanged, and it is usually concentrated on the streams or callers whose traffic changed form. If the ratio of processor time to bytes accepted has jumped, the node is doing something new per record, not more of the same.
  • Is buying more processor capacity ever the right answer here?
    Yes, when the re-encode is required rather than accidental — a policy about the form records are held in, or a check that genuinely must run per record. Then it is a standing cost proportional to record count and belongs in the capacity model. What is wrong is reaching for capacity before establishing which of the two cases you are in.

saying these in an interview costs you the question

  • Assumes a node always decompresses everything it receives.
  • Reaches straight for more capacity without finding what changed.
  • Expects processor cost to track bytes rather than record count.
  • Believes the writer's compression choice can never cost the cluster.
  • Blames organic growth despite a clear step change in the signal.