skip to content

In the Dataflow job UI, what do the data freshness and system latency graphs measure?

level: seniorimportance: should knowfreq 48%

answer

  1. one measures staleness, one measures stuckness
  2. one is now minus the watermark
  3. one is the oldest element still in flight
  4. per-stage views localise the culprit
  5. a 45-degree climb means nothing is completing

basics

~20 s

Data freshness is the gap between now and the job's output watermark — how far behind real time fully-processed data is. System latency is the longest time any element currently in the pipeline has spent being processed or waiting to be processed.

solid answer

~50 s

They answer two different questions. **Data freshness** is wall-clock now minus the output watermark, so it tells you how far behind real time the pipeline's *completed* work is; it is the number to compare against a freshness SLA. **System latency** is the maximum duration an element has currently been processing or awaiting processing inside the pipeline, so it tells you whether something is *stuck* right now. Dataflow charts both per job and per stage, and exports them to Cloud Monitoring (`job/data_watermark_age` and `job/system_lag`). Reading them together is the diagnostic: both climbing points at a slow or blocked stage, which the per-stage view localises. Freshness climbing while system latency stays flat usually means the watermark itself is held back — a quiet or lagging source, or a source partition producing nothing — rather than a slow transform. Pair them with backlog metrics to tell a saturated job from a stalled one.

code

text · 4 lines
text
healthy      freshness ~flat low   system latency ~flat low   backlog ~flat
saturated    freshness rising      system latency rising      backlog rising, at max workers
stuck stage  freshness rising      system latency high on ONE stage
source lag   freshness rising      system latency ~zero       backlog low

go deeper

for a junior

Recall that a Dataflow streaming job exposes both a data freshness and a system latency graph, and that freshness is about how far behind real time the output is.

for a middle

Define each precisely — freshness as now minus the output watermark, system latency as the longest an element has currently been in the pipeline — and know they are charted per stage as well as per job.

for a senior

Show the diagnostic read: both climbing means a slow or blocked stage you localise per stage; freshness climbing alone means the watermark is held back at the source; and backlog tells you whether the job is saturated rather than stuck.

for a principal

Own the alerting contract: page on the freshness SLA as the customer-visible symptom, warn on system latency and persistent backlog at the worker ceiling as leading indicators, and exclude draining jobs so redeploys do not distort the signal.

## Two clocks, two questions Dataflow's Job Metrics tab for a streaming job shows several time series; the two that matter for operating the job are **data freshness** and **system latency**. Candidates who have only read the docs describe them as "both measure lag". They do not, and telling them apart is the whole point of the question. **Data freshness** = current wall-clock time − the output watermark. The watermark is the pipeline's claim about how far along in event time it has completely processed data. So freshness answers: *how stale is the output?* If freshness is four minutes, the job is asserting that everything up to four minutes ago has been fully processed and emitted. This is the metric that maps onto a business SLA — "the dashboard is never more than five minutes behind" is a freshness statement. **System latency** = the current maximum duration that an element has been processing or awaiting processing inside the pipeline. It answers a different question: *is anything stuck right now?* If system latency is 30 seconds, some element has been in the works for 30 seconds. Rising system latency means work is piling up inside a stage, not merely that time has passed. Both are charted for the job overall and **per stage**, and both are exported to Cloud Monitoring, where they appear as `job/data_watermark_age` and `job/system_lag` (with per-stage variants). The per-stage view is where diagnosis actually happens. ## Reading the pair The interesting information is in the combination: - **Both climbing together.** The classic overload or blockage. Elements are sitting in a stage, so the watermark cannot advance and freshness follows system latency upward. Open the per-stage graphs: the first stage whose system latency jumps is your suspect. Typical causes are a slow external call in a `DoFn`, a hot key concentrating work on one worker, a throttled sink, or simply not enough workers for the input rate. - **Freshness climbing, system latency flat and low.** Nothing is stuck; the *watermark* is being held back. Common causes: a source with an idle or lagging partition or subscription, so the watermark estimate cannot advance even though everything received has been processed; a source whose oldest unacknowledged message is old; or a very low-traffic stream where the watermark heuristic waits for evidence that time has moved on. Adding workers here does nothing — this is a source-side or watermark-configuration problem. - **System latency spiking, freshness roughly stable.** A transient straggler that the pipeline is absorbing. Worth watching, not usually worth paging. - **Both flat and low while the backlog grows.** Check the backlog metrics (backlog bytes/elements per stage, and the estimated backlog processing time on Streaming Engine). The job may be keeping up with *watermark* progress while the source accumulates work faster than the job drains it. ## Why backlog belongs in the same conversation Freshness and system latency describe time; backlog describes volume. Dataflow's autoscaler for streaming reasons about backlog and CPU utilisation, so backlog is also the metric most closely tied to *why* the job did or did not add workers. A three-signal read — freshness, system latency, backlog — separates the three states you actually care about: healthy, saturated-but-scaling, and stuck. ## Turning them into alerts A workable alerting policy for a streaming job: - **Page on data freshness** exceeding the SLA for a sustained window (not on a single spike). This is the symptom your consumers feel. - **Warn on system latency** rising above its usual band, as the leading indicator that freshness is about to break. - **Warn on backlog** growth that persists after the job has reached its maximum worker count — that combination means the scaling ceiling, not the pipeline, is the constraint. - Alert on **watermark not advancing at all**, which the freshness metric expresses as a straight upward line at a 45-degree slope: freshness growing exactly as fast as wall-clock time is the signature of a completely stalled watermark. That last pattern is the most useful single visual in the UI. A freshness line that climbs at exactly the rate of real time means *nothing* is completing — the watermark is frozen. A freshness line that climbs slower than real time means the job is falling behind but still making progress. ## A caution about drain Draining a job advances the watermark to infinity, which makes data freshness collapse to zero as everything fires. That is an artefact of shutdown, not a sign of health, so exclude draining jobs from freshness alerting. ## What to say in an interview Define both precisely in one sentence each, then immediately give the discriminating case: freshness high with system latency low means the watermark is held back by the source rather than by a slow transform. That single observation is what separates someone who has operated a streaming job from someone who has read the metrics list.

  • Data freshness on a Dataflow streaming job is climbing steadily but system latency stays near zero. What does that suggest?
    Nothing is stuck inside the pipeline — no element has been waiting long — so the output watermark itself is being held back. Look at the source: an idle or lagging partition or subscription, an old oldest-unacknowledged message, or a very low-traffic stream where the watermark heuristic cannot advance. Adding workers will not help.
  • How would you use these metrics to decide whether a job is saturated or stalled?
    Read them with backlog. Saturated looks like growing backlog with system latency and freshness climbing while the job sits at its worker ceiling. Stalled looks like freshness climbing at the same rate as wall-clock time — the watermark frozen — often with one stage's system latency dominating in the per-stage view.
  • Why should freshness alerts exclude jobs that are draining?
    Draining advances the watermark to infinity so every open window fires, which drives data freshness to zero regardless of the job's real health. Alerting that treats that as a signal will both miss real problems and fire spuriously around planned redeploys, so filter on job state.

saying these in an interview costs you the question

  • Treats data freshness and system latency as the same measurement
  • Thinks data freshness measures how long a single record took
  • Adds workers whenever freshness rises, without checking system latency
  • Ignores per-stage views and only looks at job-level graphs
  • Confuses backlog size with watermark lag

context