skip to content

An agent must react within a second when a key changes in Consul's KV store, without polling the API in a loop. Explain how a blocking query on `GET /v1/kv/config/app` works, what the `X-Consul-Index` response header is for, and the two index values a client has to handle defensively.

level: middleimportance: should knowfreq 50%

answer

  1. it is a long poll, not a subscription
  2. the server needs to know where you left off
  3. a response header carries that position
  4. the same value back means nothing happened
  5. a smaller value means start over

basics

~20 s

A blocking query is a long poll. You send the last X-Consul-Index value back as ?index=N with ?wait, and Consul holds the connection open until that path's index advances or the wait expires, then you repeat with the new index.

solid answer

~50 s

Every read endpoint that supports blocking returns an `X-Consul-Index` header — the Raft index of the data you just read. Send it back as `?index=<N>&wait=30s` and the server parks the request until the underlying data changes past that index, or the wait elapses. If it returns with the *same* index, that was a timeout and nothing changed; if it returns a *larger* index, re-read the payload and store the new value. Two defensive rules matter in practice. If the returned index is **smaller** than the one you sent — which happens after a restore or a state reset — reset to `0` and start over, or you will block forever waiting on an index that will never arrive. And treat any index below `1` as `1`, because Consul does not issue `0` as a valid data index. Also diff the payload: on a prefix query the index is the maximum across the prefix, so an unrelated sibling key wakes you up.

code

bash · 24 lines
bash
#!/usr/bin/env bash
# Long-poll a Consul KV prefix and act only on real changes.
set -euo pipefail
ADDR=http://127.0.0.1:8500
API='/v1/kv/config/app/?recurse'
index=0

while true; do
  body=$(mktemp); headers=$(mktemp)
  curl -s -D "$headers" -o "$body" "${ADDR}${API}&index=${index}&wait=30s" || { sleep 2; continue; }

  new=$(awk 'tolower($1) == "x-consul-index:" { print $2 }' "$headers" | tr -d '\r')
  [ -z "$new" ] && { sleep 2; continue; }
  [ "$new" -lt 1 ] && new=1               # 0 is never a valid data index
  if [ "$new" -lt "$index" ]; then        # state reset: restart from scratch
    index=0
    continue
  fi
  if [ "$new" -gt "$index" ]; then
    echo "changed at index $new"          # re-render / reload here
  fi
  index=$new
  rm -f "$body" "$headers"
done

go deeper

for a junior

Know that Consul supports a long poll rather than requiring timed polling, and that the X-Consul-Index header is the value you send back on the next request to be woken on change.

for a middle

Walk the loop end to end: index in, wait, compare the returned index, act only when it advanced. Name both defensive cases — an index moving backwards, and treating a value below one as one — and say what breaks without them.

for a senior

Show that a wake-up is not a change: diff the payload before acting, keep handlers idempotent, and reason about connection cost when thousands of watchers are open. Know when a stale read is the right trade for read scale.

for a principal

Own the watch topology across the platform — how many long-lived connections the agents carry, whether services watch prefixes or individual keys, and whether change delivery goes through a rendering tool rather than every service implementing its own loop.

## The index is a position in the log, not a timestamp Every write in Consul goes through Raft and lands at a monotonically increasing log index. A KV entry carries the index of its last change as `ModifyIndex`, and a read response carries `X-Consul-Index` — for a single key that is the key's `ModifyIndex`; for a prefix read it is the maximum across everything under the prefix. Because it is a log position and not a clock, two parties can agree on "has anything changed since I last looked" without any shared time source and without the server keeping per-client state. ## The long-poll loop A blocking query is just an ordinary read with two extra parameters: ``` GET /v1/kv/config/app?recurse&index=1904&wait=30s ``` The server compares the current index for that path with the one you supplied. If it is already higher, it answers immediately. Otherwise it holds the request open. It returns when the data changes past your index, or when the wait expires — whichever comes first. The default wait is 5 minutes and the maximum is 10 minutes; Consul also adds a small random jitter (up to a sixteenth of the wait) so that a thousand watchers that started together do not all time out and re-arm in the same millisecond. The client loop is therefore: ``` index = 0 loop: resp = GET path?index=index&wait=30s newIndex = resp.header["X-Consul-Index"] if newIndex < 1: newIndex = 1 # 0 is never a valid data index if newIndex < index: index = 0; continue # state was reset; start over if newIndex > index: apply(resp.body) # real change index = newIndex ``` Both guards are load-bearing. Without the reset guard, a client that saw index 90,000 before a snapshot restore rolled the cluster back to 40,000 will send `index=90000` forever and never be woken again — a silent, permanent stall that looks like "the watcher stopped working" with no error anywhere. Without the floor, a `0` sent back as `?index=0` means "return immediately", which turns your long poll into a hot loop against the agent. ## An unchanged index is not a change The most common integration bug is treating *any* return from the blocking query as a change and acting on it — re-rendering a file, bouncing a process, republishing config. Returns are frequent and mostly are timeouts. Compare indexes first, and then compare content: on a `?recurse` watch, a write to *any* key under the prefix advances the maximum index, so a watcher on `config/app/` wakes for a change to a key it does not care about. Idempotent handling — render, compare with what is already on disk, act only on a real diff — is what makes this safe. ## Watches, and what does the polling for you You rarely write the loop by hand. `consul watch -type=key -key=config/app/db_url -- /usr/local/bin/on-change` runs the loop inside the agent and invokes a handler on change; the agent configuration file has an equivalent `watches` stanza supporting types such as `key`, `keyprefix`, `services`, `nodes`, `checks` and `event`. Tools built on the same mechanism — consul-template and envconsul — hold a blocking query per dependency referenced in their templates and re-render when any of them advances. ## Consistency and cost By default a read is served from the leader's committed state, which is why responses also carry `X-Consul-KnownLeader` and `X-Consul-LastContact`. Adding `?stale` lets any server answer from its own replicated copy, which scales reads horizontally at the cost of possibly lagging the leader — `X-Consul-LastContact` tells you how long ago that server last heard from the leader, so you can reject an answer that is too stale. `?consistent` goes the other way, forcing a round trip that confirms leadership before answering: slower, but no window. Cost matters at scale. Each outstanding blocking query holds an open connection and server-side resources for the life of the wait, so ten thousand watchers each watching five individual keys is fifty thousand long-lived connections. Watch a prefix once and fan out locally rather than opening one query per key, and give each watcher a bounded wait so a wedged connection is eventually retried rather than waited on forever.

  • Your watcher stopped firing after the cluster was restored from a snapshot, with no errors logged. What happened?
    The restore rolled the Raft index backwards, so the client is still sending an index higher than anything the cluster will now issue and the query blocks until the wait expires, every time, forever. That is exactly the case the "if the returned index is smaller than the one you sent, reset to 0" rule exists for. Restarting the watcher clears it too, because it starts from index 0.
  • You watch config/app/ with ?recurse and your handler fires when an unrelated key under it changes. Is that a bug in Consul?
    No. On a prefix query the returned index is the maximum across the prefix, so any write beneath it advances the index and wakes every watcher of that prefix. Consul tells you something changed, not what. The handler must diff the fetched content against what it already applied and do nothing when the parts it cares about are identical.
  • When would you add ?stale to a blocking query?
    When read volume is high and a small lag is acceptable — stale reads let any server answer from its replicated state instead of funnelling every watcher through the leader, which is the standard way to scale thousands of watchers. Pair it with the X-Consul-LastContact header so a follower that has fallen far behind can be rejected rather than trusted.
  • How long will Consul hold a blocking query open?
    The wait parameter controls it, defaulting to 5 minutes and capped at 10; the server adds jitter of up to wait/16 so watchers armed at the same moment do not re-arm in lockstep and stampede the agent. Clients usually pick a shorter wait — tens of seconds — so a half-dead connection is detected and retried on a bounded schedule.

saying these in an interview costs you the question

  • Treats every returned response as a change without comparing indexes
  • Forgets to send the index back, turning the watch into a busy loop
  • Never handles an index that moves backwards after a restore
  • Assumes the index is a timestamp or a version counter
  • Opens one blocking query per key across thousands of keys

context