skip to content

Consul's KV store holds the runtime configuration for every service on your platform, and all of the teams' tooling authenticates with a single shared token. What is your plan for limiting who can write where, and for getting data back after someone runs a recursive delete on the wrong prefix?

level: principalimportance: nice to knowfreq 22%

answer

  1. one credential, whole-platform blast radius
  2. the key layout decides the grant
  3. default deny, then widen deliberately
  4. the backup restores more than you wanted
  5. the record of truth can live elsewhere

basics

~20 s

Scope tokens to prefixes under a default-deny ACL system, one token per workload rather than one shared token, and keep the desired state outside Consul. Snapshots restore the whole cluster, not one prefix, so schedule prefix-level exports for recovery.

solid answer

~50 s

Two moves. First, authorization: enable ACLs with `default_policy = "deny"`, then write policies granting `key_prefix` write only on each team's own namespace, so the key layout becomes the permission boundary. Issue a token per workload — ideally minted by an auth method rather than a long-lived static secret — so a compromised or buggy client can damage only its own prefix. Second, recovery, where the trap is that `consul snapshot save`/`restore` operate on the entire Raft state: KV, sessions, ACL tokens and the catalog together. Restoring yesterday's snapshot to undo one bad delete also rolls back every other team's changes and every token issued since, so it is a disaster-recovery tool, not an undo button. For prefix-level recovery, run scheduled `consul kv export` dumps and restore with `consul kv import`. The strongest version is that KV is not the source of truth at all: the desired state lives in a repository a job syncs into Consul, so recovery is a re-run rather than a restore.

code

hcl · 7 lines
hcl
# Agent: ACLs on, deny by default, survive a control-plane blip.
acl {
  enabled                  = true
  default_policy           = "deny"
  down_policy              = "extend-cache"
  enable_token_persistence = true
}

go deeper

for a junior

Know that Consul access is controlled by tokens attached to policies, that policies grant read or write on key prefixes, and that a token should be scoped to what one service actually needs.

for a middle

Explain a default-deny ACL setup and how prefix rules are written, and be clear that KV keeps no history — a delete or an overwrite leaves nothing behind to roll back to.

for a senior

Show the operational consequence: snapshots restore the entire cluster state, so prefix-level recovery needs its own scheduled export, and an untested restore is only a hypothesis. Tie the key layout to the ACL boundary deliberately.

for a principal

Own the position that KV distributes state whose source of truth is elsewhere, so recovery is a pipeline re-run rather than a restore, and set the credential model — a token per workload, minted short-lived — along with the drift and rehearsal checks that keep it honest.

## Why one shared token is the real finding A single token used by every pipeline and every service means every consumer of configuration is also a potential destroyer of it. Consul KV keeps no history — a write replaces, a delete removes, and nothing records the previous value. So the blast radius of one mistyped recursive delete or one leaked token is the entire platform's configuration, with no per-key rollback to fall back on. Both halves of the answer follow from that: shrink what any one credential can touch, and make sure the data exists somewhere other than Consul. ## Prefix layout is the authorization model Consul's ACL rules are prefix-based, so the key namespace *is* the permission boundary and has to be designed as one before anything else: ```hcl # consul agent configuration acl { enabled = true default_policy = "deny" down_policy = "extend-cache" } ``` ```hcl # policy for team-a's deployment pipeline key_prefix "config/team-a/" { policy = "write" } key_prefix "config/shared/" { policy = "read" } key_prefix "" { policy = "deny" } ``` Three decisions are embedded here. `default_policy = "deny"` means anything not granted is refused, which is the only defensible starting point — the alternative allows everything you forgot to think about. `down_policy` decides what happens when a client agent cannot reach the servers to resolve a token: `extend-cache` keeps honouring cached decisions so a control-plane outage does not immediately break the data plane, at the cost of briefly honouring a revoked token. And a `config/<team>/<service>/` shape means a team's grant is one line, which is only true if you chose that shape early; retrofitting a prefix layout onto keys that grew organically is expensive. Runtime consumers should be narrower still: a service that reads its own configuration gets read on its own prefix and nothing else. The write path belongs to the pipeline, not to the application. ## Tokens per workload, and preferably short-lived Static tokens pasted into CI variables are the practical failure mode — they outlive the person who created them, they get copied between pipelines, and rotating them means finding every copy. Roles let you attach the same policy set to many tokens without duplicating rules, and auth methods let a workload exchange an identity it already has (a Kubernetes service-account token, for instance) for a short-lived Consul token, which removes the long-lived secret entirely. Whatever the mechanism, the target property is the same: every token is attributable to one workload, and revoking it affects only that workload. Audit the read side too. Anything in plain KV is readable by every token granted read on that prefix, and base64 in the API response is encoding, not protection. Material that must not be broadly readable does not belong in plain KV at all. ## Recovery, and the snapshot trap `consul snapshot save backup.snap` captures the servers' entire replicated state machine — KV, sessions, prepared queries, ACL tokens and policies, catalog registrations — and `consul snapshot restore` replaces the cluster's state with it. That is the right tool for "we lost server quorum" and the wrong tool for "team-a deleted their prefix at 14:05", because restoring rolls back *everyone*, including every configuration change and every token issued since the snapshot was taken. Reaching for it to fix one prefix converts a one-team incident into a platform incident. So run two backup mechanisms with different granularities: - **Cluster level.** Scheduled `consul snapshot save` to durable off-cluster storage, tested by actually restoring into a scratch cluster — an untested restore is a hypothesis. Automated periodic snapshotting via `consul snapshot agent` is an Enterprise feature; on OSS it is a scheduled job you own. - **Prefix level.** Scheduled `consul kv export <prefix>` producing a JSON dump per namespace, versioned and retained. `consul kv import` puts it back, and because it is scoped you can restore one team without touching another. ## The stronger position: Consul as distribution, not as record The best answer to "how do we get the data back" is that the data was never only there. Keep the desired configuration in a repository, have a pipeline apply it to Consul on merge, and Consul becomes the *distribution* mechanism for state whose source of truth is version-controlled and reviewable. Recovery from a bad delete is then a pipeline re-run, and you also get change review, history, attribution and diffs — none of which the KV store provides. Only the genuinely dynamic values, the ones a machine writes and a human never should, remain Consul-native, and those are exactly the ones worth exporting on a schedule. ## What to measure afterwards A policy nobody can see the effect of decays. Worth having: an alert when a token authenticates outside its expected prefix; a periodic diff between the repository's desired state and what is actually in Consul, which catches both drift and out-of-band writes; a restore rehearsal on a calendar; and a rule that any recursive delete in automation is conditional or gated, since the unconditional recursive delete is the single operation people most regret running.

  • Why is consul snapshot restore the wrong response to one team deleting their prefix?
    Because it replaces the whole replicated state machine — every team's KV, the ACL tokens and policies, sessions and catalog data — with the snapshot's version. Undoing one prefix would roll back every change made across the platform since the snapshot, including tokens issued in between, turning a single-team incident into a platform-wide one. Prefix-scoped export and import are the proportionate tool.
  • What does down_policy control, and why does the choice matter?
    It decides how a client agent behaves when it cannot reach the servers to resolve a token. extend-cache keeps honouring previously cached decisions, so a control-plane outage does not instantly break every data-plane request; the cost is that a revoked token may be honoured a while longer. The strict setting converts a Consul outage into an authorization outage. Pick according to the failure you fear more.
  • How do you avoid long-lived static tokens in CI?
    Use an auth method so a workload exchanges an identity it already holds — a Kubernetes service-account token, for example — for a short-lived Consul token scoped by a role. That removes the static secret that gets copied between pipelines and outlives its owner, and it makes revocation meaningful because the credential is bound to a workload rather than to a variable in a settings page.
  • When should configuration not live in Consul KV at all?
    When it needs history, review or attribution, none of which the KV store provides — a write overwrites and a delete leaves nothing. Keep that state in a repository and let a pipeline sync it in, so Consul distributes rather than records. Reserve KV for values written by machines and read by machines, where the current value is genuinely the only one that matters.

saying these in an interview costs you the question

  • Uses one shared token for every pipeline and service
  • Assumes a snapshot can restore a single prefix
  • Grants a blanket write policy on the empty key prefix
  • Believes KV keeps previous versions of a value
  • Never rehearses a restore and assumes the backup works

context