skip to content

Replicated Object Cache

Some Gatekeeper rules cannot be decided from the object in hand - unique ingress hostnames need every other one. Interviewers probe the cache that allows it and the race it quietly introduces.

on this pageshow

questions

4

In Gatekeeper, what do the Config and SyncSet objects do, and where does the synced data appear to a rule?

level: juniorimportance: must knowfreq 55%

answer

  1. admission shows one object only
  2. the rule needs siblings it cannot fetch
  3. kinds chosen by group, version, kind
  4. Config's syncOnly, or a SyncSet
  5. cluster and namespace branches of one document

basics

~20 s

Gatekeeper's Config (spec.sync.syncOnly) and SyncSet objects choose which resource kinds are replicated into its in-memory cache. Those objects then appear to Rego under data.inventory, so a constraint can read cluster objects other than the one being admitted.

solid answer

~40 s

An admission request only carries the object being created or updated, so a rule that needs a *sibling* object — say, "this namespace must already hold a default-deny NetworkPolicy" — has nothing to read. Gatekeeper solves that by replicating chosen kinds into its own cache. You list them either in the singleton `Config` object (named `config` in `gatekeeper-system`, under `spec.sync.syncOnly`) or in a cluster-scoped `SyncSet` (under `spec.gvks`); several SyncSets may exist and the sets are unioned, so a policy bundle can ship its own data requirement without editing one shared object. Each entry is a group, version and kind. Replicated objects show up in Rego at `data.inventory.cluster[<groupVersion>][<kind>][<name>]` and `data.inventory.namespace[<ns>][<groupVersion>][<kind>][<name>]`. It is a watch-fed replica held in memory by every Gatekeeper pod, not a live query — so sync only what a rule actually reads.

go deeper

for a junior

Be ready to say that an admission request carries only the object under review, and that Gatekeeper's Config or a SyncSet lists which other kinds get copied into data.inventory for a rule to read.

for a middle

Explain the address shape — cluster versus namespace branches keyed by groupVersion, kind and name — and that the union of every SyncSet and the Config's syncOnly defines what is replicated.

for a senior

Show you treat the inventory as a replica you operate: only the kinds a rule reads, memory sized for what you replicate, RBAC to list and watch them, and a deliberate check that the sync is actually working.

for a principal

Own the standing rule for what a shared cluster is allowed to replicate, and who reviews a new syncOnly entry, so one team's policy does not silently add the whole Pod population to every Gatekeeper pod's heap.

## Why a replicated cache exists at all When the API server calls Gatekeeper's admission webhook, the request it sends carries a narrow slice of the world: the object being created or updated, the previous version of that object on an update, who is making the request and which groups they belong to, and whether this is a dry run. It carries **no other object, no cluster state and no history**. That is a property of the admission contract itself, not a Gatekeeper limitation. Rego evaluated inside a ConstraintTemplate cannot go and fetch the missing object either — there is no API client in the evaluation. So a rule like *"a workload may only be created in a namespace that already holds a default-deny NetworkPolicy"* is unanswerable unless the NetworkPolicy is **already present in the data document** when the rule runs. Populating that document is exactly what the sync configuration does. ## Choosing the kinds: Config and SyncSet Two objects control replication, and they do the same job: - **`Config`** — a singleton, named `config` in the `gatekeeper-system` namespace. Its `spec.sync.syncOnly` is a list of `{group, version, kind}` entries. This is the original mechanism, and because there is exactly one of these objects, every team that needs a kind replicated has to edit the same resource. - **`SyncSet`** — a cluster-scoped resource whose `spec.gvks` is the same list of group/version/kind entries. You may create as many as you like. The effective set of replicated kinds is the **union** of every SyncSet plus the Config's `syncOnly`. That lets a policy bundle ship its own data requirement alongside its templates and constraints, instead of coupling every bundle to one shared object. A core resource has an empty group, so a Pod is `{group: "", version: "v1", kind: "Pod"}`; a NetworkPolicy is `{group: "networking.k8s.io", version: "v1", kind: "NetworkPolicy"}`. ## Where the data lands Gatekeeper writes replicated objects into the Rego data document under `data.inventory`, split by scope: ``` data.inventory.cluster[<groupVersion>][<kind>][<name>] data.inventory.namespace[<namespace>][<groupVersion>][<kind>][<name>] ``` `<groupVersion>` is the string form: `"v1"` for core resources, `"networking.k8s.io/v1"`, `"apps/v1"` and so on. The leaf value is the whole object as JSON, exactly as the API server serves it — so `.spec`, `.metadata.labels` and the rest are all reachable. Note the two different things a rule reads. `input.review.object` is **the object under admission**, delivered by the API server for this request. `data.inventory` is **everything else**, delivered by the sync controller ahead of time. Confusing them is the classic first mistake. ## It is a replica, and replicas have properties The sync controller watches the configured kinds and mirrors them into memory. Three consequences follow, and interviewers probe all three: 1. **It is eventually consistent.** An object created a moment ago may not be in the cache when the next admission request is evaluated. A rule that requires a sibling to exist can therefore reject a batch that was applied in the wrong order — and the same apply succeeds on retry. 2. **It costs memory.** Every replicated object is held by every Gatekeeper pod. Syncing a small, low-cardinality kind such as NetworkPolicy costs almost nothing; syncing Pods, Secrets, ConfigMaps or Events in a large cluster can dominate the controller's footprint and push it into restarts. The discipline is to sync only the kinds a rule genuinely reads, and to reach for a narrower kind when one will do. 3. **It needs permission.** Gatekeeper can only replicate kinds its ServiceAccount is allowed to `list` and `watch`. If you have narrowed its ClusterRole, an entry for a kind it cannot watch simply yields an empty branch of the inventory — no error at the point where the rule runs. ## Verifying it works Because a missing sync is silent at evaluation time, verify it deliberately rather than assuming. The controller logs the kinds it starts watching and reports sync activity through its exported metrics, and the cheapest end-to-end check is a throwaway non-blocking constraint whose message reports what it can see in `data.inventory` for a known namespace. If it reports nothing, the problem is the sync configuration or the RBAC, not the rule.

  • Two policy bundles each ship a SyncSet, and the old Config still lists syncOnly kinds. Which wins?
    None of them wins — the sets are unioned. A kind listed in any SyncSet or in the Config's `syncOnly` is replicated. That is the point of SyncSet: separate bundles declare their own data needs without coordinating edits to a single shared object. The flip side is that nothing shrinks the set implicitly; removing a kind means removing it everywhere it is listed.
  • You add Secrets to the synced kinds and the Gatekeeper pods start restarting. What happened?
    Every replicated object is held in memory by every Gatekeeper pod, so a high-cardinality or large-payload kind can blow the container's memory limit and get it OOM-killed. Secrets, Pods, ConfigMaps and Events are the usual offenders. Sync only kinds a rule actually reads, prefer the narrowest kind that answers the question, and size the controller's memory request against the objects you have chosen to replicate.
  • How would you confirm a kind is really being replicated before you rely on it?
    Do not infer it from the Config being applied. Check that Gatekeeper's ServiceAccount may list and watch that kind, look at the controller logs for the watch starting, and deploy a throwaway non-blocking constraint whose violation message reports what it finds under `data.inventory` for a known namespace. An empty result there points at sync configuration or RBAC, not at your rule logic.

It is a read replica, not a database query. The rule reads a copy that someone arranged to have on hand, and everything true of a replica — lag, memory cost, needing the right credentials to fill it — is true here.

saying these in an interview costs you the question

  • Says the Rego rule queries the API server during admission
  • Assumes every cluster object is in data.inventory by default
  • Confuses data.inventory with input.review.object
  • Treats syncing extra kinds as free
  • Thinks Config and SyncSet override each other rather than combining

context

open as a page

How does a Gatekeeper rule read a sibling object from data.inventory, and what happens if that kind was never synced?

level: middleimportance: should knowfreq 48%

basics

~10 s

A rule indexes data.inventory by scope, groupVersion, kind and name. An unsynced kind makes that lookup undefined, not false — so the surrounding logic decides whether the rule blocks everything or waves everything through.

open as a page

A Gatekeeper constraint blocked a Deployment for a missing NetworkPolicy applied seconds earlier — why, and what do you do?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Replication lag. A Gatekeeper rule reads data.inventory, a watch-fed replica rather than a live query, so the NetworkPolicy existed in the API server but had not reached the cache when the decision was made.

open as a page

In Gatekeeper, what does an external data Provider add to a rule, and what does a slow one do to admission?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A Provider registers an HTTPS service Gatekeeper may call during evaluation, letting a rule resolve a fact no cluster object holds — an image tag to its digest. The call runs inside the admission request, so provider latency becomes request latency.

open as a page