In Gatekeeper, what does an external data Provider add to a rule, and what does a slow one do to admission?
answer
- some facts are not cluster objects
- a registered service, not a free HTTP client
- url, timeout and a CA to trust
- keys go out, values or errors come back
- the call sits inside the admission request
basics
~20 sA Provider registers an HTTPS service Gatekeeper may call during evaluation, letting a rule resolve a fact no cluster object holds — an image tag to its digest. The call runs inside the admission request, so provider latency becomes request latency.
solid answer
~50 sReplication only helps with facts that are cluster objects. For anything else — resolving an image tag to a digest, asking an inventory service who owns a repository — Gatekeeper's external data feature lets a rule call out at decision time. You register a `Provider` naming the service URL, a timeout in seconds and a CA bundle to trust its certificate; the rule then calls the `external_data` builtin with the provider name and a list of keys, and the service answers with a value or a per-key error for each key, plus a system-level error field. Two things bite. The call happens inside the admission request, so every matching create pays the provider's latency, and the API server's own webhook deadline decides what happens beyond that. And the rule must handle the error fields explicitly — ignoring them means an unresolvable key quietly produces no violation.
go deeper
Know that a rule can be given facts from outside the cluster by registering a Provider service, and that the rule asks for them by key during the admission decision.
Explain the Provider's url, timeout and CA bundle, that the rule calls out with a batch of keys, and that the response carries per-key errors as well as values.
Show you handle the failure path deliberately — unresolved keys and provider outages as explicit policy outcomes — and that you treat the provider as a latency and availability dependency on every matching create.
Own the call about putting a synchronous external dependency on the workload-create path at all: which facts justify it, what availability the provider must hold, and which way the gate fails when it does not.
## What the feature is for A replicated cache answers questions about *cluster objects*. Plenty of policy questions are not about cluster objects at all: what digest does this image tag currently resolve to, is this repository registered to a team, is this base image on the approved list maintained outside the cluster. None of that is a Kubernetes resource, so no amount of syncing will make it visible. Gatekeeper's external data feature covers that gap by letting a rule call an HTTPS service during evaluation. It is deliberately narrow — this is not a general HTTP client dropped into Rego, and the ordinary Rego HTTP builtin is not available inside a ConstraintTemplate. ## The pieces **The Provider resource.** A cluster-scoped object in Gatekeeper's external-data API group that registers one service. The fields that matter operationally are: - `url` — the service endpoint. HTTPS, and typically an in-cluster Service. - `timeout` — how long, in seconds, Gatekeeper will wait for a response before giving up on the call. - `caBundle` — the certificate authority Gatekeeper should trust for that endpoint. Get this wrong and every call fails on TLS verification, which surfaces as a provider error rather than as anything resembling a certificate message in the rule. The feature is not on by default; the controller has to be started with it enabled. **The builtin.** Inside a rule you call `external_data` with an object naming the provider and the list of keys you want resolved — for an image-digest lookup, the keys are the image references pulled out of the object under admission. Batching matters: send all the images in one call rather than one call per container. **The response.** The service replies with a structured response that pairs each requested key with a value or with a **per-key error**, alongside a **system-level error** field for a failure that is not about any particular key. The Rego value you get back exposes both: the resolved responses, the per-key errors, and the system error. ## Handling the errors is the policy decision The most common review finding on an external-data rule is that it reads only the success path. If a key could not be resolved, and the rule only produces a violation when it positively sees a bad value, then an unresolvable key produces no violation — the request is admitted precisely in the case where you know least about it. So the rule has to state, explicitly, what an unanswered question means: - Check the per-key errors and decide whether an unresolved key is itself a violation. For a rule that requires a resolvable digest, it should be. - Check the system-level error and decide whether a provider-wide failure denies or allows. This is the fail-closed/fail-open decision, made in your rule rather than in the platform, and it deserves to be a conscious choice rather than an omission. Write the messages so the two cases read differently: *"image tag could not be resolved to a digest"* is a very different conversation from *"image resolved to a digest that is not permitted"*. ## What a slow provider does This is the operational half of the question, and the part that gets asked from the platform team's chair. The call is made **during** the admission request. There is no background refresh and no prefetch: the API server is waiting, the client is waiting, and the create does not complete until the provider answers or the timeout expires. So the effects stack up: 1. **Every matching create pays the latency.** A provider that normally answers in 20 ms and degrades to 2 s turns every workload create matched by that constraint into a two-second operation. 2. **The Provider timeout bounds one call, not the request.** It stops a hung provider from pinning the evaluation forever, but a provider sitting at its ceiling still costs the full timeout on every request. 3. **The API server has its own deadline for the webhook.** Exceed it and the call fails from the API server's point of view, and what happens then is governed by the webhook's configured failure policy — fail the request or ignore the webhook. That is the setting on which "a slow provider" becomes either a cluster-wide outage or a silently unenforced control. The operating discipline follows from that. Keep the provider's scope tight so few requests trigger it. Run it in-cluster and highly available, because you have just put a network dependency on the create path of your own workloads. Set the timeout to something you would actually be willing to add to every request, not to something generous. Monitor its latency the way you would monitor a database on a request path — and decide in advance which way you want the gate to fail when it is down, because the choice is being made either way. ## The comparison an interviewer is fishing for Replicate a fact when it is a cluster object, small, and read by rules often. Call out for it when it is not a cluster object, or when freshness at the moment of decision genuinely matters more than the availability you are giving up. The cache costs memory and gives you staleness; the call costs latency and availability and gives you a current answer.
- Your external-data rule denies unapproved digests but admits everything while the provider is down. Why?Because the rule only produces a violation on the success path. When the provider fails, the per-key errors and system error are populated and no value is resolved, so nothing matches the deny condition and the request sails through. Handle the error fields explicitly and decide whether an unanswered question denies — for a rule of this kind it should.
- Would you replicate a fact into the inventory or resolve it through a provider?Replicate it when it is a cluster object, small enough to hold in memory, and read frequently — you trade memory and a little staleness for no request-time cost. Call a provider when the fact is not a Kubernetes object at all, or when the answer must be current at the moment of decision, accepting latency on every matching request and a new availability dependency on the create path.
saying these in an interview costs you the question
- Treats external_data as a general HTTP client for any rule
- Ignores the per-key errors and the system error field
- Assumes the call happens in the background, not in the request
- Sets a generous timeout without counting it against every request
- Forgets the provider becomes an availability dependency of workload creates