skip to content

Why can't a Gatekeeper ConstraintTemplate call http.send, and what replaces it?

level: seniorimportance: should knowfreq 44%

answer

  1. admission sits on the write path
  2. the same rule gets replayed later
  3. a credential in a cluster-wide object
  4. push the fact in, don't pull it
  5. parameters, replicated data, or a provider

basics

~20 s

Gatekeeper evaluates template Rego without http.send. A call during admission would put an outside service on every matching write's path and make results non-reproducible. Facts are pushed in instead: parameters, replicated objects, or an external data provider.

solid answer

~50 s

Admission runs synchronously on every matching write, so a call from inside a rule adds an outbound dependency, and its latency, to every Pod creation in the cluster; when it is slow you get a stalled webhook, and the failure policy then decides between blocking writes and letting them through unchecked. Evaluation also has to be reproducible: the same rule is replayed over live objects and in offline tests, and a rule whose answer depends on what a remote service said at that instant cannot be replayed or explained afterwards. A template is a cluster-wide object too, which is a poor home for a credential. So the design flips from pull to push: put the fact in the Constraint's parameters, have Gatekeeper replicate the objects you need, use a declared external data provider so the platform owns the timeout, or move the lookup to CI and check the resulting field.

go deeper

for a junior

Know that a template's rule cannot make network calls, and that facts it needs must be supplied to it — typically through the Constraint's parameters.

for a middle

Explain the mechanics: admission is synchronous on every matching write, and rules are re-evaluated later, so evaluation has to be a pure function of its input.

for a senior

Pick a replacement and defend it end to end, including the failure mode when the source of the fact is unavailable and what that means for writes cluster-wide.

for a principal

Set the policy for the whole estate: which facts are allowed to become admission dependencies at all, who operates them to what availability, and which checks are deliberately kept off the write path.

## The ban, stated plainly Gatekeeper runs the Rego in a ConstraintTemplate without the `http.send` built-in available. You cannot write a rule that fetches something while it is deciding. Candidates usually meet this the moment they try to write their first genuinely interesting policy — "reject the image unless our service says it was scanned" — and it is worth understanding *why* the answer is no, because the reasons dictate what the replacement has to look like. ## Three reasons, and they compound **1. Admission is on the write path.** Every create or update the Constraint matches waits for the decision. An outbound call means every one of those writes now also waits for a third service. The API server gives a webhook a short window, and if that window closes, the failure policy takes over: either the write is refused — so an unrelated outage now blocks deployments cluster-wide — or the write proceeds unchecked, and the guardrail you shipped is not running. Neither is a thing you want triggered by someone else's rollout. **2. The same rule is replayed.** Gatekeeper does not only decide at admission time; it also sweeps existing objects and re-evaluates the same rules over them, and you will want to run those rules offline against manifests in a rule repository's CI. All of that assumes evaluation is a pure function of its input. A rule that consults a remote service returns different answers on Tuesday and on Wednesday for the same object, which means a recorded result cannot be reproduced, and "why was this blocked?" has no answer you can reconstruct. **3. A template is public within the cluster.** A ConstraintTemplate is a cluster-scoped object that anyone with read access can look at. Any credential the call needed would live there, or be plumbed into the evaluating process for every template to use. Neither is a boundary you want. There is a fourth, quieter reason: a rule that can call out is a rule that can be made to call *anywhere*, so the policy engine becomes an outbound request tool sitting inside the control plane. ## What to do instead, in the order to try it **Push the fact into the Constraint's parameters.** Best when the fact is small and slow-moving — the approved registries, the allowed volume types, the set of valid cost-centre values. Something outside the rule (a controller, a scheduled job, a pull request) keeps the Constraint current, and the values are then visible in the cluster as ordinary reviewable data. The rule stays a pure predicate. **Have Gatekeeper replicate the objects the rule needs** so the rule reads them from its own copy rather than fetching them. This works only for Kubernetes objects, and it trades freshness for determinism: you are reading a cache, and you have to decide what a slightly stale answer means for your rule. **Declare an external data provider.** Gatekeeper supports making the call *on the rule's behalf* through a declared provider, so the timeout, the transport trust and the caching become the platform's configuration rather than something buried inside a policy string. This does not make the latency or the outage go away — you still own the answer to "what happens when the provider is down?" — but it moves that decision out of the rule and into something operable. **Move the lookup off the admission path entirely.** Often the strongest answer. Do the expensive check where it is allowed to be slow — in CI, or in a controller — and have it record the outcome on the object as a label, an annotation or a separate resource. Admission then checks a field on the object in front of it, which is exactly the shape it is good at. The trade you are making is explicit: admission verifies a *claim* someone else made, so the integrity of that claim becomes the thing you have to protect. ## How to answer this in an interview Do not stop at "it is not supported". Name the write path, name the reproducibility requirement, and then show that you know the replacement is a *design change*, not a workaround: the rule stops asking questions and starts checking answers that were pushed to it. The follow-up is almost always "and what happens when your external source is unavailable?" — have a position ready, and make it a deliberate choice between blocking writes and admitting unchecked ones rather than something you discover during the outage.

  • You go with an external data provider anyway. What is the first thing you decide?
    What happens when it is slow or unreachable. That is two separate choices: what the rule does with a missing answer, and what admission does when evaluation cannot finish in time. Pick both deliberately, write down which one you chose to fail towards, and make sure the on-call person can find that decision at three in the morning.
  • Why is caching the remote answer not a sufficient answer?
    A cache still needs a first call on someone's write, so cold start lands on a real user, and eviction puts you back there periodically. Worse, a cached answer is a stale answer: you have not removed the outside dependency, you have made its failures intermittent and harder to reproduce. If a stale answer is acceptable, say so explicitly and fetch it ahead of time instead.
  • When is moving the check into CI the wrong answer?
    When nothing stops an object reaching the cluster by another path. If the check runs in a pipeline and admission trusts a label the pipeline set, anyone who can set that label has bypassed the control. Then admission has to verify something the writer could not have forged, or the check has to stay on the admission path.

A border officer checking passports cannot phone a foreign registry for each traveller; the queue would move at the speed of that phone line. The stamp has to be applied earlier, and the officer checks the stamp.

saying these in an interview costs you the question

  • Assumes caching would make a remote call safe on the write path
  • Wants to embed a service token in a cluster-wide template
  • Ignores that every matching write pays the call's latency
  • Treats a remote-dependent decision as reproducible later
  • Has no answer for what happens when the source is down

context