skip to content

OPA is embedded in ten Go services with the policy baked into each binary — how does a new rule roll out?

level: middleimportance: should knowfreq 44%

answer

  1. no push channel for a baked-in rule
  2. ten release trains, not one deploy
  3. two rule versions live at once
  4. log the revision that was evaluated
  5. runtime bundles trade lag for a dependency

basics

~20 s

One rebuild and release per service. With Rego compiled into the binary there is nothing to push, so enforcement coverage follows ten release trains, and for a while some services enforce the new rule and some the old one.

solid answer

~50 s

You cannot flip a switch. With policy compiled into each binary, a new rule reaches production only when each service rebuilds, re-tests and rides its own release train, so coverage is a property of ten release schedules rather than one deploy. Expect a mixed fleet: for days, some services enforce the new approved-region list and some enforce the previous one. That is fine when the change is a tightening you deliberately want to stage, and dangerous when someone is telling an auditor the control is on everywhere. Make it observable — have each service report the policy revision it evaluated in its decision logs or on a status endpoint — so 'who is enforcing this?' is answered with data, not a spreadsheet. If the lag is unacceptable, configure the embedded SDK to fetch policy at runtime, which buys freshness back at the cost of a runtime dependency inside every process.

go deeper

for a junior

Know that when a rule is compiled into a service binary, changing the rule means rebuilding and redeploying that service — there is nothing to push at a running process.

for a middle

Explain why coverage becomes a per-service release property, what the fleet looks like mid-rollout, and how runtime bundle fetching changes that picture and what it costs.

for a senior

Demonstrate that you would instrument the rollout — the policy revision reported with each decision — so control coverage is a number you can produce rather than an assumption.

for a principal

Own the trade you are making for the organization: baked-in policy gives reproducible artifacts and slow uniform change; runtime fetching gives fast fleet-wide change, including of mistakes.

## Enforcement coverage becomes a build-and-release property When OPA is embedded as a Go library and the Rego is compiled into the service binary, the policy is part of the application artifact. That has one very large consequence that candidates routinely miss: **there is no rollout mechanism for the rule that is separate from the rollout mechanism for the code.** Updating the policy is a source change, a build, a test run, a review, and a deploy — ten times over for ten services. ### What actually happens on the day You merge the new rule — say the approved-region list drops a region that the company is exiting. Nothing changes in production. Service A picks the change up in its next release that afternoon. Service B is mid-freeze and ships on Thursday. Service C has not been released in four months and nobody remembers how. Service D vendored the policy months ago and has drifted. For the whole of that window the fleet is enforcing two different versions of the same control, and the interesting question is not whether that is acceptable — often it is — but whether you can *say which services are on which*. ### Staged is not the same as unknown A staged rollout is a feature: tightening a rule service by service is exactly how you avoid breaking every provisioning path at once. What turns it into a defect is losing track of the state. Two habits fix that: - **Emit the policy revision with the decision.** Have each service include the policy version or bundle revision it evaluated in its decision log entries, or expose it on a status endpoint. Then a query over decision logs answers 'which services enforced the new region list yesterday?' directly. - **Treat the fleet state as the reporting unit.** 'The rule is merged' is not a control status. 'Nine of ten services enforce revision 47, one is on 46, expected Thursday' is. This matters most when someone outside engineering depends on the answer. A control that exists in the repository but runs in six of ten services is a control with a coverage number, and the honest number is the one you can produce from telemetry rather than from intent. ### The alternative, and what it costs Embedding does not force you to bake the policy in. OPA's Go SDK can run a managed embedded instance configured the same way a standalone OPA is: it downloads policy bundles from a service on an interval and swaps them in without a restart. That restores freshness — the new rule reaches every process minutes after publication, no rebuild involved — but it changes the shape of the deployment in ways you should say out loud: - Every process is now a client of a policy-distribution endpoint, so a change in *its* availability or contents reaches everything at once. You have traded a slow, staged, per-service rollout for a fast, simultaneous, fleet-wide one — in both directions, including for a bad rule. - The service's behaviour is no longer fully determined by its own artifact. Reproducing an old decision means knowing which policy revision that process had loaded at the time, which is another reason to log the revision with the decision. - Startup now has an ordering concern: the process must have a policy before it can answer, so a cold start with an unreachable distribution endpoint is a state you have to define. ### The judgment Baked-in policy suits a small number of high-value services where you want the artifact to be self-contained, auditable and reproducible, and where a per-service release is not a hardship. Runtime bundle fetching suits a fleet where the lag between merging a rule and enforcing it everywhere is the thing you actually care about. The wrong answer in an interview is to describe embedded OPA as if policy updates were free — they are as fast as your slowest release train, and that number is the one you should be able to quote.

  • Is a staged rollout of a policy change actually a problem?
    Not by itself — staging a tightening service by service is how you avoid breaking every caller at once. It becomes a problem when nobody can say which services are on which version, because then the control has an unknown coverage number. Emit the evaluated policy revision with each decision and the staging stays deliberate rather than accidental.
  • How would you get freshness back without giving up embedding?
    Run the embedded instance through OPA's Go SDK configured to download policy bundles at runtime instead of compiling Rego into the binary. The rule then lands in every process minutes after publication with no rebuild. In exchange, every process becomes a client of the distribution endpoint and a bad rule also reaches everything at once.
  • A service has not been released in four months. What is its policy status?
    It enforces whatever rule was current four months ago, and it will keep doing so indefinitely — nothing expires. That is the clearest argument for reporting coverage from telemetry rather than from the repository: the merge date of a rule says nothing about the services that never picked it up.

saying these in an interview costs you the question

  • Says a policy update reaches embedded services automatically
  • Assumes a pod restart picks up a rebuilt rule
  • Reports the rule as live because it merged
  • Ignores that two policy versions run concurrently during rollout
  • Treats runtime bundle fetching as free of new dependencies

context