An OPA sidecar in 3,000 pods versus one shared decision service — what do you size before choosing?
answer
- everything about one sidecar times N
- base data, not the rule, dominates memory
- requests are paid even when idle
- N independent pollers, correlated restarts
- who bumps the version in 3,000 specs
basics
~20 sMemory multiplied by replica count, and policy-fetch fan-out. Each sidecar holds its own copy of the policy and any base data, and polls for updates independently, so 3,000 replicas mean 3,000 copies and 3,000 pollers.
solid answer
~50 sMemory times replica count, and poll fan-out. A sidecar holds its own parsed policy plus any base data document it loads, so a large approved-accounts list is resident 3,000 times over while a shared server holds it once — measure a single sidecar against your real data before you multiply. Every sidecar is also its own client of the policy-distribution endpoint, so that endpoint sees 3,000 pollers rather than a handful, with a spike whenever a big deployment rolls. Against that the sidecar buys scoped failure and a call that never leaves the pod, and it costs you an upgrade path: moving to a new OPA version means redeploying every workload, whereas the shared service is one rollout its owner performs. The usual honest answer is mixed — sidecars for the few latency- or data-sensitive services, one shared service for the long tail.
go deeper
Know that a sidecar means one OPA process per pod, so its memory and its policy fetching happen once per replica rather than once for the cluster.
Explain what fills a sidecar's memory, why resource requests cost capacity even when idle, and how independent polling changes the load on the policy-distribution endpoint.
Show you would measure a real sidecar against production-sized data, model correlated restarts, and segment the estate rather than declaring one topology the winner.
Own the long-lived cost: a sidecar pins an engine version into every team's manifests, so upgrading the fleet becomes an organizational programme rather than one team's rollout.
## Sizing the fleet, not the process The per-pod sidecar is an attractive default: language independent, failure scoped to one workload, the input never leaves the pod, policy updated at runtime without touching the application artifact. What makes it a real decision rather than an obvious one is that every property of a single sidecar is multiplied by the replica count, and at three thousand replicas the multiplication is the design. ### What actually multiplies **Resident memory.** An OPA process holds the parsed policy and, more importantly, any base data documents it has loaded — the approved-account list, the region map, the per-team exception table. The policy is usually small; the data is not, and the data is what you should measure. A base document that is comfortable in one process becomes a serious number when it is resident in three thousand. The correct move here is not to guess but to run one sidecar against the production-sized data document, read its steady-state footprint, and multiply. Then decide whether the data has to be in memory at all, or whether the rule can be restructured to consult a smaller slice. **Resource requests, not just usage.** In a scheduled environment you pay for what you request. A sidecar with a conservative memory request applied across every pod in the estate consumes schedulable capacity whether or not it is used, and it changes the bin-packing of every workload it is attached to. **Policy-fetch fan-out.** Each sidecar independently fetches policy from wherever policy comes from, on its own interval. Three thousand pollers is a different load profile from ten, and the sharp edge is correlated restarts: a fleet-wide rollout or a node drain sends a large fraction of them to fetch at once. The distribution endpoint has to be sized for the correlated case, not the average. **Upgrade surface.** This is the cost that shows up a year later. The OPA version in a sidecar is pinned in every workload's deployment spec, so moving the fleet to a new engine version is a change to every team's manifests and a redeploy of every workload — coordinated across teams who have their own freezes and priorities. A shared decision service is one rollout, performed by the team that owns it, on their schedule. ### What the shared service costs instead It is not free, it is differently expensive: - **Capacity you must size** against aggregate decision volume from every caller, including the caller that ships a loop and multiplies its query rate overnight. - **A shared component in everyone's path**, whose scope of impact is the whole caller set rather than one pod. That is a real operational commitment: on-call, capacity headroom, upgrade rehearsal. - **Noisy-neighbour dynamics**, where one team's traffic pattern degrades another team's decisions — something sidecars structurally cannot do to each other. - **The input flow** described elsewhere: every caller's input document now arrives at one place. ### How to actually decide Run the numbers rather than the ideology: 1. Measure one sidecar with the real policy and real base data. Multiply by replicas. If that number is a rounding error against your cluster, the sidecar's isolation is cheap and you should take it. 2. Count the pollers and model the correlated-restart case against the distribution endpoint. 3. Ask who will own the OPA version bump in eighteen months, and whether that person can realistically drive a change through every workload's manifests. 4. Segment. Very few estates need one answer for everything. The services with strict data-locality or latency constraints get a sidecar or embedding; the long tail of low-volume callers uses the shared service. A mixed topology is not indecision, provided each placement has a stated reason. ### The answer that lands What distinguishes a strong answer here is refusing the framing that one topology wins. The interviewer is checking whether you know the sidecar's costs are per-replica and mostly invisible until you multiply, whether you would measure rather than estimate, and whether you have thought a version upgrade through to the point where three thousand deployment specs need editing. Say the numbers you would collect and how you would segment, and the question is answered.
- Which part of a sidecar's memory usually dominates, the policy or the data?The data. Parsed Rego for a placement rule is small; the base documents it consults — approved account lists, region maps, exception tables — are what fill the process, and they are resident in every replica. Measure a sidecar loaded with production-sized data rather than with the rule alone, because the rule-only figure will mislead you by an order of magnitude.
- Why is the correlated-restart case the one to size the distribution endpoint against?Because sidecars fetch independently but restart together. A fleet rollout, a node drain or a cluster upgrade sends a large fraction of the pollers to fetch at nearly the same moment, so peak demand on the policy endpoint is set by that event rather than by the steady-state interval.
- What is the strongest argument for the sidecar despite all this cost?Scoped failure and a call that never leaves the pod. One wedged sidecar affects one workload, one team's query volume cannot degrade another's decisions, and the input document stays local. If those properties matter for a service, the per-replica cost is what you pay for them — which is exactly why most estates apply the sidecar selectively rather than everywhere.
saying these in an interview costs you the question
- Sizes one sidecar and never multiplies by replicas
- Assumes the policy dominates memory rather than the data
- Forgets each sidecar fetches policy independently
- Ignores that a version bump touches every workload spec
- Insists one topology must be used estate-wide