Your OPA sidecars hold a 400 MB data document and keep OOMing. What do you do?
answer
- the archive size tells you nothing
- no sharding, one copy each
- sidecars multiply by pod count
- carry only fields rules reference
- key the object by the lookup value
basics
~20 sEvery OPA instance holds the whole data document in memory, so the cost is paid per replica and per sidecar. Shrink it: ship only the fields rules read, scope bundle roots per consumer, and key the document by the lookup value.
solid answer
~50 sStart with where the cost comes from: OPA parses the data document into in-memory values, so the footprint is materially larger than the JSON on disk and far larger than the compressed bundle. There is no sharding — each instance holds all of it, so a sidecar model multiplies the document by the pod count. Three levers, in order of payoff. **Ship less**: an inventory record has thirty fields and the rule reads two, so project the document down at bundle-build time to what rules actually reference. **Ship to fewer engines**: split bundle roots so each consumer carries only its slice. **Reshape it**: an array of records forces a scan per decision, whereas an object keyed by the value the rule looks up is a single reference. Size limits for the activation peak, not steady state.
go deeper
Know that OPA keeps the data document in memory in each instance, so a big document is a real cost in every pod that runs one, not a one-off cost somewhere central.
Be able to explain why the parsed footprint exceeds the archive, that there is no sharding between replicas, and how an object keyed by the lookup value beats an array the rule must scan.
Show the diagnosis and the fix: measure the parsed size, project the document to the fields rules reference, split bundle roots per consumer, and size limits for the activation peak rather than steady state.
Own the growth curve. Decide who is allowed to add facts to a shared bundle, what size budget the build enforces, and at what point the estate stops replicating an inventory into every engine at all.
### Where the memory goes A bundle arrives as a gzipped tarball, so the number you see in a transfer log is the least informative number available. OPA decompresses it and parses the JSON into in-memory values, and the resulting footprint is substantially larger than the raw JSON text, which is itself substantially larger than the archive. Sizing a memory limit from the bundle size is the classic mistake here. The second multiplier is topology. OPA does not shard a data document across instances; each one holds the whole thing so it can answer any query locally, which is the entire point of putting the engine next to the workload. A 400 MB document in a deployment of six replicas is 2.4 GB of cluster memory. As a sidecar next to three hundred pods it is 120 GB, and the pods are now mostly policy engine. That arithmetic, not any single OOM, is usually what forces the redesign. ### Lever one: carry fewer fields The most common cause of a giant data document is that someone piped an inventory export straight into the bundle. Each record has thirty fields — owners, tags, timestamps, descriptions, links — and the rules reference two of them. Project at build time: the bundle build reads the export and emits only the fields policies actually reference. This is usually the single biggest reduction available and it costs a transformation step in a pipeline that already exists. Do the same audit for whole subtrees. Data that was added for a rule that has since been deleted stays in the bundle forever, because nothing fails when it is present. ### Lever two: carry it to fewer engines A bundle declares roots, and there is no requirement that every OPA fetch the same bundle. If the approved instance-type lists differ per business unit, or one consumer needs the node-class facts and another only needs the registry facts, split them so each engine fetches the slice it evaluates against. The estate-wide document is convenient for whoever builds the bundle and expensive for everyone who runs one. ### Lever three: shape it for the lookup How the document is shaped changes both memory and evaluation cost. An array of records forces the rule to scan: ```rego # data.approved.instance_types is an array of objects deny contains msg if { count([t | t := data.approved.instance_types[_]; t.name == input.instance_type]) == 0 msg := sprintf("instance type %v is not approved", [input.instance_type]) } ``` That comprehension is evaluated per decision and grows with the list. Keyed by the value the rule looks up, the same guardrail becomes a single reference: ```rego # data.approved.instance_types is an object: {"m7g.large": true, ...} deny contains msg if { not data.approved.instance_types[input.instance_type] msg := sprintf("instance type %v is not approved", [input.instance_type]) } ``` The reshape also forces the useful question: what does each record need to carry *besides* the key? Very often the answer is nothing, and an array of fat objects collapses into a set of keys. One correctness note to state deliberately, because it is the trap in the keyed form: if the whole subtree is missing, the reference is undefined for every input, the negation succeeds, and the rule denies everything. Cheap memory wins can turn a data-loading bug into a total outage, so pair the reshape with a check that the document is present and non-trivial before the engine is considered ready. ### Sizing for the peak Bundle activation is not free: the new document is built before the old one can be released, so the high-water mark during a swap exceeds steady state. A limit set from a quiet-period reading will hold for hours and then kill the pod the moment a new bundle lands — which looks like a random OOM and is actually a perfectly periodic one. Size for the activation peak, and watch document size as a trend so you find out that the estate's facts grew forty percent this quarter before the pods do. ### What to say when asked Name the multiplier first (per instance, not per cluster), then the three levers in order of payoff (fewer fields, fewer engines, better shape), then the operational detail that separates people who have run this from people who have read about it: measure the parsed footprint, not the archive, and leave headroom for activation.
- Beyond memory, why does keying the document by the lookup value help?It removes work from every decision. Scanning an array with a comprehension costs time proportional to the list on each evaluation, while a keyed object resolves in a single reference. It also makes the rule shorter and its intent obvious, and it exposes how much of each record was never needed once the key itself carries the answer.
- Your OPA pods have memory limits set from steady-state usage and OOM only when a new bundle lands. Why?Because activation builds the incoming document before the outgoing one is released, so the swap has a higher high-water mark than either state alone. Limits derived from a quiet-period reading are therefore under-sized by exactly the amount that matters, and the OOM tracks the publish cadence rather than traffic.
- How would you stop a shared bundle growing until it breaks every consumer at once?Give the bundle build a size budget that fails when exceeded, trend the parsed document size rather than the archive size, and split roots so each consumer fetches only its slice. A single estate-wide document couples every engine's memory ceiling to the growth of the noisiest data producer.
saying these in an interview costs you the question
- Sizing the memory limit from the compressed bundle size
- Assuming replicas share one copy of the data document
- Shipping a whole inventory export when rules read two fields
- Ignoring the higher memory high-water mark during bundle activation
- Adding replicas in the belief that the data is sharded across them