skip to content

On OPA's Data API, how do PUT and PATCH differ when writing facts?

level: middleimportance: should knowfreq 52%

answer

  1. one verb replaces, one verb edits
  2. the body is the whole truth
  3. an operation list applied in order
  4. RFC 6902 applies all or none
  5. the torn state comes from two requests

basics

~10 s

PUT replaces the entire document at the given path, so anything previously under it is gone. PATCH takes a JSON Patch operation list and edits in place, changing only the paths those operations name.

solid answer

~50 s

`PUT /v1/data/approved/instance_types` writes the document at that path wholesale: whatever was there is replaced by the body, and OPA creates any missing parent objects on the way down. That makes it idempotent and easy to reason about — the writer must always send the complete value. `PATCH /v1/data/approved` takes a JSON Patch array (RFC 6902) of `add`, `remove`, `replace`, `move`, `copy` and `test` operations applied in order to the subtree at that path, so you can add one approved instance type without shipping the other four hundred. JSON Patch is all-or-nothing: if an operation is invalid — for example `add` under a parent that does not exist — the whole patch is rejected and the document is untouched. The real partial-write hazard is not a torn patch but a sync job that issues several requests and dies between them.

code

json · 4 lines
json
[
  {"op": "add",    "path": "/instance_types/m7g.large", "value": true},
  {"op": "remove", "path": "/instance_types/m4.large"}
]

go deeper

for a junior

Know that both verbs write facts into OPA, that PUT sends the whole value for a path and PATCH sends a list of small edits, and that GET on the same path reads back what the engine holds.

for a middle

Be ready to state the semantics precisely: PUT replaces the subtree and creates missing parents; PATCH applies RFC 6902 operations in order and is rejected as a whole if one is invalid.

for a senior

Show you have thought about the writer. Talk about validating the payload before a full replace, making the sync job re-drivable, and why a sequence of requests is where torn state really comes from.

for a principal

Own the contract between the system of record and the engine: who may write which subtree, what a write must prove before it lands, and how a bad fact is rolled back as fast as a bad policy.

### PUT: replace the subtree `PUT /v1/data/<path>` creates or overwrites the document at that path with the request body. The important word is *overwrites*: it is not a merge. If `data.approved` previously held `instance_types` and `regions`, and you PUT an object containing only `instance_types` to `/v1/data/approved`, the regions are gone. OPA will also create missing ancestor objects, so a PUT to a path that does not exist yet succeeds rather than 404s. This makes PUT the natural verb for a reconciling sync job: compute the full current value of the fact, write it, and the engine's state now matches the source of truth regardless of what it held before. It is idempotent — replay it and nothing changes — and there is no drift to reason about. Its danger is the mirror image of its strength. A PUT is a statement that the body *is* the truth, so a bug upstream ships straight through. If the job that assembles the approved instance-type list gets an empty result from its source — an API returning an empty page, a query with a bad filter — the PUT faithfully writes an empty list, and from the next decision onwards every workload is denied because nothing is on the approved list. Guard the write, not the rule: validate the payload before sending it, and refuse a write that shrinks the document beyond a sane threshold. ### PATCH: edit in place `PATCH /v1/data/<path>` takes a JSON Patch document — an array of operations with `op`, `path` and, where relevant, `value` — applied in order to the subtree at that path. `add` inserts or, for an existing object key, sets; `remove` deletes; `replace` requires the target to exist; `test` asserts a value and fails the patch if it does not match. Paths inside the operations are JSON Pointers relative to the path in the URL, and `-` as the last segment of an array pointer means "append". PATCH is the verb for incremental updates from an event stream: a class was approved, so add one key; a class was retired, so remove one. You do not have to hold or transmit the whole document, which matters when the document is large. Its constraints are the RFC's. `add` requires the parent container to exist — you cannot add `/instance_types/m7g.large` if `instance_types` is absent. `replace` requires the target itself to exist. And application is atomic: if any operation is invalid the entire patch fails and the store is left exactly as it was. That is the answer to "what does a half-applied patch leave behind" — nothing, by design. ### Where partial state actually comes from The atomicity is per request, not per intention. A sync job that issues a PATCH to add three new classes and a second request to remove two retired ones, and then crashes between them, has left the engine in a state that never existed in the source of truth: additions applied, removals not. Decisions made in that window are made against a document nobody wrote deliberately. The defences are ordinary distributed-systems hygiene rather than anything OPA-specific. Prefer one request that carries the whole intent — either a single PUT of the reconciled value, or a single PATCH holding every operation — over a sequence. Make the job re-drivable so a crash is repaired by the next run rather than by a human. And use `test` operations when a patch is only valid against a particular prior state, so a stale patch fails loudly instead of applying to the wrong document. ### Reading back and deleting `GET /v1/data/<path>` returns what the engine currently holds at that path. Use it before you theorise about a wrong decision — what the bundle service serves and what a given replica has activated are different questions. `DELETE /v1/data/<path>` removes the document at a path, equivalent to a `remove` operation targeting it. ### Choosing If the producer can cheaply compute the whole fact, PUT it and stop reasoning about drift. If the fact is large and the producer sees changes as events, PATCH it, and periodically PUT a full reconciliation anyway so that a missed event does not live in the engine forever.

  • A sync job PUTs the full approved list every five minutes. What is its worst failure mode?
    A bad upstream read that returns nothing. PUT means the body is the truth, so an empty list is written faithfully and every workload is denied from the next decision onwards — a total outage caused by a data job, not by a policy change. Validate the payload before writing and refuse writes that shrink the document implausibly.
  • Your patch of ten operations fails on the seventh. What is in the store?
    Exactly what was there before. JSON Patch is applied atomically, so an invalid operation rejects the whole document and you can retry safely. The state to worry about is the one produced by a job that sends several separate requests and dies between them, leaving additions applied and removals pending.
  • When would you use a test operation in a patch?
    When the patch is only meaningful against a particular prior state. A test operation asserts a value at a pointer and fails the entire patch if it does not match, so a stale or out-of-order update from an event stream is rejected loudly instead of being applied to a document it was never computed against.

saying these in an interview costs you the question

  • Thinking PUT merges into the existing document
  • Assuming a failed JSON Patch leaves some operations applied
  • Writing an empty list from a failed upstream read, then blaming the rule
  • Using add on a path whose parent object does not exist

context