skip to content

You are adding ETag-based preconditions to a resource with many independently edited fields and embedded sub-objects. How do you decide what the ETag value should be computed over, and what breaks if you choose badly?

level: seniorimportance: should knowfreq 34%

answer

  1. Validator scope must cover the write scope
  2. Too coarse = false conflicts; too fine = lost updates
  3. Split the resource, don't narrow the validator
  4. Version columns, not serialized-byte hashes
  5. Keep volatile derived fields out of the version

basics

~20 s

Compute the validator over exactly the state the endpoint protects. Whole-resource validators are simple but cause false conflicts when clients edit unrelated fields; finer granularity means exposing those parts as their own resources with their own ETags, not narrowing the token on the same URL.

solid answer

~50 s

Granularity is a conflict-rate decision. - **Whole-resource validator** (one version per aggregate): simple and always correct, but any change anywhere invalidates every client's token. On a hot object edited by many people on disjoint fields you get constant 412s that are not real conflicts. - **Finer validators**: if two parts are genuinely independent, model them as their own resources (`/orders/42/shipping-address`, `/orders/42/notes`) with their own ETags. A note edit then does not invalidate an address edit. The rules I apply: the validator must cover everything the write can change and everything the client's decision depended on. Never compute it over a subset of the fields a `PUT` replaces, or you reintroduce silent lost updates. And do not derive it from serialized bytes if the representation varies by content negotiation, sparse fieldsets or serializer version — every deploy would invalidate all tokens. So: split the resource, do not narrow the validator.

go deeper

for a junior

Know that the ETag represents the version of the whole resource you fetched, and that any change to it invalidates your token.

for a middle

Explain the two failure modes — false conflicts when coarse, missed conflicts when narrower than the write scope — and that the validator must cover everything the write can change.

for a senior

Argue for splitting resources over narrowing validators, discuss PATCH-scoped validators, and reject byte-hash ETags on representation-stability grounds.

for a principal

Treat granularity as a product decision driven by measured conflict rates and the editing model, and set a house rule so services do not each invent their own.

## What the validator is promising An ETag is a promise: "if this token is unchanged, nothing you care about has changed." The design question is what *this* covers. Two failure directions exist, pulling in opposite ways. - **Too coarse** → false conflicts. One version per aggregate means an edit to a comment invalidates a pending edit to the shipping address. Clients see 412s for changes that could never have conflicted, and each costs a re-read, a re-merge, and possibly a user prompt. - **Too fine** → missed conflicts. If the validator covers only the fields you *think* the caller edits, a `PUT` that replaces the whole representation can still wipe fields outside the validator's scope, and the precondition happily passes. You have built ceremony with no safety. ## The rule that resolves it The validator must cover **at least** the union of (a) everything the write can modify and (b) everything the client's decision depended on. For a `PUT` that replaces the representation, that is the entire representation, so a whole-resource validator is the only correct choice. If that produces too many false conflicts, the fix is not a narrower validator on the same URL — it is **splitting the resource**. ## Splitting instead of narrowing If `note` and `shippingAddress` are independently edited, promote them: `PUT /orders/42/shipping-address` with its own ETag over the address only. Now write scope and validator scope match exactly, so the safety property holds while conflicts narrow. The parent `/orders/42` keeps a validator over the whole aggregate, which is right for callers that replace the whole thing. The cost is more URLs, more round trips to assemble a view, and a parent ETag that must change whenever a child does, or a cached parent goes stale. `PATCH` deserves special mention. A patch that only touches `note` conflicts with another patch only if that one also touched `note`, yet a whole-resource validator rejects both. Some teams accept the false conflicts for simplicity; others define field-scoped validators for a patch-only endpoint, which is defensible precisely because the write scope is narrow and explicit. If you do that, document it loudly — it changes what the 412 means. ## How to compute the value Prefer a monotonic version (or a hash over the version columns of the covered rows) rather than a hash over serialized response bytes. Byte hashes are seductive because they are automatic, but they bind the token to representation details: content negotiation, sparse fieldsets (`?fields=id,status`), field ordering, and serializer upgrades all change the bytes without changing state. A deploy then invalidates every outstanding validator and conditional writes fail all at once. Also decide about derived and volatile fields. If the representation embeds a computed `viewCount` or a recomputed price, including them makes the token churn on activity nobody edited. Keep such fields out of the aggregate's version, or move them to their own sub-resource. ## Judging the result in production Instrument it: measure the 412 rate and, where you can, how often the conflicting writes actually touched overlapping fields. A high 412 rate with low true overlap means the granularity is too coarse and the resource should be split. A near-zero rate alongside reports of lost edits means the validator is not covering what the write changes.

  • Why is hashing the JSON response body a risky way to generate an ETag?
    The hash tracks the representation, not the state. Content negotiation, sparse fieldsets, key ordering and serializer upgrades change the bytes while the resource is unchanged, so every client's stored validator dies at deploy time and conditional writes start failing en masse. A version column derived from stored state is stable across all of those.
  • If a sub-resource gets its own ETag, what should happen to the parent resource's ETag when the sub-resource changes?
    The parent's validator must change too, because the parent representation embeds the child and a stale parent token would let a whole-resource PUT overwrite the child's newer state. Usually you bump an aggregate version on any child write, or compute the parent ETag from the child versions.

saying these in an interview costs you the question

  • Computing the validator over only the fields the client is 'expected' to edit while allowing a full-representation PUT
  • Assuming finer validators are always better, without noticing they weaken the safety guarantee
  • Hashing the serialized response body and being surprised when a serializer change invalidates all tokens
  • Including volatile counters or computed prices in the version so tokens churn constantly
  • Giving a sub-resource its own ETag but never bumping the parent's

context