You are designing a JSON REST API and want conditional GET support so repeat readers can be answered with HTTP 304 Not Modified. What does the server have to implement for those 304s to actually be cheap, where do the savings land, and when would you decide the feature is not worth adding at all?
answer
- validator cheaper than payload, or it's bytes-only
- version column, not a hash of the rendered body
- check before you render — short-circuit
- saves egress + client parse, never the round trip
- tag must cover expands, scope, serializer version
basics
~20 sA 304 pays off only if deciding "unchanged" is far cheaper than building the response — derive the validator from a stored version column or a hash saved at write time, and check it before rendering. Savings are egress bytes and client parsing; origin CPU only if you short-circuit.
solid answer
~50 sFirst decide where the win has to land. If I build the ETag by rendering the JSON and hashing it, I have already done the queries and the serialization — I save egress and client parsing, and nothing at the origin. To make 304s genuinely cheap the validator must come from something readable without building the representation: a monotonic version column, a per-aggregate change counter, or a content hash stored at write time. The handler checks it first and short-circuits. The validator must cover everything that varies the bytes — embedded children, `include`/sparse-fieldset params, locale, the caller's auth scope, and my own serializer version so a deploy invalidates old copies. Weak ETags are the honest label for JSON. Worth adding for large payloads, read-heavy resources, metered or CDN-fronted clients. Not worth it for small responses, or resources that change on nearly every read.
code
http · 7 linesHTTP/1.1 200 OK
Content-Type: application/json
ETag: W/"order-8f21:v47:s12"
Vary: Accept-Language, Authorization
Cache-Control: private, no-cache
{"id":"8f21","status":"SHIPPED","lines":[ ... ]}go deeper
Know that 304 means "nothing changed, reuse your copy" and that the body is what gets saved. Being able to say the server compares a validator it sent earlier is enough at this level.
Name the three validator sources — regenerate-and-hash, hash stored at write, version column — and say which of them actually spares the database. Mention that the tag must change when embedded data or the payload shape changes.
Lead with where the savings land and what the short-circuit has to look like in the handler. Cover validator correctness across expands, auth scope and serializer version, collections being the hard case, and give concrete conditions under which you would not add 304 support.
Frame it as a cost decision: measure payload size, read/write ratio and egress spend before implementing, compare against freshness lifetimes and delta endpoints that beat 304 outright, and weigh the standing staleness risk a validator introduces. Decide it per resource class and per cache topology, not as a blanket API policy.
## The design question, not the wire question A conditional GET lets a server answer "your copy is still current" instead of resending the representation. The header-level handshake that produces the 304 belongs to HTTP caching; the API-design question is different and much more consequential: **how much work does your server do before it can decide?** If the answer is "all of it", you have bought a bandwidth optimisation and nothing else — which is sometimes fine, but you should know that is what you bought. ## Three ways to source the validator **1. Regenerate and compare.** Run the query, build the DTO, serialize it, hash the bytes, compare with what the client sent. Correct and trivial to bolt on (most frameworks offer it as a filter), but the origin has done 100% of its normal work. You save egress and client-side parsing. Nothing else. **2. Hash stored at write time.** Every write recomputes and persists the content hash of the canonical representation. Reads fetch one small column. Cheap on read, but the hash is now tied to a particular rendering, so it must be recomputed whenever the serializer changes, and it is awkward when the representation depends on request parameters. **3. Version or state token.** Each row (or aggregate) carries a monotonically increasing version — an optimistic-locking counter, a sequence-backed change id, or the storage engine's own row version. The validator is derived from it, is readable in one indexed lookup, and is naturally correct across any rendering. This is the design that makes 304s cheap at origin. Whatever the source, the handler must **check the validator before assembling the payload**. A version column read after you have already loaded the aggregate saves nothing but bytes. ## Where the savings land Saved: response bytes on the wire (matters most on metered, mobile, or intercontinental links); client-side JSON parsing, object allocation and re-render; and, with strategies 2 and 3 plus an early short-circuit, origin database and CPU time. Not saved: the TCP/TLS connection, DNS and routing, authentication and authorization, rate-limit accounting, request logging — and above all the round trip. Conditional GET does not make a chatty API fast; it makes a fat API thin. If latency is your problem, a freshness lifetime that removes the request entirely, or a coarser endpoint that returns more per call, is the actual answer. ## Making the validator correct The validator must be a function of everything that can change the bytes you would send: - the resource's own version; - **embedded children** — an order that inlines its line items must change its tag when a line item changes, so you need an aggregate version or a max over the graph; - **request-shaped payloads** — `fields=`, `include=`, `expand=`, sort and page parameters all produce different bytes at what may be the same URL path; - **negotiated variants** — language, media type, API version header; pair these with `Vary`; - **the caller's authorization view** — if two callers with different scopes get different JSON, one validator per URL is a data leak waiting for a shared cache; - **your own serializer/schema version** — fold a build or schema id in, so a deploy that renames a field invalidates every stored copy instead of leaving clients on last week's shape. Weak versus strong matters here as a promise you can keep. A strong validator asserts byte-for-byte identity; a weak one (`W/` prefix) asserts semantic equivalence. JSON APIs rarely control their bytes — compression, key ordering, float formatting — so weak is usually the honest label. Strong tags are needed for byte-range requests, which most JSON APIs never serve. ## Collections are the hard case A single resource has an obvious version. A collection page must reflect inserts, updates, deletes and reordering: `max(updated_at)` silently misses deletes. Practical options are a tenant- or table-level change counter bumped on every write (cheap to read, over-invalidates, and that is usually acceptable), or a monotonic "last change id" per tenant. Often the better API answer is to skip conditional GET on collections entirely and offer a delta/since-cursor endpoint, which returns only what changed and beats an all-or-nothing 304. ## CDN interplay If a CDN or reverse proxy fronts the API, the economics shift. The edge holds the copy and revalidates with the origin, so one cheap origin revalidation can serve many clients and you save origin egress as well as client egress. But that only works for shareable responses: per-user payloads must be marked private, the edge drops out, and your 304 saves one client's bytes again. Conversely, for genuinely public content, a plain freshness lifetime (optionally with stale-while-revalidate) removes the request altogether and outperforms any conditional GET — reach for 304s when correctness requires the client to confirm each time. ## When to skip it Skip when payloads are a few KB after compression (the 304's own headers are not much smaller), when the resource changes on nearly every read so the validator almost never matches, when computing the validator costs as much as the answer, and for internal service-to-service calls on a fat network where the round trip dominates. Adding ETags "because REST" buys a new class of staleness bug — a validator that fails to change — for no measurable win.
- Would you emit a weak or a strong ETag for a JSON endpoint, and why?Weak, in almost every case. A strong validator promises byte-for-byte identity, and a JSON API does not control its bytes: gzip/brotli settings, map key ordering, float and timestamp formatting can all shift without the resource changing. Weak asserts only semantic equivalence, which is exactly what a version-derived tag can honestly guarantee. Strong validators are needed for byte-range requests, which JSON APIs generally do not serve.
- You derive the validator from an updated_at timestamp with one-second resolution. What can go wrong?Two writes inside the same second produce the same validator, so a client that fetched between them caches the older representation and is told it is still current — indefinitely, until the next write. Timestamp-based validators also miss deletes in collections and depend on clock behaviour across nodes. A monotonic version counter or a sequence-backed change id avoids all three problems.
- Your API renders different JSON for the same URL depending on the caller's scope. How does that change the validator design?The validator must identify the representation the caller actually receives, not the underlying row, so scope (or a hash of it) has to be part of the tag, and the response must be marked private or Vary declared on the negotiating header. Otherwise a shared cache can match one caller's validator against another caller's stored copy and hand back a view they are not entitled to.
It is the difference between a warehouse checking one revision number in its ledger to tell you nothing changed, and recounting every shelf first and then telling you nothing changed. Both spare you the truck; only one spares the warehouse.
saying these in an interview costs you the question
- Claiming that adding ETags automatically reduces server load — with regenerate-and-compare the origin does all its normal work and only the bytes are saved.
- Saying a 304 saves the round trip; it saves the payload, not the connection, auth, or latency.
- Hashing the fully rendered response body and describing that as a cheap validator.
- Emitting a strong ETag for JSON whose bytes are not stable under compression, key ordering or serializer changes.
- Deriving the tag only from the root row, so an embedded child change, an include/fields parameter, or a deploy that renames a field leaves clients pinned to a stale copy.
- Putting per-caller responses behind a shared cache without private/Vary, so one user's validator matches another user's view.