You own a large Next.js App Router site whose routes are all cached. For each kind of content, how do you decide between a time-based revalidation window and event-driven invalidation, and how do you choose how fine-grained your cache tags should be?
answer
- budget first, mechanism second
- cost by traffic or cost by change
- fast path plus backstop
- tags are an interface, not decoration
- both extremes of granularity fail
basics
~20 sDecide per content type from an explicit staleness budget: windows where being minutes old is acceptable and no change event exists, event-driven invalidation where a change must be visible immediately or a reliable publish event already exists. Size tags to real change events.
solid answer
~60 sI start from a staleness budget agreed with the business per content type, not from a technical default: how wrong may this be, for how long, before it costs us? Content with a tolerant budget and no trustworthy change event gets a time-based window, because it needs no integration and its cost is bounded by traffic. Content with a tight budget, or that already emits a publish or mutation event, gets on-demand invalidation — cost then scales with the rate of change rather than the rate of reads, which is the right axis for read-heavy content. In practice almost everything gets both: the event as the fast path, a generous window as the backstop for events that get lost. Tag granularity follows the shape of the change events: one tag per entity so a single edit refreshes only what read it, plus collection tags for genuinely collection-wide changes. Too coarse and one edit re-renders the site; too fine and some page silently never refreshes because nobody attached the tag.
go deeper
Know that some content can be a few minutes out of date and some cannot, and that Next.js offers both a timed window and an explicit invalidation call to match those cases.
Be able to say which mechanism you would pick for a given page and why, and to explain that a window costs work proportional to traffic while invalidation costs work proportional to changes.
Argue for running both together — event as the fast path, window as the backstop — and describe the monitoring that makes a silently dead invalidation path visible before users report it.
Own the staleness budget as a negotiated product input, the tag taxonomy as an interface other teams build against, and the platform consequences such as shared caching and bulk-import behaviour that your policy commits the organisation to.
## Start from a budget, not a default The failure mode in large codebases is that every route inherits whatever number the first engineer typed. The useful discipline is to write down, per content type, how stale it may be before someone is misled or annoyed. Legal and pricing copy: near zero. Editorial articles: minutes. Marketing pages: hours. A help centre: a day. That budget is a product decision, and once it exists the technical choice mostly falls out of it. ## The two cost curves The choice between window and event is a choice between two scaling axes. A **time-based window** makes revalidation cost proportional to *traffic*, capped at one regeneration per route per window. It requires no integration and no cooperation from anyone else, and it degrades gracefully — nothing breaks if a system upstream goes quiet. Its ceiling is that it can never be immediate, and pushing the number down to chase freshness multiplies regenerations on exactly the hottest routes. **Event-driven invalidation** makes cost proportional to the *rate of change*. For read-heavy content that changes rarely, that is dramatically cheaper and also immediate. Its price is a dependency: someone must reliably emit the event, the endpoint receiving it must be authenticated and available, and the invalidation must reach every instance serving the app. Every one of those is a way to silently stop refreshing. So the rule of thumb: **windows are cheap to build and expensive to make fresh; events are cheap to make fresh and expensive to make reliable.** Choose by which cost you can afford, and then run both — event as fast path, window as backstop — wherever the content matters. The backstop converts "stale forever because a webhook died" into "stale for at most the window", which is the difference between an incident and a nuisance. ## Designing the tag taxonomy Tags are the interface between the systems that change data and the pages that read it, and they deserve to be designed once rather than invented per feature. - **Entity tags** — `post-42`, `author-7`. The default. They match the granularity of real edits: one document changed, so the pages that read that document refresh. - **Collection tags** — `posts`, `navigation`. For changes whose blast radius genuinely is the whole collection: a new item appearing in a list, a reordering, a global navigation change. - **Avoid a single global tag.** A `content` tag on everything makes each trivial edit invalidate the whole site. Every route then re-renders on its next request, hitting your data sources with a cold-cache burst — a self-inflicted thundering herd triggered by a typo fix. - **Avoid tags nobody invalidates.** The mirror failure is tag proliferation: labels attached at read sites that no change event ever names. They cost nothing at runtime but create false confidence that a page is event-refreshed when only the window is actually keeping it honest. The practical enforcement is that tag strings are derived from one shared module (`postTag(id)`, `collectionTag('posts')`) that both the read path and the invalidation path import. Tags are plain strings with no type checking and no registry; a mismatched literal produces no error anywhere, just a page that quietly never updates. Centralising the construction is the only thing that makes a typo a compile-time problem instead of a support ticket six weeks later. ## Operational commitments that come with the policy A policy is only real if it is observable. At minimum: log every invalidation with its tag and source, alert when a content type's invalidation rate drops to zero for longer than it plausibly should, and track how often regenerations fail. On the deployment side, on-demand invalidation is only correct across multiple instances if they share a cache — that is a platform requirement created by the policy, not an implementation detail, and it belongs in the decision rather than being discovered later. Bulk operations deserve an explicit answer too. A content migration that fires ten thousand webhooks will invalidate a collection tag ten thousand times and then re-render everything at once on the next traffic. Deciding up front whether such imports coalesce, or bypass the webhook and rely on the window, is the difference between a routine migration and a paged incident. ## How to present it A strong answer names the budget as the input, names the two cost curves as the reason for choosing, commits to running both mechanisms together for anything important, and then treats tag granularity as an interface design problem with a failure mode at each extreme. Weak answers pick one mechanism as universally better, or describe tags as free labelling with no cost to getting the size wrong.
- If event-driven invalidation is immediate and cheaper for read-heavy content, why keep a time-based window at all?Because events are a dependency that can fail silently — a dropped webhook, a mid-deploy 500, a record edited straight in the database. Without a window, any of those means stale forever with no signal. With one, the worst case is bounded by the window. The window is not there for freshness; it is there so a broken fast path degrades instead of breaking.
- How would you catch a tag that pages read but no change event ever invalidates?Construct tag strings from one shared module used by both readers and invalidators, so the set of tags in use is enumerable rather than scattered through string literals. Then compare tags attached at read time against tags actually invalidated in production logs over a period; anything read but never invalidated is being kept fresh only by its window, which may or may not be intentional.
- What is the operational risk when a bulk content import fires one webhook per record?Two compounding effects: thousands of invalidations of the same collection tag, then a cold-cache burst as every affected route re-renders on its next request, all hitting the data sources at once. Decide in advance whether imports coalesce into a single broad invalidation, or skip the webhook entirely and let the window absorb the change more gradually.
saying these in an interview costs you the question
- Declares one mechanism universally better than the other
- Picks revalidation numbers with no staleness budget behind them
- Treats tag granularity as arbitrary labelling with no cost
- Uses one global tag so every edit invalidates everything
- Ignores that events fail silently and need monitoring