skip to content

How do you decide what belongs in session attributes rather than being re-derived on each request?

level: principalimportance: should knowfreq 38%

answer

  1. every attribute costs on every request
  2. re-derivable from identity means re-derive
  3. a cached copy nothing invalidates
  4. two builds read one stored shape
  5. additive-only, versioned, or primitives only

basics

~20 s

Store only what cannot be re-derived and must survive between requests: small, stable, self-contained values. Everything derivable is cheaper to re-fetch, because each attribute costs serialization on every save and couples two deployed versions to one shape.

solid answer

~40 s

Three costs decide it. Every attribute is serialized and deserialized around the requests that touch the session, so size shows up as latency on a remote store. Every attribute is a cached copy, so it can go stale against the source of truth and there is no invalidation. And every attribute's shape becomes a contract between code versions, because during a rolling deploy the old and new builds read each other's entries. That leaves a narrow rule: keep identifiers and short workflow markers, keep derivable data out, and where a value is expensive to re-derive, cache it outside the session where it can be invalidated and where its shape is versioned deliberately.

go deeper

for a junior

Default to storing as little as possible: who the client is, and a marker or two. Anything you can look up again from that identity should be looked up again rather than kept.

for a middle

Explain the per-request serialization cost and the staleness of a copied value, and give the re-derivability test you would apply before adding any new attribute.

for a senior

Bring the deployment angle: two builds read one stored shape, so attribute changes are add-then-migrate-then-remove, and an unreadable entry presents as users being signed out.

for a principal

Set the policy — a size budget, an additive-only or versioned-shape rule, and where derived data is cached instead — and be explicit that a staleness window on cached permissions is a product decision.

## Three costs, paid per attribute A session attribute looks free in a handler. It is not, and naming the costs is most of the answer. 1. **Serialization cost.** Attributes are turned into a stored form on save and back on load. On a remote store that is a round trip whose size you chose; on a large attribute it is also work on every request that touches the session, and that work lands in request latency. 2. **Staleness cost.** An attribute is a cached copy of something the system knows better elsewhere. Nothing invalidates it when the source changes, so a value copied in at sign-in can be wrong for as long as the entry lives. 3. **Shape-coupling cost.** The stored form is written by one build and read by another. That makes the shape of every attribute a compatibility contract, and the contract is enforced across a deployment window rather than at compile time. ## What belongs, and what does not | Keep in the session | Keep out | |---|---| | The identity reference — who this client is | Whole records loaded from a data store | | Short workflow markers — step reached, accepted terms | Anything derivable from the identity plus a query | | A tiny pending outcome for the next request | Large collections, search results, rendered fragments | | A per-session preference not worth persisting | Values that must be correct rather than merely recent | The test that resolves most cases is: **can this be re-derived from the identity plus durable data?** If yes, re-derive it. A query answered in a millisecond is almost always cheaper than an attribute that inflates every save, goes stale invisibly, and pins a shape across deploys. If a re-derivation is genuinely expensive, that is an argument for a cache with an invalidation story — not for the session, which has no invalidation story at all. ## The rolling-deploy problem During a rolling deploy two builds serve at once, and a client's requests are spread across both. An entry written by the new build will be read by the old one and vice versa, so: - **Adding a field** is safe only if both builds tolerate its absence. - **Removing a field** is safe only once no live build requires it. - **Renaming or retyping a field** is never a single-step change; it is add, dual-write, migrate readers, remove. - **A stored form tied to a declared type** is stricter still: a shape that no longer matches can fail to deserialize, and a failure to deserialize an entry is a signed-in user thrown back to an unauthenticated state. Three viable strategies, in increasing order of discipline: 1. **Additive-only.** Never remove or retype; tolerate unknown and missing fields on read. Cheapest, and it accumulates dead fields. 2. **Versioned entries.** Carry a version marker in the entry and have readers upgrade older shapes in place. Explicit, and it costs an upgrade path per change. 3. **Store primitives only.** Keep identifiers and short markers, so there is nearly no shape to break. The most robust, and it works precisely because nothing large lives there. The failure mode that motivates all three is worth stating plainly: a deploy that changes an attribute's shape without a strategy signs out everyone whose entry cannot be read, which looks like an outage and is rarely traced back to the deploy. ## Judgment calls with no single right answer - **Cached permissions or profile fields.** They remove a lookup per request and guarantee a window in which a change has not taken effect. Whether the window is acceptable is a product decision, not a technical one. - **Where the derived cache lives.** Outside the session it can be shared, sized and invalidated; inside it is per client and invisible. The first is more machinery, the second is more surprise. - **Attribute size budget.** Setting one — a few hundred bytes, say — makes the tradeoff explicit and catches the multi-record attribute in review rather than in a latency graph. - **Whether to keep a session at all** for a given surface. Some flows carry everything they need per request and need no entry; that is a design position to take deliberately. ## What a strong answer sounds like It starts from the costs rather than from a list of allowed values, it names the rolling deploy as the reason shapes matter, and it ends with a rule the team can apply without asking: identifiers and markers in, derivable data out, anything else justified in review.

  • What actually breaks when a deploy changes the shape of a stored session attribute?
    Two builds serve at once, so entries written by one are read by the other. A reader that cannot interpret the stored shape either sees a missing value or fails outright, and a failure usually presents as the client being treated as brand new — everyone signed out mid-deploy, with the cause hidden one release behind.
  • Why is caching an expensive lookup in the session worse than caching it elsewhere?
    Because the session offers no invalidation, no sharing and no size control. The copy is per client, lives as long as the entry, and cannot be cleared when the source changes. A cache outside it can be keyed, sized, invalidated and shared across clients, at the cost of being explicit machinery.
  • How would you enforce a size budget on session attributes in practice?
    Measure the stored form on save and record it, alert on the upper percentiles rather than the mean, and fail loudly in non-production when an entry exceeds the budget. Pair that with a review rule that any new attribute states why it cannot be re-derived.

saying these in an interview costs you the question

  • Treats the session as free storage because it is just a map
  • Caches a full user record there and never invalidates it
  • Changes an attribute's stored shape in one deploy with no migration
  • Assumes only one build reads the store during a release
  • Argues size does not matter because the store is fast