skip to content

Which paths populate a shared identifier-keyed data-access cache, and which reads never touch it?

level: middleimportance: should knowfreq 52%

answer

  1. only identified instances create entries
  2. a miss on load by key fills it
  3. queries fill it, are never served
  4. projections and raw statements leave nothing
  5. set-based writes void without refilling

basics

~20 s

A missed load by key populates a shared identifier-keyed cache, as does resolving a to-one link, and many layers store rows a query materialised. Projections, statements sent outside the mapper, and other processes populate nothing.

solid answer

~40 s

Population happens wherever the layer builds an identified row instance: a load by key that missed, a to-one link resolved through its stand-in, a deliberate warm-up load, and in many layers the rows a predicate query materialised. That last path is the asymmetric one — a filtered query can fill entries it can never be served from, and layers differ on whether they do it at all, so it is worth checking rather than assuming. Nothing is populated by a statement the application sends over the connection itself, because no row instance is built; nor by a projection of a few columns or into a transfer model, because there is no identifier to key an entry by; nor by reads issued by another process. A set-based write voids entries and puts nothing back.

go deeper

for a junior

Remember the rule that decides everything: an entry exists only where the layer built a row instance with an identifier. No instance, no entry.

for a middle

Explain the asymmetry — a predicate query can fill identifier entries yet can never be served from them — and name the bypass paths: raw statements, projections and writes issued outside the mapper.

for a senior

Show that you verify population empirically from the statement log rather than trusting configuration, and that you separate cold-start behaviour from steady state when you read hit rates.

for a principal

Judge whether a workload's read mix can ever hit at all before spending memory on it, and set the expectation that search-heavy traffic gains nothing from an identifier-keyed tier.

A shared identifier-keyed cache only helps when entries are actually in it, and the paths that put them there are narrower than most people assume. Population and service are also asymmetric: some reads fill the cache without ever being answerable from it, and some reads pass through the layer leaving no trace at all. ## The paths that populate it - **A load by key that missed.** The layer fetches the row, builds the instance and writes the snapshot into an entry. This is the main path. - **Resolving a to-one link through its stand-in.** Following the link is itself a load by key, so it consults the cache and, on a miss, populates it. - **Rows materialised by a query.** A predicate query cannot be *answered* from an identifier-keyed store, but every row it builds has an identifier, so the layer can store each one. Layers differ here: some write every materialised row of a cached type, some only for types configured for it, some only when the individual query asks for it. Do not assume it happens. - **A replace-on-write strategy.** A strategy that writes the new state into the entry at commit, instead of dropping it, is a population path — the next reader hits rather than misses. - **A deliberate warm-up.** Loading a known key set at start-up fills entries before real traffic arrives. ## What silently bypasses it - **Statements the application sends over the connection itself.** Rows come back as raw records; nothing is mapped and nothing carries an identifier the layer tracks, so nothing can be keyed. The cache never sees them. - **Projections.** A query selecting three columns, or selecting into a transfer model, produces no identified row instance. No identifier means no entry, even though the underlying rows belong to a cached type. - **Reads issued by anything else.** Another application, a scheduled job inside the database, an operator at a console — all touch rows the layer has no visibility into. - **A set-based write.** One statement changing many rows has no instance in hand; the layer can void the affected entries but has nothing to put back, so the following reads pay misses. How that voiding is ordered belongs with invalidation and with bulk writes. - **Anything filtered.** Worth restating because it surprises people most: the cache does not answer a `WHERE` clause. A workload that reads exclusively through search filters shows a hit rate near zero no matter how much memory you give it. ## Why the asymmetry matters Population through queries is what makes the cache useful in an application that searches and then navigates. The search runs, materialises rows and fills entries; the *navigation* that follows — opening one result, walking its links, reloading it on the next request — is then served from the cache. Where the layer does not populate from queries, the same workload keeps missing until each row happens to be loaded by key at least once. The opposite mistake is expecting the fill to make the search itself cheap. It never does: the statement runs every time, and only the by-key traffic behind it gets faster. ## Measuring instead of assuming 1. Count, per mapped type, how many reads arrive **by identifier** and how many arrive through a predicate. Only the first group can ever hit. 2. Instrument hits, misses and evictions per type rather than one number for the whole store; one hot type easily hides a dozen useless ones. 3. Read the statement log after a representative run. Loads by primary key that you expected to be hits still emitting `SELECT` mean the population path you assumed does not exist in this layer or is not switched on for that type. 4. Watch the first minutes after a deploy separately. A cold store turns every path into a miss, and an application tuned against a warm cache can fall over exactly then. ## What an interviewer is listening for - that you know a query may fill the cache yet can never be served by it - that raw statements and projections leave nothing behind, and why: no identified instance, no key - that you would confirm whether population from queries actually happens rather than assuming it - that you would measure hit rate per type before defending the cache at all

  • Why can a projection into a transfer model not create an entry?
    Because no identified row instance is built. The entry key is the mapped type plus the identifier, and a projection returns loose column values with nothing to key on. The underlying rows may well be of a cached type, but the layer never held an instance of one.
  • An application mostly searches, then opens one result. Where does the cache actually help?
    On the second half. Each search still runs a statement, but the rows it materialises can fill identifier entries; opening a result, walking its to-one links and re-reading it on the next request are all lookups by key, and those are the reads a hit removes.
  • How would you check whether your layer populates entries from query results?
    Run a query that materialises known rows, then load those rows by key in a fresh unit of work with the statement log on. If the loads emit no primary-key `SELECT`, population from queries happens; if they emit one each, it does not, and the by-key traffic is what has to warm the cache.

saying these in an interview costs you the question

  • Believes every read through the layer fills the cache
  • Expects a filtered query to be answered once its rows are cached
  • Thinks raw statements over the connection keep entries current
  • Assumes a projection caches the rows behind the selected columns
  • Reports one global hit rate and never looks per type