In a lookup service that caches 'not found' results, how should a negative entry be represented and expired?
answer
- Absent versus not cached
- Reserved marker, never empty
- Only a successful no-row query
- Seconds or minutes, not hours
- Create path replaces it
basics
~10 sStore a reserved tombstone value that readers can tell apart from a cache miss, write it only for a confirmed no-row result, and give it a short TTL so a later-created item appears quickly.
solid answer
~50 sA negative entry must be **distinguishable from a miss**. Many cache clients report a missing key as an empty or null result, so storing an empty value can make the reader treat the tombstone as a miss and query the database anyway, or fail when decoding it. Use a reserved sentinel or a typed envelope such as `{"state":"absent"}`. Write it **only for an authoritative not-found**, meaning a query that succeeded and found no row: never for a timeout or error, and never for a result that depends on who is asking unless the caller's scope is part of the key. Give it a **short TTL**, seconds to a few minutes against hours for positive entries, because the item may be created later and tombstones cost memory. The create path should also overwrite or clear the tombstone.
code
json · 12 lines[
{
"key": "item:48213",
"ttlSeconds": 3600,
"value": { "state": "present", "item": { "id": 48213, "title": "Desk lamp" } }
},
{
"key": "item:99999",
"ttlSeconds": 60,
"value": { "state": "absent" }
}
]go deeper
Remember that caching absence needs a special marker value and a short expiry, so repeat lookups for missing items stop reaching the database.
Explain why the marker must differ from a miss, which database outcomes may become tombstones, and why the negative TTL is shorter than the positive one.
Show the production traps: errors cached as absence, caller-dependent visibility under a shared key, and the create path forgetting to replace the tombstone.
Discuss how you would set negative-TTL policy across many lookup services, weighing staleness for newly created items against database load and cache memory.
## What a negative entry is In a cache-fronted lookup service, a request for an ID with no database row normally stores nothing, so every repeat goes back to the database. A **negative entry** (also called a **tombstone** or **null sentinel**) fixes that by caching the fact of absence: "`item:99999` has no row". The next request for the same key becomes a hit that returns not-found without touching the database. The idea is simple. Most of the design work is in three details: how absence is **represented**, which results are allowed to **become** tombstones, and how long they **live**. ## Representing absence unambiguously The reader has to tell three states apart: *present* (a value), *absent* (a tombstone) and *unknown* (a real miss). Many cache clients report a miss as an empty or null result, so the encoding matters. | Encoding | Problem | |---|---| | Empty string or null value | Often indistinguishable from a miss, so the reader queries the database anyway; some serialisers fail on it | | Reserved sentinel string | Works, as long as no legitimate value can ever equal the sentinel | | Typed envelope (`state: present` / `state: absent`) | Clearest; every entry says what it is, and adding states later is easy | Whichever you choose, the read path must check for the sentinel **before** treating the value as a row. A tombstone that leaks into business logic as a real object is a bug that is hard to trace. ## What is allowed to become a tombstone A tombstone asserts that the item does not exist. Only cache that when you actually know it: - **A query that completed and found no row**: cache it. - **A timeout, connection error or overload rejection**: do not cache it. The row may well exist, and caching absence would turn a short database problem into not-found responses for the whole negative TTL. - **A row that exists but is hidden from this caller**: do not store a shared tombstone. Either cache the caller-independent fact and apply permissions after the cache, or include the permission scope in the key. - **Soft-deleted rows**: decide explicitly whether they count as absent for readers, and cache whatever the API would return. ## Choosing the TTL Negative entries normally get a **much shorter TTL** than positive ones: for illustration, 30 to 300 seconds against an hour. Three forces push it down: - **Staleness in the dangerous direction.** If the item is created after the tombstone is written, readers see not-found until the tombstone expires or is replaced. The negative TTL is the worst case when invalidation fails. - **Memory.** A tombstone holds no row, but its key, the sentinel and the cache's per-entry overhead still take space, and absence can be requested for an unbounded number of keys. - **Low value per entry.** A not-found query on an indexed key is usually cheap to repeat, so a long tombstone saves little. Against that, a TTL that is too short stops absorbing repeats. A reasonable target is the shortest TTL that still turns most repeat requests for the same missing key into hits. Measure it: log how soon repeat requests for a missing key arrive, and pick a TTL that covers most of those gaps. If repeats mostly arrive within seconds, a few seconds of absence is enough. ## Keeping the create path honest The TTL is only the backstop. When an item is created, the write path should **overwrite the key with the new value** (or at least delete it) after the database commit, so a tombstone does not hide a real item for the rest of its life. Races between that write and a slow reader need extra care, but the principle is the same: absence is a cached claim, and the code that makes it false must correct it. ## Checklist 1. Encode absence with a sentinel or envelope the reader cannot mistake for a miss. 2. Write it only after a successful query that returned no row. 3. Keep caller-dependent results out of shared tombstones. 4. Give it a TTL far shorter than positive entries. 5. Overwrite or clear it when the item is created. Following these steps turns repeat requests for missing keys into cheap hits without hiding real items for long or letting errors pass as absence.
- Should a negative entry be written when the database query times out?No. A timeout says nothing about whether the row exists. Caching it as absence would hide a real item for the whole negative TTL and turn a short database problem into not-found responses. Only a query that completed and returned no row is reliable enough to cache.
- How does the key design change when visibility depends on the caller?If a row can exist but be hidden from some callers, a shared tombstone would tell authorised callers that the item does not exist. Either cache only the caller-independent fact and apply permissions after the cache, or put the permission scope in the key. Caching 'not visible to this user' under a shared key is a correctness bug.
saying these in an interview costs you the question
- An empty string or null is a fine way to cache 'not found'
- Cache a database timeout as not-found to protect the database
- Negative entries should live as long as positive ones for consistency
- Tombstones cost nothing because they hold no data
- A shared tombstone is fine even when visibility depends on the caller