Job postings expire hourly but the posting-vector index rebuilds nightly — how does the retrieval stage stay fresh?
answer
- arrivals and departures differ
- small fresh segment beside the base
- deletes are lazy in the structure
- drop tombstones on the read path
- dead entries eat result slots
basics
~20 sSplit additions from removals. New postings are embedded on arrival into a small fresh segment searched beside the nightly-built base index; filled or expired postings are dropped on the read path from a tombstone set, because the structure only reclaims them at rebuild.
solid answer
~50 sIndex freshness has two halves with different urgency. **Additions** matter because a posting's best day is its first: an event-driven path embeds each new posting as it is ingested and upserts it into a small fresh segment, which the retrieval stage searches alongside the base index and merges. **Removals** matter more, because showing a filled role is worse than missing a live one: the retrieval stage consults a tombstone set on every request and drops those candidates before the scorer sees them, since index structures differ in whether a delete is a physical removal or just a mark, and the space is usually only reclaimed by the nightly rebuild. Tombstones still consume result slots, so the effective depth falls as they accumulate — watch the tombstone ratio and trigger an off-cycle rebuild when it crosses a threshold.
code
pseudocode · 15 linesfunction retrieve(seeker_vector, depth):
base = base_index.search(seeker_vector, depth) // rebuilt nightly
fresh = fresh_segment.search(seeker_vector, depth) // upserted on posting events
merged = merge_by_similarity(base, fresh)
live = empty list
for each posting in merged:
if posting.id in tombstones: continue // filled or withdrawn
if posting.expires_at <= now: continue // aged out since the build
live.append(posting)
if size(live) == depth: break
emit_metric("effective_depth", size(live) / depth)
return livego deeper
Recall that an index built on a schedule does not know about postings created or closed since the build, and that the retrieval stage has to handle both cases itself.
Explain the two-segment shape: a base index rebuilt on a schedule plus a small fresh segment receiving upserts, searched together, and a tombstone check applied to the merged results.
Show that removal must be enforced on the read path because the structure reclaims entries lazily, and that accumulated tombstones silently reduce effective depth — then name the metrics that expose it and the trigger for an off-cycle rebuild.
Decide whether the freshness the product needs justifies the two-segment complexity at all: with a small catalogue a faster rebuild is simpler, and complexity should follow the rebuild cost, not habit.
## What actually goes stale "The index is stale" hides three separate clocks. Keep them apart: - **Index freshness** — whether the postings currently open are the postings in the index. This is what a nightly build costs you. - **Vector staleness within a posting** — a posting whose text was edited after it was embedded. Usually tolerable, and repaired by re-embedding on the edit event. - **Encoder version** — what the vectors mean at all. A different clock entirely, and a much sharper failure. This leaf's operational problem is the first one, in both directions: postings arriving and postings leaving. ## Additions: a fresh segment beside the base A nightly build gives a posting an invisible window of up to twenty-four hours, which on a job board is the window where it gets most of its applications. The standard shape is two segments searched together: - the **base index**, rebuilt on a schedule over the whole catalogue, compacted and well-formed; - a **fresh segment**, small, receiving upserts from an event-driven embedding path as postings are ingested, and searched on every request beside the base. The retrieval stage queries both and merges by similarity. Because the fresh segment holds hours rather than years of inventory, it stays small enough that searching it adds little latency; at the next rebuild its contents are folded into the base and it starts empty again. The price is a second search per request and a merge, which is a fixed and modest cost. ## Removals: enforce on the read path Removal is the asymmetric case. A seeker who applies to a role that was filled last week has a worse experience than a seeker who never saw it, and the damage lands on the employer too. But most index structures make deletion lazy — a delete marks an entry and the structure reclaims it at compaction — and structures differ in exactly how. The tier therefore cannot wait for the index to forget: 1. The posting lifecycle publishes a close, withdraw or expiry event. 2. That id enters a **tombstone set** the retrieval stage can check in constant time. 3. Every candidate coming out of either segment is checked against the tombstone set and its own expiry time, and dropped before it reaches the scorer. 4. The next rebuild removes the vector physically and the tombstone entry retires with it. The rule to state plainly: **deletion is enforced on the read path; the index catches up later.** ## Tombstones cost depth, not correctness A dead entry still occupies a slot in the search's result set, so as tombstones accumulate the stage returns fewer live candidates than the depth it asked for. If a fifth of every result set is dead, a search at depth 900 yields roughly 720 live candidates before eligibility filters take their share — which is a recall loss no error will announce. Two responses: - raise depth to compensate, which costs latency and merge work; - treat the **tombstone ratio** as a build trigger: past a threshold, rebuild off-cycle rather than keep paying. ## What to monitor - **Invisible window**, at p95: elapsed time from posting created to first retrievable. This is the number the business cares about. - **Tombstone ratio** in returned result sets, and time since the last full rebuild. - **Effective depth**: live candidates surviving the tombstone check, against the depth requested. - **Fresh-segment size and lag**: a stalled embedding path shows up as a segment that stops growing, which looks like nothing at all on request-level metrics. ## The trade being made A more frequent full rebuild is the simplest answer and is sometimes right: if the catalogue is small enough that a rebuild takes minutes, the two-segment design is unnecessary complexity. The segment split earns its keep when the rebuild is expensive relative to the freshness the product needs — which on a multi-million-posting board with hourly turnover it almost always is.
- Why not simply rebuild the posting index every fifteen minutes and drop the fresh segment?Because a rebuild over millions of vectors costs real compute and produces a whole new structure to validate and swap, so doing it every fifteen minutes spends heavily to refresh inventory that mostly did not change. The segment split localises the churn: a small structure absorbs the day's arrivals while the expensive build stays on its own schedule.
- What happens if the event-driven embedding path stalls for six hours?Requests keep succeeding and latency is unchanged, because the base index still answers. The only symptom is that postings from the last six hours are unreachable, so the invisible-window metric and fresh-segment lag are what detect it; nothing in the request path will.
saying these in an interview costs you the question
- Deleting a posting removes its vector from the index immediately
- The scoring stage will rank closed postings low anyway
- Tombstones are harmless until the next rebuild
- A fresh segment exists so rebuilds can run without downtime
- Index freshness and encoder version are the same staleness