skip to content

questions

5

For a product catalog's search index, what is the difference between a batch rebuild and near-real-time incremental indexing?

level: juniorimportance: must knowfreq 60%

answer

  1. schedule versus stream
  2. cost follows size or change rate
  3. clean slate versus accumulating drift
  4. hybrid: hot fields streamed

basics

~20 s

A batch rebuild periodically re-reads the whole catalog and builds a fresh index, so results lag by up to the rebuild interval. Near-real-time indexing applies each change as it commits, giving seconds of lag but needing an always-running pipeline.

solid answer

~40 s

A **batch rebuild** is a scheduled job that exports the whole catalog, transforms it, builds a complete new index and switches searches to it. It is simple and starts from a clean state every run, but a change waits until the next run, and the cost grows with catalog size rather than with how much changed. **Near-real-time indexing** captures each committed change and applies it to the live index as an upsert or delete, so lag is usually seconds, and cost follows the change rate. The price is a streaming pipeline that must handle ordering, duplicates and deletes, and that can slowly drift. Most catalogs combine them: near-real-time updates for price, stock and delisting, plus a periodic full rebuild that resets drift and picks up schema changes.

go deeper

for a junior

Know the two strategies by their lag: a rebuild on a schedule versus changes applied as they happen, and name one field that needs the faster path.

for a middle

Explain what each costs, what drives that cost, and why incremental pipelines drift while rebuilds reset it; mention the engine visibility delay.

for a senior

Describe a hybrid you would run: which fields stream, how often the full rebuild runs, and how the rebuild avoids losing changes that arrive while it builds.

for a principal

Frame the choice as a staleness budget per field against pipeline cost and on-call burden, and decide whether teams share one change stream or build their own.

## Two ways to keep a search index current A product-search index is a copy of catalog data reshaped for retrieval: text fields prepared for matching, filters for category and brand, and numeric fields for price and ranking. The **source of truth** is the transactional database, so the index is always a derived copy that trails it. The two broad strategies differ in how the copy is refreshed and therefore in how far behind it can fall. That distance is called **index lag** or staleness. ## Batch rebuild A **batch rebuild** recreates the index from scratch on a schedule. - A job reads a consistent snapshot of all products from the database or from an export. - It transforms and enriches each record, joining categories, prices, ratings and computed popularity. - It builds a complete new index, then switches readers to it and discards the old one. - **Staleness** is up to one schedule interval plus the build time: with a nightly run, a change made just after the export waits roughly a day. - Every run starts clean, so drift, missed deletes and half-applied updates vanish. - Cost scales with **catalog size**, not change volume. Illustratively, a catalog of 50 million products where 2 million change per day still rewrites all 50 million documents each night. ## Near-real-time incremental indexing **Near-real-time (NRT)** indexing keeps one live index and applies changes as they happen. - A pipeline captures each committed change, usually through an outbox or the database's commit log, and puts it on a queue. - An indexer turns each event into an upsert or delete on the live index. - Lag is typically seconds: capture delay, queue wait, indexing time, and the engine's **visibility delay**. Many engines buffer newly written documents and expose them to searches after a short, configurable interval, often around a second, which is why this is called *near* real time. - Cost scales with the **change rate**: in the example above, 2 million writes a day instead of 50 million. - The pipeline must handle ordering, duplicate delivery, retries and deletes. Mistakes accumulate as silent drift. ## Side by side | Aspect | Batch rebuild | Near-real-time | |---|---|---| | Freshness | Hours to a day | Seconds | | Cost driver | Catalog size | Change rate | | Moving parts | A scheduled job | Change capture, queue, indexer running all the time | | Drift | Reset every run | Accumulates unless reconciled | | Schema change | Natural: the next build uses the new schema | Needs a separate full reindex | | Failure mode | A failed run keeps yesterday's index | A stalled pipeline lets lag grow silently | ## Why most catalogs use both Pure batch is too stale for fields shoppers notice: an item shown as in stock after it sold out, or an old price. Pure NRT never gets a clean slate. A common **hybrid** is: - NRT updates for price, stock, availability and delisting, where staleness costs money or trust; - batch jobs for expensive derived signals such as sales rank, recomputed on a schedule; - a periodic or on-demand **full rebuild** into a new index that resets drift and applies schema changes, while replaying the NRT changes that arrive during the build so none are lost. ## Choosing between them 1. Decide how stale each field may be before users or the business are hurt. 2. Compare the daily change rate with the catalog size. 3. Estimate the cost of a wrong result: an oversold item, a wrong price, a delisted product still shown. 4. Check whether the team can operate and monitor a streaming pipeline. A small catalog that changes weekly is well served by batch alone. A marketplace with constant price and stock changes needs the NRT path and keeps the rebuild as its safety net.

  • Why is it called near-real-time rather than real-time indexing?
    Even after a change reaches the index, many engines do not expose new documents immediately. They buffer writes and make them searchable after a short, configurable interval, often around a second, because exposing every write instantly is expensive. Add capture delay and queue wait, and the end-to-end lag is seconds rather than zero.
  • If you already have near-real-time updates, why keep a full rebuild at all?
    Incremental pipelines drift: a dropped event, an indexer bug or a rejected document leaves a stale entry that nothing revisits. Some changes, such as a new field type or a different text analysis, cannot be applied to existing documents in place and need every document rewritten. A periodic or on-demand rebuild resets the index to exactly what the source of truth says.

A batch rebuild is reprinting the whole store catalogue every night; near-real-time indexing is updating individual shelf labels the moment a price changes.

saying these in an interview costs you the question

  • Near-real-time indexing makes a change searchable the instant it commits.
  • A batch rebuild only costs as much as the number of changed products.
  • Once near-real-time updates exist, a full rebuild is never needed again.
  • Batch rebuilds are obsolete and never the right choice for a catalog.
  • Incremental indexing cannot drift because every change is applied.
open as a page

Why do dual writes from application code to both a product database and its search index drift out of sync?

level: middleimportance: must knowfreq 64%

basics

~20 s

The two writes share no transaction: a crash or timeout between them leaves one side stale, and concurrent writers can reach the index out of order. Deriving index updates from committed database changes, through an outbox or change capture, closes the gap.

open as a page

In a search-indexing pipeline consuming product change events, how do you stop a stale event from overwriting a newer document?

level: middleimportance: should knowfreq 46%

basics

~20 s

Give every event a monotonically increasing version from the source of truth, such as a row version or commit-log position, and write to the index only if that version is newer than the stored one. Partitioning events by product ID keeps each product's changes in order.

open as a page

How would you define and measure an index-lag SLO for a product-search index fed from a transactional database?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Define lag as the time from a change committing in the database to it appearing in search results, and target a percentile, such as 99% of changes within 30 seconds. Measure it with commit timestamps carried through the pipeline plus synthetic canary products.

open as a page

When a product-search index is fully rebuilt into a shadow index, how do you keep the writes that arrive during the rebuild?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Record the change stream's position before the bulk load starts, load a snapshot into the shadow index, then replay every change from that position with version-guarded writes until it catches up. Only then switch the alias; keep the old index for rollback.

open as a page