For an interactive search box served by scatter-gather over many shards, how would you decide whether to return partial results when some shards miss the deadline?
answer
- latency promise vs completeness promise
- who consumes the answer
- random slices vs routed shards
- chance a top-10 hit is lost
- flag, floor, don't cache
basics
~20 sReturn partial results only if the missing shards hold little of the likely answer and the consumer can tolerate gaps. Flag the response as partial, keep it out of the result cache, and fail instead where completeness is the requirement.
solid answer
~50 sThe decision trades a latency promise against a completeness promise. Base it on three things. First, the **consumer**: a person scanning a ranked list rarely notices one lost result, but exact counts, exports, compliance searches and pagination do. Second, the **data layout**: when documents sit on shards at random, 2 missing shards out of 100 leave the top 10 untouched about 82% of the time (0.98^10). When shards are routed by tenant or time, or tiered by quality, one missing shard can erase a whole slice of the answer. Third, the **guardrails**: mark the response `partial` with shard counts, set a minimum coverage below which the request fails, never cache a partial answer as if it were complete, alert on the partial rate, and treat replicas and hedging as the first defence so partial answers stay rare.
go deeper
Recall that a deadline lets a search answer from the shards that replied, and that the response should say it is partial.
Explain how to estimate what a missing shard costs when documents are spread randomly, and why tenant or time routing makes a missing shard much more damaging.
Show the operational guardrails: a coverage floor, a partial flag, keeping partial answers out of caches, monitoring the partial rate, and retrying only the missing shards within the deadline.
Own the policy per endpoint: which surfaces trade completeness for latency, where the coverage floor sits, and how layout choices such as tiering change that trade.
## The trade being made A scatter-gather query that waits for every shard is as slow as its slowest shard. A **deadline** caps that wait. When it fires, the coordinator can fail the request, or it can answer from the shards that did reply, which is a **partial result**. No choice is right everywhere. The answer depends on what the result is used for and how documents are laid out across shards. Because this trades a product promise (complete answers) against an SLO (fast answers), and the right line differs by endpoint, it is a call leads are expected to own. ## What a missing shard actually costs If documents are assigned to shards at random, each document in the true global top 10 sits on a missing shard with probability equal to the fraction of shards missing. With 2 of 100 shards absent: - each top-10 document is lost with probability 2%; - the expected number of top-10 documents lost is 10 × 0.02 = 0.2; - the probability that the top 10 is completely unaffected is 0.98^10 ≈ 82%. When a document is lost, the next-best one takes its place, and a person scanning the list rarely notices. That statistical argument breaks down when shards are **not** random slices: - **Tenant routing.** Documents placed by tenant mean a missing shard drops some tenants' results entirely. - **Time placement.** Documents placed by time mean a missing shard can remove a whole week. - **Quality tiers.** In a tiered index where one small shard holds the best documents, losing that shard costs far more than 1/N of the answer. ## When partial is acceptable | Use case | Partial acceptable? | Why | |---|---|---| | Web-style keyword search box | usually yes | people feel latency; a missing tail result rarely registers | | Typeahead suggestions | yes | incomplete suggestions do little harm and the next keystroke retries | | Exact counts or facets used for billing or reports | no | wrong numbers look authoritative | | Compliance or legal discovery search | no | completeness is the requirement | | Exports, or paging through many pages | no | missing shards leave silent gaps | | A no-results page | with caution | an empty partial answer may be a false negative | ## Guardrails when partial answers are allowed - **Mark the response.** Include a `partial` flag with shards answered and shards total, so clients, logs and dashboards can tell. - **Minimum coverage.** Fail the request if fewer than a set share of shards replied, for example 95%. - **Keep partial answers out of the result cache.** A cache would keep serving the degraded answer after the shards recover. Cache only complete responses, or give partial ones a very short lifetime and a separate key. - **Monitor the partial rate** as an SLO signal. A steady trickle usually points to a capacity problem or a hot shard. - **Use replicas, replica selection and hedging first**, so partial answers are the last line of defence. - **Retry narrowly.** Re-running the whole fan-out for one missing shard doubles the load; retry only the missing shards, and only if the deadline allows. ## A decision procedure 1. **Classify the consumer.** Is it a person reading a ranked list, or a program that needs every match? 2. **Check the layout.** Are shards random slices, or are they routed by tenant, time or quality tier? 3. **Set the deadline** from the page's latency budget, minus merge and fetch time. 4. **Pick a coverage threshold** and the behaviour below it: return an error, retry the missing shards, or degrade openly. 5. **Decide the caching rule** for partial responses. 6. **Instrument** coverage and partial rate, and review both after every incident. The result is usually a policy that differs per endpoint. The interactive search box answers at the deadline with a coverage floor. Counting, export and compliance endpoints wait longer or fail loudly. The result cache accepts only complete responses.
- In a scatter-gather search service, why should partial responses stay out of the result cache?A result cache serves the stored answer to everyone who sends the same query until the entry expires. If a partial answer is cached, the degraded result keeps being served after the missing shards recover, and nothing marks it as incomplete. Cache only responses with full coverage. If partial answers must be cached for load reasons, give them a very short lifetime and a key that records their coverage.
- Which signals tell you that partial search results have become a real problem rather than rare noise?Track the share of responses flagged partial and the average shard coverage per query, broken down by endpoint and by shard. A rising partial rate across many shards points to capacity or overload. A partial rate concentrated on one shard points to a hot or unhealthy shard. Set an SLO on coverage alongside latency so the deadline cannot quietly trade away completeness.
- How does a tiered search index change the partial-results decision?In a tiered index, a small tier holds the documents most likely to rank highly and a larger tier holds the long tail. Missing a long-tail shard costs little, because its documents rarely reach the top 10. Missing the top tier can gut the answer. So you can allow partial answers when only long-tail shards are late, but require the top tier to be complete, or give it more replicas and a longer deadline.
saying these in an interview costs you the question
- Partial results are always wrong; a search must wait for every shard.
- A missing shard never matters because results are spread evenly.
- A partial result can be cached like any other response.
- Clients do not need to know that a response was partial.
- Raising the deadline until no shard ever misses it costs nothing.