skip to content

Your LlamaIndex query engine misses its p95 latency budget — which knobs do you trade?

level: principalimportance: should knowfreq 40%

answer

  1. measure the halves before touching either
  2. candidates cost more than survivors
  3. count the LLM calls the mode implies
  4. first token is not total time
  5. decide in advance what you shed

basics

~20 s

Attribute the time first: measure retriever.retrieve() alone against the full query() to see whether retrieval, postprocessing or synthesis dominates. Then trade deliberately — candidate pool size, reranker placement, response mode and streaming each buy latency back at a different quality cost.

solid answer

~50 s

Start with attribution, not with knobs. Time `retriever.retrieve(q)` on its own and subtract it from `engine.query(q)`; that splits retrieval from postprocessing plus synthesis, and the split tells you which half to touch. Retrieval is usually milliseconds, so the time is almost always in a cross-encoder reranker scoring a large candidate pool, or in the synthesizer making more LLM calls than you assumed. From there the trades are: lower `similarity_top_k`, which cuts reranker work proportionally but costs recall; lower the reranker's `top_n` or drop to a cheaper model, which cuts prompt size and rerank cost; move from a multi-call response mode toward `compact`, which usually collapses synthesis into one call; enable `use_async` where the mode has independent work to parallelize; and turn on `streaming=True`, which does not reduce total time but converts a long wait into a fast first token. Budget them explicitly per stage and set a degradation path — skip the reranker under load rather than time out.

code

python · 18 lines
python
import time

retriever = index.as_retriever(similarity_top_k=25)
engine = index.as_query_engine(
    similarity_top_k=25,
    node_postprocessors=[reranker],
    response_mode="compact",
)

q = "what is the refund window for annual plans?"

t0 = time.perf_counter()
retriever.retrieve(q)
t1 = time.perf_counter()
engine.query(q)
t2 = time.perf_counter()

print(f"retrieval {t1 - t0:.3f}s, rerank+synthesis {t2 - t1:.3f}s")

go deeper

for a junior

Know that a slow query engine has separable stages — search, reranking, and the LLM call — and that the LLM step is usually the largest.

for a middle

Be able to time retrieval alone against the full query, and explain how the response mode's call count and the candidate pool size drive the total.

for a senior

Show a disciplined process: attribute per stage, move one knob at a time, know that reranker cost tracks candidates and synthesis cost tracks LLM calls, and account for cold models and provider tails.

for a principal

Own the budget as a contract — milliseconds allocated per stage, a pre-agreed degradation ladder that sheds quality instead of failing, and an explicit position on whether the SLO measures first token or full response.

## Attribute before you tune The failure mode in latency work is guessing. A LlamaIndex query engine has three stages and each can dominate, so measure them apart: - **Retrieval**: time `retriever.retrieve(q)` directly. This is one embedding call plus a vector-store round trip — typically tens of milliseconds. - **Postprocessing**: the difference between the retriever's time and the engine's time, minus synthesis. In practice you isolate it by running the engine with and without `node_postprocessors`. - **Synthesis**: everything left. Sanity-check it by counting LLM calls — the response mode determines that count, and it is the number people most often get wrong. A useful trick during attribution is a `no_text` response mode, which runs retrieval and postprocessing and skips generation entirely, giving you the front half's cost cleanly. ## The knobs, and what each one actually costs **`similarity_top_k`.** Almost free in the vector store, expensive everywhere downstream. It is the multiplier on reranker work and, without a reranker, on prompt size. Cutting it is the fastest win and the one that most directly costs recall. Never cut it blindly — check the recall@k curve first, because if the curve flattened at 15 and you are running 50, the first 35 are pure latency. **Reranker choice and placement.** A cross-encoder does one forward pass per candidate, so its cost is `similarity_top_k` times per-pair inference. Options in descending quality: hosted reranker (adds a network hop and a rate limit), local cross-encoder (adds CPU or GPU contention), smaller cross-encoder, no reranker. Reranking is also the most natural thing to shed under load, because the pipeline still returns an answer without it. **Response mode.** The dominant synthesis term. `refine` issues one sequential LLM call per node — ten nodes is ten round trips, and that alone will miss most interactive budgets. `compact` usually collapses to a single call. `tree_summarize` makes more calls but its batches at each level are independent, so `use_async=True` turns node count into tree depth. Check what mode you are actually running before optimizing anything else; a mode set months ago for a summarization endpoint and copied into a chat endpoint is a classic finding. **Model choice for synthesis.** A smaller or faster model on the synthesis call often moves p95 more than any retrieval tuning, and the quality loss is bounded when the reranker has already handed it three highly relevant passages. Retrieval quality and synthesis model size are substitutes to a degree. **`streaming=True`.** It does not reduce total latency at all — it changes what the user experiences, moving the perceived measure from full-response time to time-to-first-token. If your SLO is written against total time, streaming does not help you meet it; if it is written against perceived responsiveness, it may be the whole fix. Decide which SLO you are actually defending. ## Budget per stage, then defend it Write the budget down as a split, for example: 50 ms retrieval, 150 ms rerank, 800 ms synthesis, with 2 s p95 end to end. A per-stage budget makes regressions attributable — when p95 moves, one line moved — and it makes the tradeoff conversation concrete, because raising top-k is now visibly spending someone else's milliseconds. Then decide the degradation path in advance. Under pressure a query engine can shed work gracefully: drop the reranker and pass fewer retrieved nodes straight through; fall back to a cheaper synthesis model; cap `similarity_top_k`. Each is a quality reduction that still returns an answer, which is almost always better than a timeout. Choose the order deliberately and make it observable so a degraded response is visible in traces rather than a mystery. ## The things that are not latency problems Two confounders masquerade as engine slowness. First, the hosted dependencies — embedding endpoint, reranker API, LLM provider — have their own tail behaviour, and a p95 that is fine at the median and terrible at the tail often reflects a provider's tail rather than your configuration; timeouts and retries there need explicit policies, and a naive retry doubles the worst case. Second, cold state: a local cross-encoder loading its model, or an in-memory lexical index rebuilding at boot, makes the first requests after a deploy dramatically slower. Warm those on startup rather than accepting a latency spike on every rollout. ## What good judgment looks like here The principal-level answer is not a list of knobs; it is the discipline of attributing before tuning, holding one knob fixed while moving another, expressing the target as a per-stage budget, and pre-deciding what quality you will trade when the budget is breached in production rather than improvising during an incident.

  • Does enabling streaming help you meet a p95 latency SLO?
    Only if the SLO is written against perceived responsiveness. Streaming does not reduce total generation time; it delivers the first tokens sooner, so time-to-first-token collapses while end-to-end time is unchanged. If the SLO measures the complete response, streaming buys nothing and you need a real reduction — fewer LLM calls, fewer prompt tokens, or a faster model.
  • Which single setting most often turns out to be the hidden cost when synthesis is slow?
    The response mode. A mode that refines node by node issues one sequential LLM call per retrieved node, so a top-k of ten becomes ten round trips. Compacting usually collapses that to one. Confirm which mode the engine is actually running before tuning retrieval — the count of LLM calls is a configuration fact, not a mystery.
  • How do you shed load without failing requests?
    Pre-define a degradation ladder and make each step observable: skip the reranker and pass the top retrieved nodes straight through, cap the candidate pool, fall back to a smaller synthesis model. Each returns a slightly worse answer instead of a timeout. Deciding the order in advance and tracing when it triggers is what makes it operable rather than improvised.
  • Why can retrieval quality substitute for synthesis model size in a latency budget?
    Because a strong retrieval and reranking stage hands the model three highly relevant passages, and answering from clean context is a much easier task than reasoning over noisy context. That widens the range of smaller, faster models that produce acceptable answers — so investment in precision upstream buys you latency headroom downstream.

saying these in an interview costs you the question

  • Tuning knobs before measuring which stage is slow
  • Assuming streaming reduces total response time
  • Treating a higher top_k as free because the vector store is fast
  • Retrying a slow hosted call without bounding the worst case
  • Having no planned degradation and letting requests time out

context