skip to content

Retrieval Integration

Getting the right passages into the window: hybrid BM25 plus embeddings, reranking, contextual chunk summaries, grep-style agentic search. Interviewers ask why good context still gives wrong answers.

on this pageshow

questions

4

When should an agent grep a repository instead of querying an embedding index?

level: middleimportance: must knowfreq 64%

answer

  1. exact string versus unknown wording
  2. who pays to keep the index fresh
  3. completeness matters more than similarity here
  4. pointers cost less than chunks
  5. one tool among glob and list

basics

~20 s

Grep when the query is an exact string — a symbol, an import, an error code — because lexical search returns every match and reads the current files. Embedding search earns its cost only when the wording is unknown.

solid answer

~50 s

The decision is about guarantees, not about which technique is fancier. Lexical search over the filesystem matches characters: it returns **every** occurrence of a literal pattern, from the bytes on disk right now. Dense retrieval returns the k nearest neighbours by meaning: it tolerates unknown wording, but it is approximate, capped by k, and answers from an index built at some earlier moment. A dependency-upgrade agent that must move a 12,000-file repository off `import requests` needs completeness and freshness, so it greps; twenty semantically relevant files are not an answer when one missed call site breaks the build. Embeddings win the moment the query is conceptual and the vocabulary is unknown — "where do we decide whether a refund is allowed" against code that calls it `eligibility_gate`, or a prose corpus of incident write-ups. Treat retrieval as one tool among glob, grep and directory listing, and let the agent choose per query, with a fallback ladder rather than one fixed pipeline.

code

bash · 5 lines
bash
# Shape first: how many files even touch it?
rg --count-matches '^import requests' --glob '*.py' .

# Then pull the exact lines from the narrowed set
rg -n '^import requests' --glob 'services/**/*.py' .

go deeper

for a junior

Know that grep matches exact text while embedding search matches meaning, and be able to say which one you would reach for given a specific query such as finding every file that imports a library.

for a middle

Explain the mechanics: approximate top-k versus exhaustive literal matching, and the fact that an index reflects the corpus at build time. Be ready to describe a fallback ladder from glob to grep to semantic search.

for a senior

Show production judgement — index staleness on a repository that changes every commit, silent false negatives from narrow patterns, and capping grep output so it does not flood the window. Justify keeping both modes on a prose corpus.

for a principal

Own the strategic call: whether an embedding index is worth building and maintaining at all for a given corpus, what its rebuild and drift costs are, and how tool descriptions steer the model's per-query choice at scale.

## Two search primitives, two guarantees Lexical search — grep, ripgrep, glob, a directory listing — matches characters. Ask it for `import requests` and it returns every line in the tree containing that string, with path and line number, read from the bytes on disk at this moment. It knows nothing about meaning: `import httpx` will not match however similar the intent is. Its recall on the literal pattern is total; its precision is whatever the pattern deserves. Dense retrieval maps text into vectors and returns the k nearest neighbours of the query vector. It finds passages that mean something similar even when they share no words — but it can never promise completeness. Approximate nearest-neighbour search is approximate by construction, k is a cutoff you chose, and the vectors were built from a snapshot of the corpus taken at index time. Those two guarantees — completeness on a literal string versus tolerance of unknown wording — are the whole decision. ## Why coding agents drifted toward grep As of mid-2026, filesystem and grep-based search has largely displaced embedding indexes as the default retrieval path inside coding agents. Four properties of a code repository drive that. **The queries are literal.** Symbol names, imports, error strings, environment-variable names, config keys. The agent usually knows the exact token it is hunting, or can guess a small family of spellings. **The task demands completeness.** An agent moving a 12,000-file repository off one HTTP client cannot accept "here are the twenty most relevant files". A missed call site is a broken build, not a slightly worse answer. Grep is also auditable: you can count matches, count edits, and compare. **The corpus mutates constantly.** Every commit, every branch switch, every edit the agent itself makes invalidates part of an index. Keeping vectors current means a rebuild pipeline whose failure mode is silent — stale hits that look entirely plausible. **The tree is already an index.** Directory layout, naming conventions, tests sitting beside sources: a glob narrows the search space cheaply, with no build step and no drift. ## Where embedding search still wins The moment the query is conceptual and the vocabulary is unknown, lexical search collapses. "Where do we decide whether a refund is allowed" finds nothing when the code names it `eligibility_gate`. Prose corpora — policy documents, incident write-ups, support tickets, meeting notes — express one idea in many phrasings, which is exactly the case dense retrieval was built for. Paraphrase, synonymy, typo tolerance and cross-language matching all belong to it. Serious systems over prose keep both lexical and dense scoring rather than choosing. ## Choose per query, not per system The context-engineering framing is that retrieval is one tool among several the model may call: glob to narrow, grep for exact matches, a directory listing to orient, semantic search when the wording is unknown, a file read to pull the span it actually needs. Each tool's description has to say *when it wins*, because the model's selection accuracy is only as good as the descriptions it is given. A practical ladder looks like: glob the likely subtree, grep the exact identifier, widen the pattern, fall back to semantic search, then read only the lines that matter. ## The token-budget argument Grep output is a list of pointers — path, line number, the matched line — not content. That fits the just-in-time principle: carry a lightweight identifier and load the material when it is needed. Dense retrieval returns whole chunks, and you pay for every one of them whether the model uses it or not. Across a long session that is the difference between a few hundred tokens spent to locate work and tens of thousands spent to maybe locate it. ## Failure modes on both sides Grep fails loudly and quietly. Loudly when a broad pattern floods the window with thousands of hits — count first, list filenames only, then narrow. Quietly when a narrow pattern returns nothing and the agent concludes the code does not exist, having missed aliased imports, dynamic dispatch, generated files or a different spelling. Dense retrieval fails by staleness, by chunk boundaries slicing a function in half, and by returning confident near-matches with no signal that the true answer was never in the index. ## What the interviewer is listening for Not a slogan. The rule: exactness and completeness point to lexical search; unknown wording and paraphrase point to dense retrieval; a volatile corpus penalizes any index; the choice is made per query with an explicit fallback ladder; and on a prose corpus the default flips. A candidate who says "we embed everything" for a repository, or "grep is enough" for a policy corpus, has answered a query-level question at system level.

  • The agent greps for a symbol, gets zero hits, and reports that the function does not exist. What went wrong?
    Zero lexical hits mean the pattern did not match, not that the concept is absent. Aliased imports, dynamic dispatch, generated or vendored code, a different spelling, or a search scoped to the wrong subtree all produce false confidence. The fix is procedural: widen the pattern, drop to a substring or case-insensitive match, list the directory to check scope, and only then fall back to semantic search before concluding anything.
  • How would you stop a broad grep from flooding the agent's context window?
    Make the first pass return shape rather than content: match counts per file, or filenames only. The agent reads that summary, narrows the pattern or picks a subtree, and only then requests matching lines — and even then with a hard cap on returned lines. This keeps the expensive tokens tied to a decision the model has already made, instead of paying for a result set it will mostly discard.

saying these in an interview costs you the question

  • Claiming embedding search returns all matching occurrences
  • Assuming a vector index stays current as the repository changes
  • Treating grep as obsolete because embeddings are newer
  • Choosing one retrieval mode for the whole system rather than per query
  • Feeding thousands of raw grep hits straight into the context window

context

open as a page

How do you set chunk size and top-k when retrieved passages share the context window?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Chunk size times top-k is the token bill. Pick the smallest k at which end-to-end answer accuracy stops improving — not the k that maximises retrieval recall — and deduplicate near-identical passages before they crowd out the one that matters.

open as a page

In late chunking, why is the whole document encoded before chunk vectors are pooled?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Late chunking encodes the entire document first so every token attends to the rest of it, then pools token embeddings per chunk. Each chunk vector therefore carries document context that independent chunk-by-chunk embedding throws away.

open as a page

When is agentic chunking worth its cost for an internal document corpus?

level: principalimportance: should knowfreq 26%

basics

~20 s

Agentic chunking pays when documents lack reliable structural markers, the corpus is small and slow-changing, and retrieval failures trace to boundaries cutting through procedures. Otherwise its per-document model call, nondeterminism and re-embedding cost outweigh cheaper structural splitting.

open as a page