skip to content

Where does learned sparse retrieval like SPLADE fit between BM25 and dense embeddings?

level: middleimportance: nice to knowfreq 28%

answer

  1. neural, but still sparse over vocabulary
  2. weights predicted, not counted
  3. expansion adds terms never written
  4. served by an inverted index
  5. denser posting lists cost latency

basics

~20 s

SPLADE is a third leg: a transformer predicts a sparse weight for vocabulary terms, expanding a document or query with related terms it never literally contained. It keeps exact-term matching and inverted-index serving while fixing much of BM25's vocabulary mismatch.

solid answer

~50 s

Learned sparse retrieval keeps the sparse representation but learns it instead of counting it. A model projects text onto the vocabulary, producing a small set of weighted terms, and a sparsity regularizer keeps that set short. Crucially the projection performs **expansion**: a catalogue entry listed under an IUPAC name can pick up its common trade synonyms as weighted terms, so a query using either wording still matches through the inverted index. That closes BM25's biggest weakness without giving up literal token matching or interpretability — you can still read off which terms drove the score. The costs are real: a model inference at index time and usually at query time, and posting lists that are longer and denser than BM25's because expansion adds terms to many documents, which raises index size and query latency. Efficiency variants push the expansion to the document side to keep query-time work small. In hybrid setups it is fused just like any other ranked list.

go deeper

for a junior

Know that learned sparse retrieval produces weighted vocabulary terms from a neural model rather than raw counts, and that it still runs on an inverted index.

for a middle

Explain the expansion mechanism — terms that never appear literally get non-zero weight — and name the concrete cost, longer posting lists and a model pass in the indexing path.

for a senior

Show you would decide empirically whether it replaces the BM25 leg or adds a third one, and that you have thought through re-indexing, model versioning and the sparsity/latency dial.

for a principal

Own whether the organization takes on a trained retrieval component at all: who retrains it as corpora drift, what the re-encode cost is at your corpus size, and whether the recall gain justifies a model in the write path.

## The gap it fills Retrieval representations have historically come in two families. **Sparse** representations are vectors over the vocabulary — one dimension per term, almost all zeros — and are served by an inverted index, which is fast and exact because it only visits documents containing a query term. BM25 is the canonical hand-crafted sparse scorer. **Dense** representations are short continuous vectors from a neural encoder, served by approximate nearest-neighbour search, strong on paraphrase and weak on rare literals. Learned sparse retrieval sits between them: neural like dense, sparse and inverted-index-served like BM25. ## How SPLADE builds its representation SPLADE (SParse Lexical AnD Expansion) runs text through a transformer and uses its masked-language-modelling head to score every vocabulary term for the passage — essentially asking "which vocabulary words does this text activate?" — then pools and rectifies those scores into a non-negative weight per term. A sparsity regularizer, penalizing expected posting-list cost, pushes the great majority of those weights to zero, leaving a compact set of weighted terms per document and per query. Retrieval is then a dot product between two sparse vectors, exactly the operation an inverted index already performs. Two consequences follow: - **Term weights are learned, not counted.** BM25 derives a weight from term frequency, document frequency and length. SPLADE predicts it from context, so an incidental mention scores low and a topically central term scores high even at the same raw frequency. - **The vector is expanded.** Terms that never literally occur in the text can carry non-zero weight. This is the mechanism that closes vocabulary mismatch, and it is why the representation is described as lexical *and* expansion. ## A concrete case: a chemical reagent catalogue Entries carry systematic IUPAC names, common trade names, registry identifiers, purity grades and hazard phrasing, and users search with whichever they happen to know. BM25 alone fails when the query names the substance differently than the entry does — zero term overlap, zero score. A dense retriever bridges the naming gap but blurs the identifiers, retrieving chemically adjacent compounds, which is exactly the wrong error in a catalogue where the wrong reagent is a real-world hazard. Learned sparse handles the middle: the entry's representation is expanded with the synonymous names, so a synonym query matches through literal terms, while the identifier term itself remains a first-class dimension that must actually match. ## What it costs 1. **Inference in the indexing path.** Every document must pass through the model, and re-indexing after a model change means re-encoding the corpus. 2. **Query-time inference**, unless you use a document-side-expansion variant that keeps the query representation cheap — that is precisely the tradeoff those efficiency variants exist to make. 3. **Denser posting lists.** Expansion adds terms to many documents, so common expanded terms end up with long posting lists and queries touch more of the index than BM25 would. Index size and latency both rise, and the sparsity regularizer is the dial that trades effectiveness against that cost. 4. **Domain fit.** The model was trained on some distribution. On a specialized corpus its expansions may be unhelpful or wrong, and adapting it is a training project, not a config change. ## What it keeps Interpretability survives, which matters more than it sounds. You can inspect the weighted terms for a document and a query and see exactly which shared terms produced the score — impossible with a dense vector. Operationally you also keep the inverted-index machinery: exact filtering behaviour, incremental updates, and well-understood serving characteristics. ## Where it lands in a hybrid design Learned sparse does not replace fusion. Treated as a third ranked list, it merges alongside dense and BM25 through the same rank-fusion step, and the practical question is whether it earns its slot: it overlaps heavily with BM25 (it is lexical) and partly with dense (it handles synonymy), so the incremental recall it adds over the existing two legs is an empirical question on your corpus. Many teams run it as a *replacement* for the BM25 leg rather than an addition, keeping two legs rather than three, and reserve the three-leg configuration for corpora where measurement justifies the extra index. Be able to say plainly that the decision is measured, not assumed.

  • If SPLADE handles synonymy, why keep a dense leg at all?
    Because expansion works within the vocabulary. It adds related terms, but relevance that depends on sentence-level meaning, compositional phrasing or cross-lingual matching is not expressible as a bag of weighted terms. Dense embeddings encode that. Measure the overlap on your own corpus — sometimes learned sparse plus dense is the right pair and plain BM25 is the leg you drop.
  • What is the sparsity regularizer actually trading off?
    Effectiveness against serving cost. Every extra non-zero term in a document's representation lengthens some posting list, so queries touch more of the index. Penalizing expected posting-list cost during training keeps representations short and queries fast, at the price of dropping some expansion terms that would have helped recall. It is the main dial between a SPLADE that is accurate and one that is affordable.
  • How does deploying learned sparse change your indexing pipeline compared with BM25?
    It puts a model in the write path. Every document needs an inference pass before indexing, throughput is bounded by that model rather than by tokenization, and a model version change forces a full re-encode of the corpus — so you need a re-index strategy and version tracking on the representation. BM25 has none of these; its analyzer is deterministic and cheap.

saying these in an interview costs you the question

  • Describes SPLADE as producing a dense fixed-length vector
  • Says it replaces the inverted index with ANN search
  • Assumes expansion removes any need for a dense leg
  • Ignores that expansion lengthens posting lists and latency
  • Treats it as a drop-in config change with no re-indexing

context