skip to content

Your filter API gains an optional parameter every quarter - how do you decide between one composed builder, named statements, or a search service?

level: principalimportance: should knowfreq 42%

answer

  1. measure the combinations, not the surface
  2. concentrated or flat distribution
  3. pin the head, builder for the tail
  4. a second store buys lag
  5. each filter commits to an access path

basics

~20 s

Decide from the measured distribution of filter combinations, not the API surface. A composed builder suits a small orthogonal filter set; pinned statements suit a few dominant combinations; a search store suits text and facet filters, and costs a lagging copy.

solid answer

~50 s

Start with evidence: record a signature per call and rank the combinations actually requested. Three answers then compete. **One composed builder** stays right while filters are few and orthogonal and the existing access paths serve every combination acceptably. **Named statements** win when a handful of combinations carry nearly all traffic and each wants its own shape and index - the builder can survive alongside them for the tail. **A dedicated search store** wins when the filters are text or facet-shaped, when no single column is selective, or when the shape family has outgrown what indexing can serve; it costs you a second copy of the data, its sync lag, and a read-your-writes story. Most systems end up hybrid. The governance point matters as much as the choice: each new optional filter is a commitment to an access path and to shapes nobody will review, so make adding one require naming both.

go deeper

for a junior

Note the headline: the right implementation depends on which filter combinations users actually send, and that has to be measured rather than guessed from the list of parameters.

for a middle

Be able to compare the options concretely - one builder, pinned statements for hot combinations, or a separate search store - and name the access-path and freshness costs of each.

for a senior

Show the operating evidence you would gather first, the guard that bounds the worst request, and the contract tests that stop a pinned statement drifting from the builder it shadows.

for a principal

Own the governance: what adding one optional filter commits the organisation to, when narrowing the contract beats scaling the implementation, and how the freshness cost of a second store is priced before it is built.

## Decide from traffic, not from the API surface The number of optional filters an endpoint exposes tells you almost nothing. What matters is the **distribution of combinations actually requested**. Record a normalised signature per call - the sorted list of supplied filter tokens plus the sort token - and rank it. Two very different worlds hide behind the same API: - **Concentrated.** Three or four signatures carry the overwhelming majority of calls, and the tail is used a few times a week by internal users. - **Flat.** Hundreds of signatures are each used a little, typically because a UI exposes free-form facets. A concentrated distribution can be served by pinning a few statements. A flat one cannot, and no amount of index work will make it so. ## Three shapes of answer | Approach | Fits when | Principal costs | |---|---|---| | One composed builder | Filters are few and orthogonal; every combination is served acceptably by existing access paths | Unreviewed shapes; worst member unbounded; index coverage per combination is invisible | | Named statements for the head, builder for the tail | A handful of signatures dominate; each wants its own shape and index | Two code paths to keep semantically identical; drift between them is a real defect class | | A dedicated search store | Text or facet filters; no single selective column; the shape family exceeds what indexing can serve | A second copy of the data, its lag, reindexing, and a read-your-writes story to design | A fourth answer is sometimes the right one and is usually forgotten: **restrict the contract**. Replacing a free-form filter API with a small set of named searches - each with its own statement, index and test - removes the combinatorial problem instead of managing it. That is a product conversation, which is precisely why it needs a lead to have it. ## What each one really costs **The composed builder** is cheapest to extend and most expensive to reason about. Its failure mode is not a bug on day one; it is that after eight quarters nobody can say which combination is slow, which index serves which shape, or whether a given combination has ever run. Its bounded form requires a guard - at minimum, one selective filter is mandatory, list filters are length-capped, and page size is clamped. **Pinned statements** buy reviewability for the traffic that matters, and cost you duplication. The two paths must agree not just on rows but on the pager's total, the ordering and the tri-state meaning of every field. Contract tests that run the same request through both paths and compare are the cheapest insurance; without them, drift is discovered by a user. **A separate search store** solves the shape problem by changing the access model, and imports a distributed-systems problem in exchange: the copy is behind by some amount, so a user who has just edited a record may not find it; deletions must propagate or you serve tombstones; a reindex is an operation with a runtime and a rollback. None of that is a reason to avoid it - it is a reason to price it, and to insist the team can state the freshness contract out loud. ## Governance: what one more filter commits you to The recurring failure is that adding a filter looks like a one-line change. Make the true cost visible by requiring, for every new optional filter: 1. **The access path.** Which index or path serves the new combinations, or the explicit statement that the filter is only ever used together with an already-selective one. 2. **The expected selectivity.** A filter that removes 2% of rows is a filter that will never be worth planning for on its own, and knowing that up front changes the design. 3. **The grain declaration.** Whether the filter reaches through a collection and therefore changes the row grain, and which form it uses. 4. **A retirement trigger.** If the signature log shows the filter unused after a quarter, it goes. Removing an unused filter is cheaper than every future decision made in its presence. ## How the decision usually lands Most mature systems end up **hybrid, deliberately**: a builder for the long tail, two or three pinned statements for the dominant signatures, a guard that refuses unbounded requests, and a search store only where the workload is genuinely text- or facet-shaped. The signal to move a signature from the builder to a pinned statement is empirical - it appears in the top few by volume, or its latency distribution is the one users complain about. What separates a good answer here from a plausible one is refusing to make the choice from the API's appearance. The API's appearance is what the last four quarters of feature requests happened to produce; the traffic distribution is what the system is actually being asked to do.

  • What is the strongest argument against splitting out pinned statements for the dominant combinations?
    Duplication that drifts. Two paths must agree on rows, totals, ordering and the meaning of every optional field, and they diverge quietly when one is edited. It is manageable with contract tests that run the same request through both and compare, but if the team will not maintain those, one builder with a guard is the safer choice.
  • How would you make the freshness cost of a separate search store explicit?
    State it as a contract: how far behind the copy may be, what a user sees immediately after their own write, how deletions propagate, and how long a full reindex takes. Then check the product can live with those numbers before building, rather than discovering the answer through support tickets.
  • What would make you restrict the API instead of scaling the implementation?
    A flat distribution where the tail is used by a handful of internal users, or a filter set that grew from requests rather than from a model of what callers need. Replacing free-form filters with a small set of named searches removes the combinatorial problem outright and is usually cheaper than serving it well.

saying these in an interview costs you the question

  • Picks the approach from the filter count rather than measured usage.
  • Adds a second data store without stating a freshness contract.
  • Treats a new optional filter as a one-line change.
  • Duplicates statements for hot combinations with no contract tests.
  • Assumes indexing can keep up with a flat combination distribution.
  • Never considers narrowing the API to a set of named searches.