skip to content

You are setting the filtering grammar for a public API that many teams will extend over years, and clients keep asking for richer queries. How do you decide between a small fixed set of filter parameters and a general query language, and how do you leave room to evolve without a breaking change?

level: principalimportance: nice to knowfreq 30%

answer

  1. shipped grammar = permanent surface
  2. 400-unsupported-filter log = prioritized demand signal
  3. reserve `$`/filter namespace before you need it
  4. publish the capability matrix, feature-detect
  5. POST /search as the escape hatch, loses GET caching

basics

~20 s

Decide on evidence, not aesthetics: adopt a general language only when many clients genuinely need OR and grouping. Start with a bounded field×operator grammar, keep it additive, reserve syntax up front, and offer a separate search endpoint as the escape hatch for queries the URL should not carry.

solid answer

~60 s

I treat it as a commitment-versus-optionality call. A bounded parameter grammar is cheap to run, safe to validate, and cost-predictable, but it can never express cross-field OR. A general language (RSQL/FIQL, OData `$filter`) buys expressiveness and, once shipped, becomes permanent surface: every operator is contract, every parser bug is a public bug, and cost stops being predictable. So: start bounded. Instrument what clients actually ask for — count `400`s on unsupported filters, and look at multi-request patterns that are clients emulating an OR client-side. Adopt a language when that evidence is broad, not when one team asks. To stay evolvable without breaking: reserve syntax early (`$`-prefixed or a dedicated `filter=` parameter left unused), make field and operator sets additive-only, expose the capability matrix as machine-readable discovery so clients can feature-detect, and version the *matrix* not the URL. Provide `POST /resource/search` with a JSON query body as the pressure valve for large or nested queries. And commit to one grammar estate-wide — two half-adopted grammars is the outcome to prevent.

go deeper

for a junior

Not expected to lead this; be able to state that a small fixed parameter set is the safe starting point and a query language is a large commitment.

for a middle

Contrast expressiveness against validation and cost, and mention picking an existing standard over inventing syntax.

for a senior

Bring the evidence loop (unsupported-filter telemetry, client-side OR emulation), day-one caps, and the POST-search escape hatch with its caching tradeoff.

for a principal

Argue permanence and reversibility: reserve syntax, publish a machine-readable capability matrix, version the vocabulary rather than the path, standardize estate-wide with a shared implementation, and plan the migration for services already diverged.

## Framing the decision The question is not "which syntax is nicer". It is: how much query surface am I willing to make permanent, and how much optionality can I keep while still serving clients today? Every filter capability you ship is a promise. Clients build on it, and in a public API you cannot see who. A bounded grammar makes small promises (these fields, these operators) that are individually removable-ish and collectively cheap to serve. A general query language makes one enormous promise: *arbitrary predicates over the declared model*. You cannot walk that back later, and its cost profile is set by the caller rather than by you. ## The case for starting bounded A field×operator matrix is: - **Validatable** — every request either matches the matrix or gets a `400`. - **Cost-bounded** — the worst possible query is a known, testable shape, so capacity planning is real rather than aspirational. - **Documentable and generatable** — one declaration produces docs, SDK methods, validation, and tests. - **Additive** — new fields and operators do not disturb existing clients. And it is usually sufficient: the overwhelming majority of real list queries are a conjunction of a few constraints plus a time range. ## When to graduate Adopt an expression language on evidence, and name the evidence you'd collect: - The rate and content of `400 unsupported filter` responses, aggregated by requested field/operator — this is your feature-request queue, already prioritized by demand. - Clients issuing N requests and unioning results client-side, which is an OR emulated in the wrong place (visible as bursts of near-identical calls from one key). - Clients pulling large pages and filtering locally — a sign that the server-side vocabulary is too poor, and a bandwidth and privacy problem too. - Repeated bespoke endpoints (`/orders/open-or-large`) accreting in the API — each one is an unmet grammar need frozen into a URL. If that evidence is broad, pick an existing standard rather than inventing syntax: RSQL/FIQL is compact and URL-friendly with off-the-shelf parsers; OData is far bigger but standardized end to end with real tooling. Inventing a grammar means owning a parser, a spec, and a decade of edge cases — occasionally right, rarely. ## Leaving room to evolve **Reserve syntax before you need it.** Declare up front that parameter names beginning with `$` (or a specific reserved `filter` parameter) belong to the API's control vocabulary and are not resource fields. Without this reservation, the day a resource gains a field called `filter` or `sort`, your control parameters and your data model collide, and both fixes are breaking. **Make the capability discoverable.** Publish the filter matrix as a machine-readable document (an endpoint, or annotations in the OpenAPI description). Clients then feature-detect instead of hardcoding, and you can add capabilities without a version bump. This is what makes "start small" safe: growth is observable to clients. **Version the capability, not the path.** Adding a filterable field or operator is additive; removing one is the breaking act. So run a deprecation process on the matrix — mark deprecated, measure usage per API key, notify, then remove — instead of minting `/v2` because the query vocabulary grew. **Keep a pressure valve.** `POST /resource/search` with a JSON query body handles queries that are too long for a URL, too nested for a query string, or too sensitive to appear in access logs and browser history. It costs GET cacheability and bookmarkability, which is exactly why it is the exception, not the default. Naming this tradeoff — and that a POSTed search is not idempotent-cacheable — is the mark of a considered answer. **Cap from day one, not after the incident.** If you do adopt a language: maximum AST depth, maximum terms, maximum distinct fields, a required selective predicate on large collections, per-query timeouts, and cost-aware quotas. Retrofitting caps onto clients who already write 40-term predicates is a breaking change you will not enjoy. ## The organizational half The grammar is a platform decision, not a service decision. If each team picks, you get four grammars, four validators, four sets of security bugs, and clients that must learn all of them. Write it into the API style guide, provide the shared library that implements the matrix, and make review of new filterable fields cheap — the friction you add is what keeps expensive fields from appearing quietly. Equally, plan the migration story: if a service already ships a different grammar, support both in parallel behind the discovery document, measure usage, and retire the old one on a published schedule, rather than declaring the standard and letting reality diverge from it. ## What separates a principal answer Not the syntax preference — the reasoning about permanence, evidence, and reversibility: start with the smaller promise, instrument demand so the decision to grow is data-driven, reserve the syntax that makes growth non-breaking, cap before the language ships, and treat the vocabulary as versioned contract owned by the platform rather than by whichever team needed a filter last quarter.

  • What would make you switch to a POST search endpoint instead of extending the query-string grammar?
    Queries that exceed practical URL length, need real nesting, or carry values that must not land in access logs, proxy logs, and browser history. The tradeoffs are explicit: you lose GET caching, bookmarkable links, and the safe/idempotent semantics of GET, so it should be an additional, documented surface for heavy queries rather than the default path for ordinary list reads.
  • How do you retire a filter field that clients already depend on?
    Measure first: usage per API key from the filter-shape metrics. Mark it deprecated in the machine-readable capability document and in responses, notify the identified callers with a replacement, then enforce with a sunset date — returning a clear 400 naming the removed field and its replacement. The point is that the capability matrix is versioned contract, so removal follows a deprecation process rather than arriving as a surprise.

saying these in an interview costs you the question

  • Adopting OData or a custom query language because it is expressive, with no evidence of client demand and no cost caps
  • Inventing a bespoke filter syntax when RSQL/FIQL or OData would do
  • Assuming a query language can be tightened later without breaking clients who already wrote complex predicates
  • Not reserving a control-parameter namespace, so a future resource field collides with `filter` or `sort`
  • Letting each service pick its own grammar and calling the inconsistency team autonomy

context