skip to content

A search endpoint for an HTTP API needs dozens of optional filters, some with lists of hundreds of ids. How do you design the parameters, and at what point do you stop using the query string?

level: seniorimportance: should knowfreq 42%

answer

  1. flat optional params, bare collection still valid
  2. repeated keys vs comma list — document one
  3. request line caps ~4-8 KB → 414
  4. POST /search loses caching; say it is a read
  5. saved-search resource = cacheable GET results

basics

~20 s

Keep filters as flat, optional, individually-named query parameters with repeated keys or comma lists for multi-values, plus standard sort and paging params. When the query outgrows practical URL limits (roughly 4-8 KB) or needs nested boolean logic, move to a POST search endpoint that takes a JSON body.

solid answer

~60 s

Design principles first: every filter is **optional**, flat, and independently combinable; the bare collection URL must still work. Names are consistent API-wide (`sort`, `page`/`cursor`, `limit`, `fields`). Multi-value filters use either repeated keys (`?status=open&status=held`) or a comma list (`?status=open,held`) — pick one and document it, since HTTP does not define repetition semantics. Ranges get suffixed names (`createdAfter`, `priceMin`) rather than an embedded mini-language, which stays parseable and validatable. The stopping point is practical, not spec-defined: servers and proxies typically cap the request line around 4–8 KB, and long URLs also blow up logs and cache keys. Once a caller must pass hundreds of ids, or nested AND/OR logic, I add `POST /orders/search` with a JSON body. I document that it is a read expressed as POST — losing GET's cacheability and safety — and I keep the simple GET form for the common cases so ordinary clients keep their caching. A middle option, if caching matters, is to persist the filter as a saved-search resource and GET it by id.

code

http · 8 lines
http
GET /orders?status=open&status=held&createdAfter=2026-01-01&sort=-createdAt&limit=50 HTTP/1.1
Host: api.example.com

POST /orders/search HTTP/1.1
Host: api.example.com
Content-Type: application/json

{"any":[{"status":["open","held"]},{"customerIds":["c1","c2"]}],"limit":50}

go deeper

for a junior

Say filters, sorting, and paging go in the query string as optional named parameters, and that very long queries do not fit in a URL.

for a middle

Give concrete conventions — repeated keys versus comma lists, range parameter naming, standard paging names — and cite the practical URL length limits.

for a senior

Weigh the POST-search tradeoff explicitly (cacheability, safety, monitoring), mention the saved-search resource alternative, and cover validation, bounding, and whitelisting.

for a principal

Define the API-wide query grammar and its limits, decide how much expressiveness the platform will support before requiring a different mechanism, and account for cache-key behaviour and gateway limits across all consumers.

## Shape the parameters before worrying about limits A good filter surface is boring: flat, optional, orthogonal parameters with stable names. - **Optionality.** `/orders` with no parameters must return something sensible (a first page of recent orders). Every filter narrows from there. - **Flatness.** `?status=open&customerId=42&createdAfter=2026-01-01` is trivially parseable, validatable, and documentable in OpenAPI. Nested mini-languages like `?filter=status eq 'open' and total gt 100` buy expressiveness at the price of a parser you now own, an injection surface, and error messages nobody can act on. - **Consistency.** Paging, sorting, and field selection appear on every collection, so fix their names once (`page`/`pageSize` or `cursor`/`limit`, `sort=-createdAt`, `fields=id,total`) and reuse them everywhere. - **Multi-value.** HTTP does not define what a repeated query key means; frameworks vary. Choose repeated keys (`?status=open&status=held`) or a comma-separated list (`?status=open,held`), document it, and be aware that comma lists break if values can contain commas. - **Ranges and operators.** Suffixed names (`createdAfter`, `createdBefore`, `priceMin`, `priceMax`) or an explicit bracket convention (`price[gte]=100`) both work; the first is simpler and validates better. ## Where the query string actually runs out There is no limit in the HTTP specification, but there are hard limits in deployed software: request-line and header caps commonly sit between 4 KB and 8 KB (nginx `large_client_header_buffers`, many gateways and CDNs similar), and some intermediaries are stricter. Exceeding them typically produces `414 URI Too Long` or `431` — and often only for *some* users, when a specific proxy is in the path, which makes it a miserable bug to chase. Size is not the only trigger. Consider moving off the query string when: - a filter needs **nested boolean structure** (groups of OR inside AND); - the payload includes **hundreds or thousands of ids**; - filter values are **sensitive** and you do not want them in access logs, browser history, or referrer data; - the filter is a **document** in its own right — a saved segment or report definition. ## The POST-search endpoint The accepted solution is a companion endpoint: `POST /orders/search` with a JSON body describing the query. Be explicit about the tradeoffs you are accepting: - POST is neither safe nor cacheable in practice, so shared caches, CDNs, and conditional requests stop helping. - Naive retry logic and prefetching treat POST differently; state clearly in the docs that this POST is a read with no side effects. - Semantic monitoring gets harder — every search looks like the same endpoint in logs. Mitigations: keep the GET form for common, small queries so the hot path stays cacheable; return the same response envelope from both; and consider caching inside the application keyed by a hash of the normalised query body. ## The saved-search alternative If cacheable, linkable results matter, make the query itself a resource: `POST /order-searches` returns `201` with `Location: /order-searches/{id}`, and `GET /order-searches/{id}/results?page=2` is then a normal cacheable, bookmarkable, paginable GET. This costs storage and a lifecycle policy but is the cleanest fit when the same complex filter is executed repeatedly, or when results must be shared between users or systems. ## Operational notes Whatever shape you pick, **validate and bound** it: cap `limit`, cap the number of ids in a list, whitelist sortable fields (an unvalidated `sort` parameter is both an injection vector and a way to force full table scans), and reject unknown parameters loudly rather than ignoring them — silently dropping a misspelled filter is how clients ship code that returns far more data than intended. Also fix parameter **ordering normalisation** if you rely on caching: `?a=1&b=2` and `?b=2&a=1` are different cache keys to a naive cache, so instruct clients to build query strings deterministically.

  • What do you lose by switching a complex read from GET to POST, and how do you compensate?
    You lose safety and cacheability: shared caches, CDNs, and conditional requests no longer apply, and generic tooling can no longer assume the call is side-effect free. Compensate by keeping a GET form for common small queries, caching inside the application on a hash of the normalised body, and documenting explicitly that the POST performs no state change.
  • Why is silently ignoring an unrecognised query parameter a bad default?
    A client that misspells a filter believes it is narrowing the result set while the server returns everything, which can mean a data-volume incident or an information exposure. Rejecting unknown parameters with 400 surfaces the mistake at integration time instead of in production.

saying these in an interview costs you the question

  • Inventing a nested query mini-language and hand-writing its parser
  • Assuming URLs have no practical length limit because the RFC sets none
  • Allowing arbitrary field names in sort without a whitelist
  • Silently ignoring misspelled or unknown filter parameters
  • Switching every read to POST search and losing all caching for no benefit

context