A client of your JSON HTTP API complains that each item in a collection response carries far more data than the screen needs. Explain what a `fields` query parameter (sparse fieldsets) does about that, and how the JSON:API `fields[type]=` form differs from Google's partial-response `fields=` syntax.
answer
- fields = caller-driven projection
- JSON:API: fields[type]=a,b — keyed by type
- Google: fields=items(id,name) — path over the document
- id/type always returned
- unknown field → 400, not silent drop
basics
~20 sA fields query parameter lets the caller list the attributes it wants, and the server returns only those. JSON:API scopes it per resource type (fields[article]=title,body); Google's partial-response syntax describes a path through the response body (fields=items(id,name)), so it can reach nested objects.
solid answer
~40 sSparse fieldsets are caller-driven projection: the request names the attributes it wants and the response omits the rest, cutting payload size and serialization work without adding a new endpoint per screen. Two common spellings: - **JSON:API**: `GET /articles?fields[article]=title,author&fields[people]=name`. The parameter is keyed by *resource type*, so one request can project the primary resource and each included type independently. It is flat — you select attributes of a type, not a path. - **Google partial response**: `GET /files?fields=nextPageToken,items(id,name,owner/email)`. The value is a path expression over the *response document*, with parentheses for sub-selection and `/` for descent, so it can trim envelope fields and nested objects. Either way, the identifier (`id`/`type`) is normally always returned so the response stays dereferenceable, and unknown field names should be a 400 rather than silently ignored.
go deeper
Know that fields= lets the caller ask for a subset of attributes to cut payload size, and give one concrete example of the syntax.
Contrast the type-keyed JSON:API form with the document-path Google form, and note validation, identity fields, and pushing projection into the query.
Talk about where projection actually saves cost, caching/ETag consequences, and strict selector validation with depth and length limits.
Frame it as a contract-surface decision: one canonical resource with caller projection versus a proliferation of screen-shaped endpoints, and what that costs in cache fragmentation and client coupling.
## The problem: over-fetching A REST resource has one canonical representation. If `GET /articles/42` returns forty attributes, every client pays for all forty — bytes on the wire, JSON serialization on the server, parsing and memory on the client — even a list screen that shows only a title. Multiply that by a 100-item collection and a mobile network, and it becomes a real latency cost. The usual bad fixes are per-screen endpoints (`/articles/mobile-list`, `/articles/summary`) which multiply the contract surface, or a `view=summary` flag with a handful of server-defined shapes, which is coarse and keeps changing. **Sparse fieldsets** invert it: the caller declares which attributes it wants and the server projects the representation down to them. ## JSON:API `fields[type]=` JSON:API defines the parameter family `fields[TYPE]`, whose value is a comma-separated list of field names for that resource *type*: `GET /articles?include=author&fields[articles]=title,author&fields[people]=name` Key properties: - It is keyed by type, not by position in the document. One `fields[people]=name` applies wherever a `people` resource appears — as the primary data or as an included related resource. - A type that has no `fields` key is returned in full. A type with an empty list returns only its identifier. - `id` and `type` are structural identity and are always present; they are not "fields" you can drop. - Relationship names count as fields, so `fields[articles]=title,author` keeps the `author` relationship linkage and drops other relationships. ## Google partial response `fields=` Many Google APIs accept a single `fields` parameter whose value is a small path grammar over the *response body*: `fields=nextPageToken,items(id,name,owner/email)` - Comma separates siblings. - Parentheses sub-select inside an object or array element. - `/` descends one level (`owner/email` is shorthand for `owner(email)`). - `*` may select all sub-fields at a level. Because it addresses the document, it can trim envelope members (`nextPageToken`, `etag`) and reach arbitrarily deep nesting — something the type-keyed form cannot express positionally. The cost is that the selector is coupled to the response *shape*: rename a wrapper and every client's selector breaks. ## What both forms buy you - **Bandwidth and parse time**, which dominate on mobile and high-fan-out internal calls. - **One canonical resource** instead of a family of near-duplicate endpoints. - **Backward-compatible growth**: you can add attributes to a resource without inflating the responses of clients that project. ## Server-side implementation notes Projection is only a real win if it reaches the data layer. Selecting `title` but still loading the full row plus three joins saves bytes and nothing else. Serious implementations push the field set down into the query (column selection, skipping expensive computed fields, skipping downstream calls that only feed omitted attributes). Validate the selector. Unknown or malformed field names should return `400` with a machine-readable error naming the offending token; silently ignoring them means a client that typos `titel` gets a response missing the field it thinks it asked for and fails obscurely. Enforce a maximum selector length and nesting depth so the parameter cannot be used as a parser-abuse vector. Always return identity (`id`, and for JSON:API `type`) regardless of the selector, plus anything the response would be incoherent without — a pagination envelope, an `ETag`-relevant version field if your concurrency story depends on it. ## Boundaries Field selection is *reduction only*: it can never return data outside the resource's canonical representation. Pulling in related resources' bodies is the separate expansion mechanism. And projection does not weaken authorization — a caller that may not see `salary` must not see it because it asked for it; the field set is filtered by permission, not the other way round.
- If a client sends `fields[articles]=titel` (a typo), what should the API do?Return `400 Bad Request` with an error that names the unknown field. Silently ignoring it produces a response missing an attribute the client believes it requested, which surfaces later as a confusing null or a crash in the client. Strict validation turns a client bug into an immediate, self-describing failure.
- Should `fields` affect the ETag or cache key of the response?Yes. Two responses for the same resource with different field sets are different representations, so the cache key must include the selector and the ETag must differ. Practically this means adding the parameter to `Vary`-equivalent cache keying, and it is why unconstrained field selection fragments shared caches.
saying these in an interview costs you the question
- Claiming sparse fieldsets can pull in fields that are not part of the resource's representation
- Silently ignoring unrecognised field names instead of rejecting the request
- Projecting only at the serialization layer while still doing all the expensive loading
- Assuming `id`/`type` can be projected away
- Treating the field list as an authorization bypass — 'they asked for it so we return it'