skip to content

How is a JSON:API document (media type application/vnd.api+json) laid out - what goes in data, relationships and included - and what does that structure buy you over ad-hoc JSON?

level: middleimportance: should knowfreq 34%

answer

  1. data / errors / meta at top level, never data plus errors
  2. resource = type + id + attributes + relationships
  3. relationships hold type-and-id identifier objects, not full resources
  4. included = deduplicated compound document
  5. include, fields[type], sort, page; errors source.pointer

basics

~20 s

Top level holds data, errors or meta. data is a resource object with type, id, attributes and relationships; relationships hold identifier objects (type plus id) and links; included carries the full related resources, deduplicated. It gives you a normalised graph plus fixed conventions for includes, sparse fields and paging.

solid answer

~50 s

A JSON:API response has a top level with data (a resource object or an array), or errors, or meta. A resource object is always type, id, attributes for its own scalar state, relationships for its edges, and links. A relationship does not inline the related resource; it holds resource identifier objects - just type and id - plus links such as self and related. The actual related resources go once each in the top-level included array, so a compound document is a normalised graph with no duplication even when fifty comments share one author. On top of that the spec fixes the query conventions everyone otherwise invents: include=author,comments for what to pull in, fields[articles]=title,body for sparse fieldsets, sort, page and filter, and an errors array whose members carry source.pointer. The price is verbosity and rigidity: everything must be a typed, identified resource, which fits entity CRUD well and non-entity operations badly.

code

json · 19 lines
json
{
  "data": [
    { "type": "articles", "id": "1",
      "attributes": { "title": "REST at scale" },
      "relationships": {
        "author": { "data": { "type": "people", "id": "9" },
                     "links": { "self": "/articles/1/relationships/author",
                                "related": "/articles/1/author" } }
      },
      "links": { "self": "/articles/1" } },
    { "type": "articles", "id": "2",
      "attributes": { "title": "Caching" },
      "relationships": { "author": { "data": { "type": "people", "id": "9" } } } }
  ],
  "included": [
    { "type": "people", "id": "9", "attributes": { "name": "Ada" } }
  ],
  "links": { "self": "/articles?page[number]=1", "next": "/articles?page[number]=2" }
}

go deeper

for a junior

Name the top-level keys and the resource shape (type, id, attributes, relationships) and say included carries related resources.

for a middle

Explain identifier objects versus included, deduplication, and the include and fields parameters.

for a senior

Weigh the payoff - conventions and client stores - against verbosity and poor fit for non-entity operations, and mention the strict content-negotiation rules.

for a principal

Treat it as a platform-level commitment: it standardises paging, filtering and errors across teams, but locks you into entity-shaped modelling and its serializer ecosystem.

## Document skeleton Every JSON:API document is an object with at least one of: - **data** - the primary data: a resource object, an array of them, or null - **errors** - an array of error objects; data and errors must never both appear - **meta** - non-standard information and optionally **included**, **links**, and **jsonapi** (version info). ## Resource objects A resource object is identified by the pair **type** and **id** - the type is part of identity, not just a hint. Its own state lives under **attributes**; anything referencing another resource lives under **relationships**; **links** typically carries self. Crucially, attributes and relationships are separate namespaces, and foreign keys do not appear as attributes. authorId is not an attribute; the author relationship is where that edge is expressed. ## Relationships and identifier objects A relationship object may contain: - **data** - one or more **resource identifier objects**, which are the minimal type-and-id pairs - **links.self** - the relationship itself, which you can PATCH to change the linkage without touching the resource - **links.related** - where to fetch the related resource - **meta** That identifier-object indirection is what makes the format a graph rather than a tree. ## Compound documents and included When the client asks for related data with an include parameter, the server puts the full resources into the top-level **included** array. Each resource appears exactly once, keyed by its type and id, no matter how many relationships point at it. Clients typically index included by type and id on receipt and resolve pointers locally. This is the format's biggest practical win: a page of fifty articles by three authors carries three author objects, not fifty copies. ## The conventions that come free - **Sparse fieldsets**: fields[articles]=title,body limits which attributes are serialised, per type. - **include**: dot paths such as comments.author walk more than one hop. - **sort**, **page**, **filter**: sorting and pagination are specified; filter is deliberately left as a reserved, unspecified strategy. - **Errors**: an errors array whose members have id, status, code, title, detail, and source - where source.pointer is a JSON Pointer into the request document and source.parameter names a query parameter. Content negotiation is strict: the media type is application/vnd.api+json and the spec dictates responding 415 when a client sends it with media type parameters the server does not support, and 406 when it appears in Accept only with unsupported parameters. ## What it buys, and what it costs Buys: one convention instead of ten bikeshed decisions per team; deduplication for graph-shaped data; off-the-shelf client libraries that give you an identity map and a store almost for free; a uniform error shape; a linkage-editing mechanism via relationship endpoints. Costs: verbosity - a two-field object becomes a nested structure with type, id and attributes; a steep fit problem for anything that is not a persistent entity, such as a calculation endpoint or a workflow command; more serializer machinery; and enough rigidity that teams end up bending the spec, at which point the client-library advantage evaporates. ## How to choose it JSON:API pays off when the domain is genuinely a graph of entities that clients traverse and cache, and when several clients would otherwise each invent their own include and paging syntax. It is a poor fit for RPC-ish or document-shaped APIs, where its ceremony buys nothing.

  • Why does a relationship carry only type and id instead of the whole related resource?
    Because the document is a normalised graph. Identifier objects express linkage, while the full resources live once each in included, so a shared author appears a single time however many articles reference it. It also lets a client ask for linkage without the payload, and lets relationship endpoints be edited independently of the resource.
  • How does a client control how much data comes back?
    With two orthogonal parameters: include names which relationship paths to materialise into included, and fields[type] limits which attributes of each type are serialised. Together they let one endpoint serve a lightweight list view and a heavy detail view without separate routes, at the cost of variable response shapes and cache keys.

saying these in an interview costs you the question

  • Putting foreign keys such as authorId in attributes instead of using relationships
  • Inlining full related resources inside relationships rather than in included
  • Returning data and errors in the same document
  • Assuming id alone identifies a resource - identity is type plus id

context