skip to content

What can XML's markup tree with attributes and mixed content model that a JSON document's plain data model cannot?

level: middleimportance: must knowfreq 62%

answer

  1. markup versus a plain data model
  2. attributes are a second channel
  3. text and elements interleaved
  4. mixed content has no JSON counterpart
  5. namespaces versus key-prefix conventions

basics

~20 s

XML models documents: an element can carry attributes alongside child elements, children are ordered, and text and markup can interleave as mixed content. JSON models data — objects, arrays and scalars — so interleaved narrative must be encoded by convention.

solid answer

~50 s

XML is a markup language for **documents**. An element has a name, an optional set of **attributes**, and an ordered list of children, and those children may interleave text with elements — `<p>Signed on <date>2026-01-05</date> by the partner.</p>` is one node whose content is text, element, text, in that order. XML also gives you namespaces, so two vocabularies can coexist in one document without colliding. JSON is a **data model**: objects of name/value pairs, arrays, strings, numbers, booleans and `null`. There is no second channel like attributes, no mixed content, and no namespace mechanism. A record maps to JSON almost symmetrically, which is why resource-shaped APIs favour it; a paragraph with inline annotation does not, and must be rewritten as an array of parts that every reader has to agree on. Choose by the shape of what you are modelling.

code

json · 11 lines
json
{
  "invoice": {
    "id": "INV-2026-0041",
    "currency": "EUR",
    "note": [
      "Signed on ",
      { "date": "2026-01-05" },
      " by the partner."
    ]
  }
}

go deeper

for a junior

Recall that one is a markup language for documents and the other is a data model of objects, arrays and scalars. Name attributes and mixed content as things markup has and the data model does not.

for a middle

Explain the capability gap concretely: interleaved text and elements, a separate attribute channel, and namespaces. Then show what a JSON producer must invent to carry interleaved content and why that convention is invisible to a parser.

for a senior

Argue from payload shape rather than fashion. Show that a resource maps almost symmetrically onto the plain data model, that a document with inline annotation does not, and what breaks when a team forces the wrong one.

for a principal

Treat the narrower model as a deliberate purchase: fewer modelling arguments and cheaper integration, paid for by pushing any document-shaped requirement into a convention your organisation must then own and document.

## Two different things that both produce text Both encodings emit readable characters with names inline, and that shared surface hides a genuine difference in what they were built to describe. - **XML is a markup language.** It came from marking up documents: you take prose and annotate parts of it. Its unit is the **element**, which has a name, an ordered list of children, and a separate slot for **attributes** — name/value pairs attached to the element itself rather than nested inside it. - **JSON is a data model.** Its units are the object (a collection of name/value pairs), the array (an ordered list), and four scalar kinds: string, number, boolean and null. There is nothing else, on purpose. ## The three capabilities XML has and JSON does not 1. **Attributes — a second channel on every node.** An element can carry metadata (`currency`, `id`, `lang`) in a slot distinct from its content. JSON has one channel: a member is a member. Whether a given fact belongs in an attribute or in a child element is a permanent modelling argument in XML and simply does not arise in JSON, which is an advantage for JSON and a loss of expressiveness at the same time. 2. **Mixed content — text and elements interleaved.** An element's children can be a sequence such as *text, element, text*. This is the capability that markup exists for, and it has no JSON counterpart: JSON's containers hold either named members or list items, never a run of prose with annotations embedded in it in order. 3. **Namespaces — two vocabularies in one document.** XML lets a document combine elements from independently defined vocabularies with a prefix mechanism that prevents name collisions. JSON has no namespace concept, so ecosystems fall back on **conventions**, such as prefixing extension keys, and nothing in the encoding enforces them. A fourth difference is subtler: **order**. XML children are ordered by definition. A JSON array is ordered too, but the members of a JSON **object** are not carried as an ordered collection by the data model, so two readers may legitimately see them in different orders. ## What mixed content costs in JSON Suppose a partner-facing API returns an invoice whose note field is a sentence with a machine-readable date embedded in it. In markup that is one node. In JSON you must invent a representation, and every representation is a convention that both sides have to implement: - an array of alternating parts (strings for the prose, objects for the annotations), which preserves order but makes every consumer walk it; - a plain string plus a side list of offsets, which is compact but breaks the moment anyone re-encodes or re-normalises the text; - a plain string with the annotation thrown away, which is what actually happens under deadline. None of these is wrong. The point for an interview is that **the encoding gives you no help**: the structure lives in a convention document, not in the grammar, so a new consumer can parse the payload perfectly and still misread it. ## When each shape is the right one | Payload shape | Natural fit | Why | | --- | --- | --- | | A resource with named fields | JSON | The data model and the record are the same shape, so the mapping is near-symmetric | | A list of like items | JSON | Arrays are ordered and untyped by design | | Prose with inline annotation | XML | Mixed content is exactly the case markup was designed for | | A document combining two vocabularies | XML | Namespaces make coexistence a grammar feature rather than a convention | | A payload a human will hand-edit | Either, with a caveat | Markup's closing tags are noisy, but its structure survives careless editing better | ## The register an interviewer is listening for A weak answer says XML is verbose and old and JSON is modern. That is a fashion claim, and it does not survive the follow-up. The strong answer names the **capability gap in one direction**: markup can express interleaving and dual-channel annotation natively, and the plain data model cannot, so anything of that shape becomes a convention layered on top. In the other direction, JSON's narrower model is precisely why it maps onto records with so little ceremony and why there is nothing to argue about when modelling one — a smaller model is a real feature, not a deficiency, when the thing you are modelling fits inside it.

  • In XML, when should a fact be an attribute rather than a child element, and why does that question never arise in JSON?
    The usual guidance is that attributes carry metadata about the element — an identifier, a language, a unit — while child elements carry the content itself, but the line is genuinely contested and teams settle it by convention. It never arises in JSON because the data model has a single channel: every fact is a member of an object, so there is no second place to put it.
  • JSON has no namespaces. How do ecosystems let two parties extend the same document without colliding?
    By convention: reserving a key prefix for extensions, or nesting all vendor additions under one agreed member. Nothing in the encoding enforces either, so a collision is caught by review or by a validator, not by the parser. XML makes the same guarantee a grammar feature through prefixed vocabularies.
  • Are the members of a JSON object ordered?
    The data model does not carry object member order as meaningful, so a producer and a consumer may see the members in different orders and both be correct. Arrays are different — they are ordered lists and their order is part of the value. Building a contract that depends on object member order is therefore a latent defect.

saying these in an interview costs you the question

  • Says XML is just verbose JSON with angle brackets
  • Claims JSON can express mixed content natively
  • Thinks JSON object members have a guaranteed order
  • Believes attributes and child elements are interchangeable in every case
  • Assumes key prefixes are enforced the way namespaces are