skip to content

schema.org data can be expressed as a JSON-LD script block, as microdata attributes (itemscope/itemtype/itemprop) on the visible markup, or as RDFa. Why do most teams choose JSON-LD, and what do they give up by doing so?

level: middleimportance: must knowfreq 58%

answer

  1. same vocabulary, different syntax
  2. microdata is glued to the DOM tree
  3. survives redesign and template churn
  4. one serializer, one block
  5. the cost is a second copy

basics

~20 s

JSON-LD keeps structured data in one contiguous block, decoupled from the DOM, so it survives template and markup refactors and can be generated straight from the server's data model. The cost is duplication: the data now exists twice and can drift from the visible page.

solid answer

~50 s

All three syntaxes express the same schema.org vocabulary, so the choice is ergonomic rather than semantic. Microdata and RDFa annotate the visible elements — `itemscope`/`itemtype`/`itemprop` for microdata, `vocab`/`typeof`/`property` for RDFa — which means the entity graph has to mirror the DOM tree, and any redesign that moves a `<span>` can silently break the markup. JSON-LD lives in one `application/ld+json` block, so the graph is independent of the layout, it can be serialized directly from the same object the template renders, and a components-based UI can emit it in one place. Search engines also state a preference for it. What you give up is co-location: the description is now a second copy of the facts, and nothing structurally stops it drifting from what the user sees — which is both a maintenance risk and a policy risk, since search engines require structured data to reflect visible content.

code

html · 8 lines
html
<!-- microdata: the graph must follow the DOM -->
<article itemscope itemtype="https://schema.org/Product">
  <h1 itemprop="name">Ceramic Mug</h1>
  <div itemprop="offers" itemscope itemtype="https://schema.org/Offer">
    <span itemprop="price">14.00</span>
    <meta itemprop="priceCurrency" content="EUR">
  </div>
</article>

go deeper

for a junior

Know that all three are ways of writing the same schema.org data, that JSON-LD sits in its own script block, and that microdata is written as attributes on the visible elements.

for a middle

Explain the decoupling: microdata forces the entity graph to mirror the DOM tree, JSON-LD does not, which is why JSON-LD survives redesigns and can be serialized from the server's data model.

for a senior

Show the tradeoff you accepted. Name the drift risk that comes with a second copy of the facts and describe the generation and validation you put in place so the block cannot quietly diverge from the page.

for a principal

Own the migration decision for an existing site: what a mixed-syntax page costs, whether any downstream consumer still needs inline markup, and who is accountable for the structured data being true.

## Three syntaxes, one vocabulary schema.org is a *vocabulary* — a set of types and properties. It says nothing about syntax. Three syntaxes carry it on the web: **Microdata** annotates existing elements with HTML attributes: ```html <div itemscope itemtype="https://schema.org/Product"> <h1 itemprop="name">Ceramic Mug</h1> <span itemprop="sku">MUG-01</span> <div itemprop="offers" itemscope itemtype="https://schema.org/Offer"> <span itemprop="price">14.00</span> <meta itemprop="priceCurrency" content="EUR"> </div> </div> ``` `itemscope` opens a new entity, `itemtype` says what kind, `itemprop` attaches a property, and `itemid` supplies a global identifier. `itemref` exists to pull in properties from elsewhere in the document when the DOM will not cooperate. **RDFa** (in practice RDFa Lite) does the same job with `vocab`, `typeof`, `property` and `resource`. It comes from the linked-data world and is common in CMS themes and government publishing. **JSON-LD** puts the whole graph in one inert `<script type="application/ld+json">` block, as covered above, using `@context`, `@type` and `@id`. ## Why JSON-LD usually wins **The graph does not have to match the DOM.** In microdata an entity's properties must sit inside the element that opened the `itemscope` — or you reach for `itemref`, which is fragile. Real designs do not nest the way a data model does: the price is in a sticky bar, the brand is in a header, the rating is in a tab. JSON-LD has no such constraint; you describe the entities in the shape the vocabulary wants and the layout goes wherever design says. **It survives refactors.** Microdata attributes are scattered across dozens of elements in a template. A redesign that swaps a wrapper, a component library upgrade, or a well-meant cleanup of "unused attributes" quietly destroys the markup, and nothing visible breaks to tell you. A JSON-LD block is one artifact you can see, diff and test. **It is generated, not hand-annotated.** The server already has the product object. A JSON-LD block is one serializer function away from that object, so the same fetch that renders the page emits the description. That is much harder when the properties must be sprinkled across the render tree. **It composes and can be injected.** One page can carry several blocks, so a breadcrumb component can emit its own `BreadcrumbList` without knowing about the `Product`. **Search engines recommend it.** Google's structured-data documentation names JSON-LD as the preferred format while continuing to parse microdata and RDFa. ## What you actually give up **Co-location is a feature you are trading away.** In microdata, the marked-up value *is* the value the user reads: if the price changes, the marked-up price changes with it because they are the same text node. In JSON-LD you now maintain a second representation. A promotional price rendered by a client-side component while the JSON-LD still carries the list price is the classic bug — invisible to QA, visible to a crawler. This is not only a tidiness problem. Search engines require structured data to describe content that is present on the page. Persistent mismatches — a rating nobody can find, a price the page does not show — are a policy violation and can attract a manual action instead of a rich result. **Some consumers still expect inline markup.** Certain CMSs, aggregators and older tooling read microdata or RDFa out of the rendered HTML. If a specific integration demands it, check before switching. **Mixing syntaxes is legal but confusing.** A page may carry both; consumers merge what they find. Two sources of truth for the same entity means two ways to be wrong, so pick one per entity. ## Answering it well The strong answer names the mechanism, not the fashion. "JSON-LD because Google prefers it" is thin. "Because the entity graph is decoupled from the DOM, so it survives redesigns and can be serialized from the same object the page renders — at the cost of a second copy of the facts that we guard by generating both from one source" shows you have maintained it. ```html <script type="application/ld+json"> { "@context": "https://schema.org", "@type": "Product", "name": "Ceramic Mug", "sku": "MUG-01", "offers": { "@type": "Offer", "price": "14.00", "priceCurrency": "EUR", "availability": "https://schema.org/InStock" } } </script> ``` The same facts as the microdata above, with the layout free to change underneath.

  • If JSON-LD can drift from the page, how do you stop that in practice?
    Generate both from one source. The route handler or template already holds the product object; render the visible markup and serialize the JSON-LD from that same object rather than hand-writing a parallel block. Then add a check — a snapshot test or a validator run in CI — so a template change that empties a required property fails the build instead of shipping.
  • Is it a problem if a page carries both microdata and a JSON-LD block?
    It is legal, and consumers merge what they find, but it gives one entity two sources of truth. Once they disagree you cannot predict which wins, and the mismatch is exactly what structured-data policy penalises. Pick one syntax per entity; carrying microdata from an old theme alongside new JSON-LD is a migration to finish, not a state to keep.
  • Which microdata attribute exists specifically because the DOM does not match the data model?
    `itemref`. It takes a list of element ids whose properties should be folded into the current item, letting you claim properties that live outside the `itemscope` element. It works, but it couples your structured data to ids scattered across the page — the fragility that pushes teams to JSON-LD in the first place.

saying these in an interview costs you the question

  • Thinks JSON-LD and microdata express different vocabularies
  • Claims microdata is deprecated or unsupported
  • Says JSON-LD is preferred without naming a mechanism
  • Ignores that JSON-LD duplicates the visible facts
  • Confuses microdata itemprop with RDFa property

context