You lead frontend for a large catalogue site where each template's JSON-LD was hand-written by whoever built it. How do you decide what to mark up, and how do you keep the structured data trustworthy as the site changes?
answer
- scope by feature, not by vocabulary
- eligibility is not traffic
- one serializer, one source of truth
- identifiers from shared helpers
- silent failure needs monitoring
basics
~20 sMark up only the entity types tied to a rich result you actually want, generate every block from the same data that renders the visible page, and treat the output as a tested artifact monitored in Search Console — not as documentation to hand-maintain per template.
solid answer
~50 sStart by cutting scope: structured data buys *eligibility* for specific search features, so mark up the types where a feature exists and matters to the business — products, breadcrumbs, articles, the organization — and stop. Marking up everything schema.org can express costs maintenance and buys nothing. Then remove the hand-writing: give each entity one serializer that takes the same object the template renders, so the JSON-LD cannot say a price the page does not show. Mint `@id` values from shared helpers so the whole site agrees on one identity per entity. Make correctness testable — a unit test per template asserting the block parses and carries the feature's required properties — and monitor at the fleet level with Search Console's enhancement and unparsable-data reports rather than spot-checking URLs. Finally, set expectations: features get narrowed and retired, so treat this as maintained infrastructure with an owner, not a one-off SEO project.
code
javascript · 20 linesconst ORG_ID = "https://example.com/#organization";
const productId = (sku) => `https://example.com/p/${sku}#product`;
function productJsonLd(product) {
return {
"@context": "https://schema.org",
"@type": "Product",
"@id": productId(product.sku),
name: product.title,
brand: { "@id": ORG_ID },
offers: {
"@type": "Offer",
price: product.displayPrice.amount,
priceCurrency: product.displayPrice.currency,
availability: product.inStock
? "https://schema.org/InStock"
: "https://schema.org/OutOfStock"
}
};
}go deeper
Know that structured data should describe what the page really shows, and that copying a generator's output into each template by hand is how it goes stale.
Be able to describe generating the block from the same object the template renders, and using shared helpers for @id values so identities stay consistent across pages.
Show the safety net: template tests that assert the block parses and carries required properties, plus Search Console enhancement reports read at fleet level, because a broken block never breaks the page.
Own the scope and the expectation. Justify which entity types earn markup by naming the feature each unlocks, name an owner for the schema, and state clearly that the investment buys eligibility and policy safety, not guaranteed traffic.
## The decision is scope before syntax The failure mode at scale is not bad JSON. It is a hundred templates each carrying a slightly different, slowly rotting description of the same company, authored by people who are gone. The lead's job is to decide how much of this the organisation should own at all. The honest framing: structured data makes a page **eligible** for a specific search feature. If no feature consumes a type, marking it up produces no visible outcome and still costs review time on every template change. So the scope test is per type: *which search feature does this unlock, and do we want it?* On a catalogue site that usually lands on `Product` with `Offer`, `BreadcrumbList`, `Organization` for the entity behind the site, and `Article`/`BlogPosting` for editorial. Everything else is speculative. Be candid about volatility too. Search features are narrowed and retired — FAQ rich results, for instance, were restricted to a narrow class of authoritative sites, and every page that had been marked up for them lost the result overnight. A portfolio built on features that may vanish should be small and cheap to maintain. ## One source of truth, mechanically enforced Hand-written JSON-LD is a second copy of the facts, and a second copy always drifts. The structural fix is a serializer per entity type that takes the domain object the page already renders: ```javascript // one function, used by every template that shows a product function productJsonLd(product) { return { "@context": "https://schema.org", "@type": "Product", "@id": productId(product.sku), name: product.title, image: product.images.map(absoluteUrl), brand: { "@id": ORG_ID }, offers: { "@type": "Offer", price: product.displayPrice.amount, // the price the page shows priceCurrency: product.displayPrice.currency, availability: product.inStock ? "https://schema.org/InStock" : "https://schema.org/OutOfStock" } }; } ``` Two things matter here beyond tidiness. First, `product.displayPrice` is deliberately the same field the UI renders — the drift bug that gets sites penalised is a JSON-LD list price beside a promotional price on screen. Second, `productId()` and `ORG_ID` come from shared helpers, so every template mints the same identity and node references resolve instead of quietly splitting one company into several entities. ## Make it testable, then monitor it Structured data fails silently: the page renders perfectly with a broken block. That rules out "we'll notice". Two layers of safety net: **In CI.** A test per template that renders it, extracts the `application/ld+json` blocks, asserts the JSON parses, and asserts the required properties for the feature are present and non-empty. This is cheap and catches the two most common regressions — a refactor that empties a field, and an escaping bug in user-supplied text. **In production.** Search Console reports at the fleet level: enhancement reports per detected type with valid/warning/error counts, the unparsable-structured-data report, and the manual-actions section. A deploy that breaks a template shows up as a step change in error count across thousands of URLs, which no spot-check would find. Give someone the job of looking. ## The organisational half **Ownership.** Structured data spans SEO, frontend and the catalogue team, which is how it ends up owned by nobody. Name an owner for the schema — the entity types, their identifiers, their serializers — and require changes to go through it rather than being edited inline in a template. **Migration.** Legacy microdata scattered through old templates should be finished, not left beside new JSON-LD. Two sources of truth for one entity is a mismatch waiting to be found by a crawler, not a belt-and-braces safety measure. **Expectation setting.** The business will ask what the traffic gain is. The truthful answer is that eligibility is a precondition, not a lever: correct markup lets a richer result be shown, and whether it is shown depends on decisions you do not control. Promise the precondition and the absence of policy risk; do not promise the click-through. ## What separates a principal answer A senior answer fixes the templates. A principal answer decides which markup deserves to exist, makes correctness a property of the build rather than of anyone's diligence, assigns ownership, and states plainly what the investment can and cannot buy — including the possibility that the right call for a rarely-consumed type is to delete it.
- A stakeholder asks you to mark up every type schema.org offers, on the theory that more data cannot hurt. How do you answer?More data does hurt: every property is a claim you must keep true across redesigns, and claims that drift from the visible page are a policy risk rather than a neutral extra. Mark up what a search feature you want actually consumes; keep the rest out. The cost is real and recurring, the benefit for unconsumed types is zero.
- How would you handle legacy microdata that is still scattered through old templates?Treat it as a migration to finish. While both exist for one entity you have two sources of truth and no way to predict which a consumer honours, so a stale microdata price can contradict a correct JSON-LD one. Remove the attributes template by template as you touch them, verifying with the Rich Results Test on representative URLs as you go.
- What single metric would you watch to know the structured data is healthy across thousands of URLs?The error and warning counts per detected item type in Search Console's enhancement reports, tracked over time rather than read once. A deploy that breaks a template produces a step change across the whole type, which is visible there and invisible in any spot-check. Pair it with the unparsable-structured-data report to catch escaping regressions.
saying these in an interview costs you the question
- Marks up every schema.org type available
- Promises a ranking or traffic increase from markup
- Leaves the JSON-LD hand-written per template
- Keeps legacy microdata alongside new JSON-LD
- Treats it as a one-off project with no owner