skip to content

Global Attributes and data-*

The attributes legal on every element — id, class, lang, dir, hidden, inert, data-* — and what each actually does to rendering, interaction, or assistive tech. Interviewers probe data-* and the dataset API to see if you know the sanctioned way to attach custom state to markup.

part ofHTMLoverview, primer and where to startread it →
on this pageshow

questions

6

In HTML, what does the data-* attribute family give you that inventing your own attribute such as userid does not?

level: juniorimportance: must knowfreq 62%

answer

  1. a prefix reserved for your own data
  2. tolerated is not the same as valid
  3. the language keeps adding attribute names
  4. dataset only sees the data- prefix
  5. strings only, and visible to everyone

basics

~20 s

data-* is HTML's sanctioned slot for custom data: it is valid markup, it can never collide with an attribute the language adds later, and browsers expose it to scripts through element.dataset. An invented attribute such as userid gets none of those guarantees.

solid answer

~50 s

HTML is a closed vocabulary, so an attribute the specification does not define is non-conforming. Browsers are forgiving, so `<div userid="42">` parses and you can still read it with `getAttribute`, but it fails validation, it tells a reader nothing, and you are betting that HTML will never give that name a meaning. `data-*` exists to end that bet: any attribute whose name starts with `data-` is legal on every element, is guaranteed to stay meaningless to the language, and is surfaced on `element.dataset`, where `data-user-id` becomes `dataset.userId`. HTML keeps growing new attributes — `loading`, `inert`, `enterkeyhint` and `popover` all became real — so the prefix is genuine collision insurance. The value is always a string, it is plainly visible in view-source, and it carries no semantics, so data-* is for application data, never for accessibility meaning.

go deeper

for a junior

Be able to say that data-* is the legal way to attach your own data to an element, that it works on any element, and that you read it back with element.dataset or getAttribute.

for a middle

Explain the naming rules and the guarantee behind the prefix: the language promises never to define a data- name, so it cannot collide with future HTML attributes the way an invented name can.

for a senior

Show judgment about what belongs in markup at all — small hooks yes, duplicated application state or secrets no — and point out that data-* carries zero semantics for assistive technology.

for a principal

Own the convention across a codebase: naming schemes that survive component renames, a rule against markup becoming a second source of truth, and lint or conformance checking that catches invented attributes in review.

## HTML is a fixed vocabulary Every element and attribute name in HTML is defined by the specification. A name the spec does not define is *non-conforming* — invalid markup. Browsers nevertheless use a permissive error-recovery model: an unknown attribute is parsed, kept on the element, and simply carries no meaning. That gap between "invalid" and "works anyway" is why teams once shipped attributes like `userid`, `role-type` or `state` straight onto elements. The problem is not that the browser complains today. It is that you have taken a name out of a namespace you do not own. HTML has added many attributes since HTML5 shipped — `loading`, `decoding`, `fetchpriority`, `inert`, `enterkeyhint`, `popover`. Any invented name can one day become a real one with real behaviour, and then your markup starts doing something you never asked for. ## The reserved namespace HTML5 solved this by reserving a whole prefix. An attribute whose name starts with `data-` is valid on **any** element, and the language promises to never give such a name meaning. ```html <article id="post-42" data-post-id="42" data-author-handle="ada" data-published> </article> ``` The naming rules are short: the name must begin with `data-` and have at least one character after it, and it must contain no uppercase ASCII letters. That last rule is mostly self-enforcing, because the HTML parser lowercases attribute names anyway — writing `data-postId` in the source produces the attribute `data-postid`. ## What you actually gain 1. **Validity.** The document passes a conformance checker, so real errors are not buried under noise. 2. **Collision safety.** The prefix is yours forever. 3. **A first-class read path.** `element.dataset` is a live map of the element's data-* attributes, and `data-post-id` is reachable as `dataset.postId`. The lower-level `getAttribute("data-post-id")` works too. 4. **Readability.** A reviewer sees `data-` and instantly knows this is application state, not browser behaviour. Attribute selectors such as `[data-post-id]` also match these attributes, which is how declarative state often gets wired to presentation. ## What data-* is not for - **Not semantics.** `data-role="button"` tells assistive technology nothing; only a real `<button>` or a real ARIA `role` does. data-* is invisible to the accessibility tree. - **Not private.** The value ships in the HTML and is readable by anyone, so no secrets, tokens or prices you do not want tampered with. - **Not a database.** Values are strings, so a large JSON blob parked in an attribute inflates the document and drifts from the real state you already hold in memory. - **Not typed.** `data-open="false"` reads back as the string `"false"`, which is truthy — a classic bug. Presence-style flags (`data-published` with no value, tested with `hasAttribute`) are safer than string booleans. ## The honest summary An invented attribute is a private convention that browsers tolerate. `data-*` is a public contract the language guarantees. The cost of using the prefix is five characters; the cost of not using it is an invalid document and a name you do not own.

  • If browsers keep an invented attribute in the DOM anyway, what concrete harm has it caused?
    Two kinds. First, a future spec can claim the name and attach behaviour to it, so your markup silently starts doing something new. Second, the document stops validating, which buries genuine conformance errors in noise and confuses tooling and reviewers. You also lose `dataset`, since it only exposes `data-` prefixed attributes.
  • When should custom state live somewhere other than a data-* attribute?
    When it is large, secret, or already owned by your application layer. Attributes are strings that ship in the HTML, so a big JSON payload bloats the document and duplicates state you must then keep in sync. Keep the truth in your app state and use data-* for the small hooks markup genuinely needs.
  • Does adding a data-* attribute change anything about how the element is rendered or announced?
    No. It has no rendering effect and no accessibility effect at all — it is inert metadata. Meaning still has to come from the element you chose and from real ARIA attributes. If a screen reader must know something, `data-*` is the wrong place to put it.

saying these in an interview costs you the question

  • Invented attributes are fine because browsers ignore them
  • data-* attributes are hidden from the user
  • data-* can convey a role to screen readers
  • Stores auth tokens or prices in data-*
  • Expects data-* to hold numbers or objects

context

open as a page

What does the lang attribute on the <html> element actually change in a browser, and when should you set lang on an element deeper in the page?

level: middleimportance: must knowfreq 54%

basics

~20 s

lang declares the document's natural language with a BCP 47 tag, and it drives screen-reader pronunciation and voice choice, hyphenation and line breaking, font and glyph selection, the spellcheck dictionary, and translation offers. Set it again on any element whose content is in a different language.

open as a page

In HTML, how do data-* attribute names map onto keys of an element's dataset property, and what type do you get back when you read one?

level: middleimportance: should knowfreq 52%

basics

~10 s

dataset drops the data- prefix and converts each hyphen-plus-lowercase-letter into a capital, so data-user-id is read as dataset.userId. Every value read back is a string, never a number, boolean or object.

open as a page

A team hides table rows with the HTML hidden attribute, but rows carrying a class whose stylesheet rule sets display: flex stay visible. Why does hidden stop working, and what is the fix?

level: middleimportance: should knowfreq 42%

basics

~20 s

The hidden attribute is not a rendering switch — browsers implement it with a user-agent stylesheet rule that sets display: none. Any author rule setting display on the same element outranks the user-agent origin and un-hides it. The fix is a stronger author rule or a state class that controls display.

open as a page

HTML requires the id attribute to be unique within a document. What actually breaks when two elements end up with the same id?

level: middleimportance: should knowfreq 46%

basics

~20 s

Nothing throws — every lookup that resolves an id to one element silently takes the first in tree order, so getElementById, fragment links and attributes that reference an id all point at the first match while selector matching still matches all of them, which is what hides the bug.

open as a page

What does the HTML inert attribute do to the subtree it is set on, and which bug does it fix that hiding the background visually or with aria-hidden does not?

level: seniorimportance: should knowfreq 38%

basics

~20 s

inert makes an element and everything inside it non-interactive: the subtree cannot be focused by Tab or programmatically, does not respond to clicks, is not selectable or findable by find-in-page, and is hidden from assistive technology. It fixes background content behind a modal that is dimmed but still reachable by keyboard.

open as a page