In HTML, what does the data-* attribute family give you that inventing your own attribute such as userid does not?
answer
- a prefix reserved for your own data
- tolerated is not the same as valid
- the language keeps adding attribute names
- dataset only sees the data- prefix
- strings only, and visible to everyone
basics
~20 sdata-* is HTML's sanctioned slot for custom data: it is valid markup, it can never collide with an attribute the language adds later, and browsers expose it to scripts through element.dataset. An invented attribute such as userid gets none of those guarantees.
solid answer
~50 sHTML is a closed vocabulary, so an attribute the specification does not define is non-conforming. Browsers are forgiving, so `<div userid="42">` parses and you can still read it with `getAttribute`, but it fails validation, it tells a reader nothing, and you are betting that HTML will never give that name a meaning. `data-*` exists to end that bet: any attribute whose name starts with `data-` is legal on every element, is guaranteed to stay meaningless to the language, and is surfaced on `element.dataset`, where `data-user-id` becomes `dataset.userId`. HTML keeps growing new attributes — `loading`, `inert`, `enterkeyhint` and `popover` all became real — so the prefix is genuine collision insurance. The value is always a string, it is plainly visible in view-source, and it carries no semantics, so data-* is for application data, never for accessibility meaning.
go deeper
Be able to say that data-* is the legal way to attach your own data to an element, that it works on any element, and that you read it back with element.dataset or getAttribute.
Explain the naming rules and the guarantee behind the prefix: the language promises never to define a data- name, so it cannot collide with future HTML attributes the way an invented name can.
Show judgment about what belongs in markup at all — small hooks yes, duplicated application state or secrets no — and point out that data-* carries zero semantics for assistive technology.
Own the convention across a codebase: naming schemes that survive component renames, a rule against markup becoming a second source of truth, and lint or conformance checking that catches invented attributes in review.
## HTML is a fixed vocabulary Every element and attribute name in HTML is defined by the specification. A name the spec does not define is *non-conforming* — invalid markup. Browsers nevertheless use a permissive error-recovery model: an unknown attribute is parsed, kept on the element, and simply carries no meaning. That gap between "invalid" and "works anyway" is why teams once shipped attributes like `userid`, `role-type` or `state` straight onto elements. The problem is not that the browser complains today. It is that you have taken a name out of a namespace you do not own. HTML has added many attributes since HTML5 shipped — `loading`, `decoding`, `fetchpriority`, `inert`, `enterkeyhint`, `popover`. Any invented name can one day become a real one with real behaviour, and then your markup starts doing something you never asked for. ## The reserved namespace HTML5 solved this by reserving a whole prefix. An attribute whose name starts with `data-` is valid on **any** element, and the language promises to never give such a name meaning. ```html <article id="post-42" data-post-id="42" data-author-handle="ada" data-published> </article> ``` The naming rules are short: the name must begin with `data-` and have at least one character after it, and it must contain no uppercase ASCII letters. That last rule is mostly self-enforcing, because the HTML parser lowercases attribute names anyway — writing `data-postId` in the source produces the attribute `data-postid`. ## What you actually gain 1. **Validity.** The document passes a conformance checker, so real errors are not buried under noise. 2. **Collision safety.** The prefix is yours forever. 3. **A first-class read path.** `element.dataset` is a live map of the element's data-* attributes, and `data-post-id` is reachable as `dataset.postId`. The lower-level `getAttribute("data-post-id")` works too. 4. **Readability.** A reviewer sees `data-` and instantly knows this is application state, not browser behaviour. Attribute selectors such as `[data-post-id]` also match these attributes, which is how declarative state often gets wired to presentation. ## What data-* is not for - **Not semantics.** `data-role="button"` tells assistive technology nothing; only a real `<button>` or a real ARIA `role` does. data-* is invisible to the accessibility tree. - **Not private.** The value ships in the HTML and is readable by anyone, so no secrets, tokens or prices you do not want tampered with. - **Not a database.** Values are strings, so a large JSON blob parked in an attribute inflates the document and drifts from the real state you already hold in memory. - **Not typed.** `data-open="false"` reads back as the string `"false"`, which is truthy — a classic bug. Presence-style flags (`data-published` with no value, tested with `hasAttribute`) are safer than string booleans. ## The honest summary An invented attribute is a private convention that browsers tolerate. `data-*` is a public contract the language guarantees. The cost of using the prefix is five characters; the cost of not using it is an invalid document and a name you do not own.
- If browsers keep an invented attribute in the DOM anyway, what concrete harm has it caused?Two kinds. First, a future spec can claim the name and attach behaviour to it, so your markup silently starts doing something new. Second, the document stops validating, which buries genuine conformance errors in noise and confuses tooling and reviewers. You also lose `dataset`, since it only exposes `data-` prefixed attributes.
- When should custom state live somewhere other than a data-* attribute?When it is large, secret, or already owned by your application layer. Attributes are strings that ship in the HTML, so a big JSON payload bloats the document and duplicates state you must then keep in sync. Keep the truth in your app state and use data-* for the small hooks markup genuinely needs.
- Does adding a data-* attribute change anything about how the element is rendered or announced?No. It has no rendering effect and no accessibility effect at all — it is inert metadata. Meaning still has to come from the element you chose and from real ARIA attributes. If a screen reader must know something, `data-*` is the wrong place to put it.
saying these in an interview costs you the question
- Invented attributes are fine because browsers ignore them
- data-* attributes are hidden from the user
- data-* can convey a role to screen readers
- Stores auth tokens or prices in data-*
- Expects data-* to hold numbers or objects