skip to content

In HTML, how do data-* attribute names map onto keys of an element's dataset property, and what type do you get back when you read one?

level: middleimportance: should knowfreq 52%

answer

  1. prefix off, hyphen becomes a capital
  2. the map is live in both directions
  3. DOMStringMap is named for its value type
  4. source attribute names get lowercased
  5. hyphenated keys are rejected outright

basics

~10 s

dataset drops the data- prefix and converts each hyphen-plus-lowercase-letter into a capital, so data-user-id is read as dataset.userId. Every value read back is a string, never a number, boolean or object.

solid answer

~40 s

`element.dataset` is a live view of just the element's `data-*` attributes. The mapping strips `data-`, then removes each hyphen that is followed by an ASCII lowercase letter and uppercases that letter: `data-user-id` becomes `dataset.userId`, `data-role` becomes `dataset.role`. It works in reverse too — assigning `dataset.userName = "ada"` creates the attribute `data-user-name`. Two things bite people. First, values are always strings, so `dataset.count + 1` concatenates rather than adds; you need `Number(...)` or `parseInt`. Second, the key is the camelCase form, not the attribute name: `dataset["user-id"] = "7"` throws a `SyntaxError` because a hyphen followed by a lowercase letter is illegal in a key. Reading a missing key gives `undefined`, and `delete dataset.role` removes the attribute itself.

go deeper

for a junior

Know that data-user-id is read as dataset.userId and that the value comes back as a string, so convert with Number before doing arithmetic on it.

for a middle

Explain the transformation in both directions, including that assigning dataset.userName creates data-user-name, and name the SyntaxError you get from a hyphenated key.

for a senior

Show that you have debugged the real failures: a lowercased source name, a truthy "false", JSON parsed on every read, and know when setAttribute is the clearer tool.

for a principal

Set the house convention — kebab-case attributes, presence-flags instead of string booleans, no serialized state in markup — and be able to justify why markup should not become a second source of truth.

## What dataset is `element.dataset` is a `DOMStringMap` — an object-like view over exactly those attributes on the element whose names start with `data-`. It is live: changing an attribute changes what `dataset` reports, and writing through `dataset` writes a real attribute back into the DOM that you can see in devtools. ```html <li id="row" data-item-id="42" data-item-count="5" data-status="open"></li> ``` ```js const el = document.getElementById("row"); el.dataset.itemId; // "42" el.dataset.itemCount; // "5" el.dataset.status; // "open" el.dataset.missing; // undefined ``` ## The name transformation, in both directions Attribute to key: remove the `data-` prefix, then for every `-` followed by an ASCII lowercase letter, delete the hyphen and uppercase the letter. So `data-item-id` becomes `itemId` and `data-a-b-c` becomes `aBC`. Key to attribute: the reverse. Every uppercase ASCII letter becomes a hyphen plus the lowercase letter, and `data-` is prefixed. Assigning `el.dataset.userName = "ada"` produces `data-user-name="ada"` in the markup — **not** `data-userName`. Two consequences catch people: - **Hyphenated keys are illegal.** `el.dataset["user-id"] = "7"` throws a `SyntaxError`, because the setter rejects any name containing a hyphen followed by a lowercase letter. Use the camelCase key, or drop to `setAttribute("data-user-id", "7")`. - **Source case is lost.** HTML attribute names are lowercased by the parser, so `data-userId` in your markup is really `data-userid`, which is `dataset.userid` — not `dataset.userId`. This is a frequent "why is it undefined" bug. ## Everything is a string `DOMStringMap` is named for exactly this reason. There is no type coercion anywhere in the pipeline: ```js el.dataset.itemCount + 1; // "51" (string concatenation) Number(el.dataset.itemCount) + 1; // 6 ``` The same applies to booleans. `data-active="false"` reads as `"false"`, and every non-empty string is truthy, so `if (el.dataset.active)` is true either way. Two safer patterns: compare explicitly (`el.dataset.active === "true"`), or use presence as the flag — write `data-active` with no value and test `el.hasAttribute("data-active")`. A valueless attribute reads back as the empty string `""`, which is falsy, so even a truthiness test happens to behave for that form. Structured data has the same problem: if you store JSON you must `JSON.parse` it on every read, and any parse failure is a runtime error in the middle of rendering. ## Writing and deleting ```js el.dataset.status = "closed"; // sets data-status="closed" delete el.dataset.status; // removes the attribute entirely ``` Setting a key to `null` or a number does not skip the string conversion — the value is stringified, so `el.dataset.count = 5` produces `data-count="5"`. ## dataset versus getAttribute `getAttribute("data-item-id")` and `dataset.itemId` read the same attribute; the difference is ergonomics and failure mode. `getAttribute` returns `null` for a missing attribute, `dataset` returns `undefined`, and `getAttribute` takes the literal attribute name so it sidesteps the camelCase rules entirely. When a name is computed at runtime, `setAttribute`/`getAttribute` with a template string is usually clearer than trying to build a camelCase key. ## Enumerating `dataset` is enumerable, so `Object.keys(el.dataset)` or `for...in` lists every data-* key on that element — useful when a component reads a bag of options off its own root element. Note that it reflects only that element's own attributes; there is no inheritance from ancestors.

  • A colleague writes data-userId="7" in the HTML and reads el.dataset.userId, which is undefined. What happened?
    The HTML parser lowercases attribute names, so the markup really produced `data-userid`. That maps to the key `userid`, not `userId`, so the camelCase lookup misses. The fix is to write the attribute in kebab-case — `data-user-id` — which maps to `dataset.userId` as intended. Uppercase letters in a data-* name are non-conforming for exactly this reason.
  • When would you prefer getAttribute over dataset for reading a data-* attribute?
    When the name is computed at runtime, or when you want the literal attribute name in the code. `getAttribute("data-" + key)` is direct, whereas building the camelCase key means reimplementing the transform. `getAttribute` also returns `null` rather than `undefined` for a missing attribute, which some codebases prefer for consistency with other attribute reads.
  • How do you model a boolean flag in a data-* attribute without the "false" is truthy trap?
    Use presence rather than value: write `data-active` with no value and test `el.hasAttribute("data-active")`, removing the attribute to turn it off. If you must keep a value, compare explicitly with `=== "true"`. Never rely on plain truthiness, since `"false"` is a non-empty string and therefore truthy.

saying these in an interview costs you the question

  • Expects dataset to return numbers for numeric-looking values
  • Writes data-userId and expects dataset.userId
  • Uses dataset["user-id"] as the key
  • Thinks dataset copies values instead of reading attributes
  • Assumes dataset inherits data-* from ancestor elements

context