skip to content

A product rule says a display name may be at most 30 characters, and it has to be enforced in a browser form, in an API, and in storage. In JavaScript, which unit of "character" do you standardise on, and how do you keep the layers from disagreeing?

level: principalimportance: nice to knowfreq 18%

answer

  1. four numbers, one string
  2. whose expectation does the rule serve
  3. one implementation, server decides
  4. normalise before you measure
  5. fairness limit is not a resource bound

basics

~20 s

Pick grapheme clusters as the user-facing unit, normalise to NFC before counting, implement the rule in one shared function that every layer calls, and add a separate generous byte ceiling so unbounded combining marks cannot exhaust storage.

solid answer

~50 s

"Character" is ambiguous in JavaScript, and each layer will silently pick a different meaning if you let it: `.length` counts UTF-16 code units, `[...str].length` counts code points, `Intl.Segmenter` counts grapheme clusters, and a store may count UTF-8 bytes. Choose the unit by whose expectation the rule serves — for a user-facing limit that is grapheme clusters, because it is what the person typing sees. Then make the rule one function, exported once and called by the form, the API handler and any batch importer, so "30" cannot drift. Normalise to NFC before counting, since the same visible name can otherwise measure differently. Finally, treat the user-facing limit as a fairness rule, not a resource bound: a single grapheme can carry unlimited combining marks, so pair it with a hard byte cap and reject past it. The server is the authority; the browser check is UX.

code

javascript · 14 lines
javascript
const graphemes = new Intl.Segmenter('en', { granularity: 'grapheme' });

function measure(str) {
  const s = str.normalize('NFC');
  return {
    codeUnits: s.length,
    codePoints: [...s].length,
    graphemes: [...graphemes.segment(s)].length,
    utf8Bytes: new TextEncoder().encode(s).length,
  };
}

console.log(measure('Zoë 👍'));
// { codeUnits: 6, codePoints: 5, graphemes: 5, utf8Bytes: 9 }

go deeper

for a junior

Know that .length counts UTF-16 code units rather than characters, so a limit written as str.length <= 30 is not the same rule the label promises to the user.

for a middle

Be able to compute all four counts for the same string and explain why each differs, and why the grapheme count is the one that matches what a person typing sees.

for a senior

Show the enforcement design: one shared validation function called by client and server, normalisation before counting, the server as the authority, and an error message that names the unit and the actual count.

for a principal

Own the separation between a fairness rule and a resource bound — a grapheme limit for users, a hard byte cap for the system, checked in the cheap-first order — and treat retro-fitting the rule onto existing data as a product decision rather than a migration script.

## The rule is under-specified "30 characters" is not a technical statement. In JavaScript alone there are four defensible readings, and they disagree on real input: ```javascript const graphemes = new Intl.Segmenter('en', { granularity: 'grapheme' }); function measure(str) { const s = str.normalize('NFC'); return { codeUnits: s.length, codePoints: [...s].length, graphemes: [...graphemes.segment(s)].length, utf8Bytes: new TextEncoder().encode(s).length, }; } console.log(measure('Zoë 👍')); // { codeUnits: 6, codePoints: 5, graphemes: 5, utf8Bytes: 9 } ``` One string, four numbers. If the form validates one and the API validates another, some names are accepted by the UI and rejected by the server — the worst failure mode, because the user has no idea what they did wrong. ## Choosing the unit Ask whose expectation the rule exists to serve. - **Grapheme clusters** match what the person typing counts. For anything a human is told about — "at most 30 characters" in a label — this is the honest unit. `Intl.Segmenter` with `granularity: 'grapheme'` provides it. - **Code points** are a reasonable cheap approximation, and defensible if segmentation is unavailable somewhere in the stack. They over-count joined emoji and accented text. - **UTF-16 code units** (`.length`) are the tempting default precisely because they are free. They penalise non-Latin users arbitrarily: an emoji costs two, an astral CJK character costs two, plain ASCII costs one. A limit expressed this way is a limit that means something different depending on the user's language. - **UTF-8 bytes** are the right unit for a resource bound, not for a user-facing rule. Text-storage limits may be defined in bytes or in characters depending on the store, so check what yours actually enforces rather than assuming. The practical answer for a display name: graphemes for the product rule, bytes for the safety cap. ## Making the layers agree Agreement is an architecture problem, not a string problem. **One implementation.** Write the rule once — a shared `validateDisplayName(str)` used by the form, the request handler, and any importer or migration script. Three hand-written checks of "30" will drift, and the drift shows up as a support ticket rather than a test failure. **Normalise before measuring.** Count `str.normalize('NFC')`, and store what you counted. Otherwise a name typed on one input method and the same name pasted from elsewhere can measure differently, and normalising *after* the check lets a value cross the limit in storage. **Server decides.** The browser check exists to give fast feedback; it is not enforcement, since anyone can call the API directly. Duplicate the rule deliberately, and make the server's answer authoritative. **Report the unit in the error.** "Display name must be at most 30 characters (you used 34)" is actionable; a bare 400 is not. **Availability check.** `Intl.Segmenter` is part of ECMA-402 and is available in current browsers and in Node 16+, so a shared JavaScript implementation can genuinely run in both places. If some consumer of your API is not JavaScript, remember that it may count differently — which is another argument for the server owning the verdict. ## The limit is not a resource bound Grapheme clusters are unbounded. A single base letter can carry an arbitrarily long run of combining marks, so "30 graphemes" can be many kilobytes of text — a cheap way to inflate rows, logs and payloads. Every user-facing limit therefore needs a second, unglamorous ceiling: - a hard cap in UTF-8 bytes or code units, generous enough never to trouble a legitimate name; - **reject** past that cap rather than truncating, because truncating a name silently corrupts data; - apply the byte cap *before* segmenting, so you never run text segmentation over a hostile megabyte. Order matters: cheap byte check first, then normalise, then segment and count. ## Reject versus truncate For input, reject and explain. Truncation belongs in display code, where a shortened name with an ellipsis is a rendering decision and the full value survives in storage. Systems that truncate on ingest lose the original irreversibly, and the loss surfaces later as "why is this person's name cut off in the export". ## What a strong answer sounds like Name the ambiguity with concrete numbers, pick the unit by whose expectation the rule serves, make it one shared implementation with the server as authority, normalise before counting, and separate the fairness rule from the resource bound. The tell of a weaker answer is jumping straight to `.length < 30` without noticing that it encodes a different rule for every language.

  • What actually goes wrong if the browser uses .length and the server counts graphemes?
    The two disagree on exactly the users you least want to fail: names with emoji, accents or astral characters. Depending on which is stricter you either get a form that accepts what the API rejects — a dead-end with no useful message — or a client that blocks a name the server would have taken. Either way the rule is unpredictable per language.
  • Why not just enforce a byte limit everywhere and be done?
    Because it encodes a different rule for every script: 30 bytes is 30 Latin letters but roughly 10 CJK characters and about 7 emoji. As a fairness rule it is indefensible. Bytes are the right unit for a resource ceiling, which is why the two limits should coexist rather than one replacing the other.
  • Where does normalization fit into the enforcement order?
    Cheap byte cap first, so hostile input never reaches expensive work; then `normalize('NFC')`; then segment and count; then store exactly what you counted. Normalising after the check lets a value that passed grow or shrink in storage, and storing a different form from the one you measured makes the limit unverifiable later.
  • How would you roll this out over data that already violates the new rule?
    Enforce on write first and leave existing rows alone, so no one is locked out of editing an unrelated field. Measure how many stored values fail, and treat migration as a product decision — prompting the affected users to shorten their own name is almost always better than a script that truncates identities on their behalf.

saying these in an interview costs you the question

  • Uses .length and calls the rule done
  • Assumes every layer counts characters the same way
  • Enforces only in the browser form
  • Truncates user input silently instead of rejecting
  • Believes a grapheme limit bounds stored size

context