In a Go signup validator, what do unicode.IsLetter, IsDigit and IsSpace actually accept?
answer
- they answer Unicode, not ASCII
- one letter test, every script
- the digit test is category Nd only
- a digit that Atoi will refuse
- zero-width space is not white space
basics
~20 sThey test Unicode categories over a whole rune, not ASCII. IsLetter accepts letters from every script, IsDigit accepts only decimal digits in category Nd including Arabic-Indic ones, and IsSpace accepts the Unicode white-space set including the no-break space.
solid answer
~40 sAll three take a `rune` and answer a Unicode question, not an ASCII one. `unicode.IsLetter` is true for any rune in the L categories — Latin, Han, Cyrillic, Devanagari, everything — so "letters only" is a far wider rule than most validators intend. `unicode.IsDigit` is narrower than people expect: it is category Nd only, so Arabic-Indic ٣ passes even though `strconv.Atoi` will refuse it, while a Roman numeral Ⅰ fails; `unicode.IsNumber` is the broader N-category test. `unicode.IsSpace` follows Unicode's White_Space property, which includes U+00A0 NO-BREAK SPACE and the ideographic space but **not** U+200B ZERO WIDTH SPACE or U+FEFF, so a name can look padded and still pass a trim. For a real validator I would combine these with `unicode.Is` against specific script tables, and add explicit rejections for marks and control characters.
code
go · 4 linesfmt.Println(unicode.IsDigit('\u0663')) // true: ARABIC-INDIC DIGIT THREE, category Nd
fmt.Println(unicode.IsLetter('\u65e5')) // true: a Han character is a letter
fmt.Println(unicode.IsSpace('\u00a0')) // true: NO-BREAK SPACE is white space
fmt.Println(unicode.IsSpace('\u200b')) // false: ZERO WIDTH SPACE is a format charactergo deeper
Recall that these predicates take a single rune and answer questions about Unicode categories, so they accept far more than the ASCII characters you might picture.
Name the categories behind each one, and give a concrete divergence such as an Arabic-Indic digit that passes the digit test but cannot be parsed as a number.
Show how you would compose them into a real validator: an allowlist, explicit rejection of control and format characters, and awareness that zero-width characters defeat a naive space check.
Own the tradeoff between abuse prevention and exclusion. Every tightened rule rejects somebody's real name, so decide who signs off on the character policy and how you measure the false rejections it causes.
## What these functions are The `unicode` package exposes the Unicode character database as predicates and range tables. The three you meet first are: ``` func IsLetter(r rune) bool func IsDigit(r rune) bool func IsSpace(r rune) bool ``` Each takes a **rune** — a single code point — so they compose naturally with rune-wise iteration over a string. They are lookups into compiled range tables; they are not locale-aware and take no configuration. ## IsLetter is much wider than "a-zA-Z" `unicode.IsLetter` reports whether a rune is in one of Unicode's letter categories: uppercase, lowercase, titlecase, modifier and other letters. That means Han, Cyrillic, Greek, Arabic, Devanagari, Hangul and thousands more all pass. For an international product that is usually right, and it is why the predicate exists. But it is not a security boundary: "letters only" still permits mixed-script names where a Cyrillic а sits where a Latin a is expected, which is the classic look-alike impersonation shape. If you care about that, test membership in a specific script table instead, with `unicode.Is(unicode.Latin, r)` or `unicode.In(r, unicode.Latin, unicode.Han)`. Related predicates you will want alongside it: - `unicode.IsMark` — combining marks. A name made of a base letter plus forty stacked marks is all "letters and marks" and renders as a vertical smear across other people's screens. - `unicode.IsControl` and `unicode.IsPrint` — the blunt instruments for rejecting things that should never be in a display name. - `unicode.IsPunct`, `unicode.IsSymbol` — where emoji actually live (they are symbols, not letters). ## IsDigit is narrower than people expect `unicode.IsDigit` is true only for category **Nd**, decimal digit number. So: - ASCII `'7'` passes. - ARABIC-INDIC DIGIT THREE U+0663 passes — it is Nd — even though parsing it with `strconv.Atoi` fails. A validator built on `IsDigit` and a parser built on `Atoi` therefore disagree, and the gap between them is a rejected-after-accepted bug. - ROMAN NUMERAL ONE U+2160 fails; it is Nl, a letter number. - SUPERSCRIPT TWO fails; it is No, other number. `unicode.IsNumber` covers all of N (Nd, Nl and No) and is broader still. If what you mean is "an ASCII digit", say `r >= '0' && r <= '9'` and be explicit — that is clearer than a Unicode predicate that quietly means something else. ## IsSpace follows the White_Space property `unicode.IsSpace` reports true for the characters Unicode marks as white space: the ASCII controls tab, newline, vertical tab, form feed and carriage return, the space, U+0085 NEL, U+00A0 NO-BREAK SPACE, and the wider set including the en/em spaces and U+3000 IDEOGRAPHIC SPACE. What it does **not** include is the trap: U+200B ZERO WIDTH SPACE and U+FEFF are format characters, not white space. A display name padded with zero-width spaces passes a space check, passes a trim, renders as if it were shorter than it is, and can be used to create a name visually identical to someone else's. Rejecting them needs an explicit rule — category Cf, or an allowlist that simply does not include them. ## Putting them together in a validator The predicates are the primitives; the policy is yours. A workable shape is an allowlist rather than a denylist: 1. Reject the string outright if it is not valid UTF-8. 2. Iterate runes, and accept only those in an explicit set: letters, decimal digits, a small punctuation set, and the plain space. 3. Reject control, format and unassigned characters explicitly rather than relying on them failing the letter test. 4. Bound the count of combining marks per base character. 5. Decide separately whether to constrain scripts, and treat that as a product policy question rather than a coding one. Each rule you add is a rule some real person's real name will fail. That tension — abuse prevention against exclusion — is the actual content of the question; the predicates are just how you express whichever answer you choose. ## Versioning The tables in `unicode` are compiled from a specific Unicode release and move with the Go release; Go 1.27 tracks Unicode 17. A rune that was unassigned in one Go version can become an assigned letter in a later one, so "what passes validation" can change under a toolchain upgrade. That is rarely a problem in practice, but it is worth knowing before you assert that a validator's behaviour is frozen. ## What an interviewer is checking That you know these are Unicode-wide rather than ASCII, that you can name a concrete case where `IsDigit` and a numeric parser disagree, and that you treat a name validator as a policy decision with real exclusion costs rather than a regex to be tightened until abuse stops.
- How do unicode.IsDigit and unicode.IsNumber differ?`IsDigit` covers only category Nd, decimal digit number. `IsNumber` covers all of category N: Nd plus Nl, letter numbers such as Roman numerals, plus No, other numbers such as superscripts and fractions. So a Roman numeral is a number but not a digit. Neither implies the rune can be parsed by `strconv`, which accepts ASCII digits only.
- How would you restrict a display name to a specific script?Use the range tables directly: `unicode.Is(unicode.Latin, r)` tests one table and `unicode.In(r, unicode.Latin, unicode.Han)` tests several. Bear in mind that shared punctuation and digits are in the Common script and will fail a strict single-script test, so a usable rule allows Common alongside whichever scripts you permit.
- Why can a name pass a space check and still look padded?Because zero-width and format characters are not white space under Unicode's definition, so `unicode.IsSpace` is false for them and `strings.TrimSpace` will not remove them. They render as nothing, so the name looks shorter than it is and can be made visually identical to another. You have to reject them by category explicitly, or use an allowlist that never admits them.
saying these in an interview costs you the question
- Assumes IsLetter means the ASCII alphabet
- Believes IsDigit implies strconv can parse it
- Thinks IsSpace covers zero-width space
- Expects the predicates to be locale-aware
- Treats emoji as letters rather than symbols