How do Scanner's delimiter and locale settings affect parsing, and what surprises can they cause?
answer
- Delimiter is a regex (escape literal '.' as \\.)
- Non-whitespace delimiter leaves spaces in tokens
- nextDouble() parses per the Scanner's locale
- Default locale = JVM default -> 'works on my machine' bugs
- Pin useLocale(Locale.ROOT) for machine numbers
basics
~20 sScanner splits input using a delimiter pattern (whitespace by default) that you can change with useDelimiter(). It also parses numbers using a locale, so in some locales nextDouble() expects a comma as the decimal point and may reject a dot.
solid answer
~40 sTwo configurable behaviors trip people up. First, the delimiter: Scanner tokenizes using a regular expression, defaulting to whitespace. useDelimiter(",") or a regex lets you parse CSV-like input, but the new pattern then applies to everything, and surprises arise because the delimiter is regex (special characters must be escaped) and surrounding whitespace may become part of tokens. Second, the locale: number-reading methods like nextDouble()/nextInt() honor the Scanner's locale, which defaults to the JVM's default locale. In locales such as German or French the decimal separator is a comma and groups use a dot or space, so nextDouble() may throw InputMismatchException on "3.14" or accept "3,14". For deterministic parsing of program-formatted numbers, call useLocale(Locale.ROOT) (or Locale.US). The takeaway: Scanner parsing is locale- and delimiter-sensitive, so pin both when correctness matters.
code
java · 10 linesimport java.util.Locale;
import java.util.Scanner;
Scanner sc = new Scanner("3.14");
sc.useLocale(Locale.ROOT); // deterministic '.' decimal regardless of host
double pi = sc.nextDouble(); // 3.14 everywhere
Scanner csv = new Scanner("a, b, c");
csv.useDelimiter("\\s*,\\s*"); // regex: comma plus surrounding spaces
while (csv.hasNext()) System.out.println("[" + csv.next() + "]");go deeper
Knows you can change the separator with useDelimiter() and that whitespace is the default; may not yet know it is locale- or regex-sensitive.
Understands useDelimiter takes a regex and that custom delimiters keep surrounding whitespace; aware numbers parse per some locale.
Diagnoses an InputMismatchException on valid-looking decimals as a locale issue and pins Locale.ROOT; escapes regex delimiters and trims whitespace deliberately.
Establishes parsing conventions that are reproducible across environments, distinguishes machine-format from localized human input, and avoids Scanner where a stricter parser is warranted.
## Two hidden knobs Scanner looks simple, but two settings silently change what it parses: the **delimiter** and the **locale**. Misunderstanding either produces bugs that only show up with certain input or on certain machines. ## 1. The delimiter ### What it is The **delimiter** is the pattern that separates one token from the next. It is a **regular expression** (a mini-language for matching text patterns). By default the delimiter matches **whitespace** (`\p{javaWhitespace}+`), so input is split on spaces, tabs, and newlines. ### Changing it ```java Scanner sc = new Scanner("alice,bob,carol"); sc.useDelimiter(","); while (sc.hasNext()) System.out.println(sc.next()); // alice / bob / carol ``` ### The surprises - **It is a regex, not a literal.** `useDelimiter(".")` does NOT split on a literal dot — `.` in regex means "any character," so it would split on every character. Use `useDelimiter("\\.")` or `Pattern.quote(".")` for a literal dot. - **Whitespace is no longer special.** Once you set a non-whitespace delimiter, leading/trailing spaces around values become part of the token. For `"a, b, c"` with delimiter `","`, you get tokens `"a"`, `" b"`, `" c"` (note the spaces). You often need `useDelimiter("\\s*,\\s*")` to also swallow surrounding whitespace. - **The change is global** to that Scanner: all subsequent reads use the new pattern until you change it again. ## 2. The locale ### What a locale is A **locale** is a set of regional conventions — language, and crucially for parsing, **number formatting rules**: which character is the decimal separator and which is the grouping (thousands) separator. In the US/UK, `1,234.56` means one thousand two hundred thirty-four point five six (comma groups, dot decimal). In Germany, the same value is written `1.234,56` (dot groups, comma decimal). ### How Scanner uses it Scanner's numeric methods (`nextInt`, `nextLong`, `nextDouble`, `nextBigDecimal`, ...) parse according to the **Scanner's locale**, which defaults to the **JVM default locale** (the machine's regional setting). This means: ```java // On a machine whose default locale uses ',' as the decimal separator: Scanner sc = new Scanner("3.14"); sc.nextDouble(); // throws InputMismatchException — '.' is a grouping char here, not a decimal ``` The *same code* works on a US-locale machine and fails on a German-locale machine. This is a classic "works on my machine" bug. ### Pinning the locale For input your own program formats (or any input you control), make parsing deterministic: ```java Scanner sc = new Scanner(System.in); sc.useLocale(Locale.ROOT); // or Locale.US — '.' decimal, no surprises double x = sc.nextDouble(); ``` `Locale.ROOT` is the neutral, culture-independent locale — a safe default for machine-readable numbers. If you are reading genuinely localized human input, set the matching locale instead. ## 3. They interact with reading code - A `nextDouble()` that throws `InputMismatchException` on perfectly valid-looking `"3.14"` is almost always a **locale** problem. - A loop that reads one giant token instead of several is almost always a **delimiter** problem (often a regex-literal mistake). ## Practical rule When parsing machine-generated or fixed-format input with Scanner, **pin both knobs**: `useLocale(Locale.ROOT)` for numbers and an explicit `useDelimiter(...)` (regex-escaped) for non-whitespace separators. When parsing localized human input, choose the locale deliberately. Leaving them at their defaults makes correctness depend on the host machine — exactly what you do not want.
- Why does useDelimiter(".") behave unexpectedly?Because the delimiter is a regular expression and '.' matches any single character, so it splits on every character. To split on a literal dot use useDelimiter("\\.") or Pattern.quote(".").
- How do you make nextDouble() parse '3.14' the same on every machine?Set the Scanner's locale explicitly with sc.useLocale(Locale.ROOT) (or Locale.US) so the decimal separator is always '.', instead of relying on the JVM default locale.
The delimiter is the pair of scissors that cuts the input into pieces, and the locale is the ruler that decides what '3.14' measures to. Hand someone the wrong scissors or the wrong ruler and the same input comes out different.
saying these in an interview costs you the question
- Treating useDelimiter's argument as a literal string instead of a regex
- Assuming nextDouble() always accepts a dot decimal regardless of locale
- Ignoring that the default locale comes from the host machine, causing non-reproducible bugs
- Forgetting that a custom delimiter leaves surrounding whitespace inside tokens