skip to content

Using lookarounds, how would you insert thousands separators into a number string in Java, and why does this work without deleting any digits?

level: seniorimportance: should knowfreq 38%

answer

  1. Pattern: (?<=\d)(?=(\d{3})+$) replace with ','
  2. Matches a POSITION, not characters - zero width
  3. Lookbehind: digit before; lookahead: multiple-of-three digits to end
  4. $ anchor aligns grouping to the right
  5. No digit in group(0) => pure insertion, nothing deleted
  6. Real formatting: use NumberFormat/DecimalFormat + Locale

basics

~20 s

Match the empty positions between digits where exactly a multiple of three digits remain, using (?<=\d)(?=(\d{3})+$), and replace each with a comma. It matches zero-width positions, so no digit is consumed or removed - commas are just inserted.

solid answer

~40 s

The pattern is replaceAll("(?<=\\d)(?=(\\d{3})+$)", ","). It matches a **position**, not characters: the positive lookbehind (?<=\d) requires a digit immediately before the position, and the positive lookahead (?=(\d{3})+$) requires that the rest of the string from this position to the end is a whole number of three-digit groups. Both are zero-width, so the match has width zero and consumes nothing; replaceAll then substitutes that empty match with a comma, effectively *inserting* it. Because no digit is ever part of group(0), none are deleted. For "1234567" it inserts commas to yield "1,234,567". Caveats: it assumes a digits-only string (strip or anchor around a decimal point and sign first), and on very long inputs the repeated (\d{3})+$ lookahead has some cost. This is the canonical demonstration of why zero-width assertions matter for replacement.

code

java · 14 lines
java
import java.util.regex.*;

String grouped = "1234567".replaceAll("(?<=\\d)(?=(\\d{3})+$)", ",");
System.out.println(grouped); // 1,234,567

// Caveat: digits-only assumption; handle decimals/sign yourself, e.g.
String in = "-1234567.5";
String sign = in.startsWith("-") ? "-" : "";
String body = sign.isEmpty() ? in : in.substring(1);
int dot = body.indexOf('.');
String intPart = dot < 0 ? body : body.substring(0, dot);
String frac = dot < 0 ? "" : body.substring(dot);
String out = sign + intPart.replaceAll("(?<=\\d)(?=(\\d{3})+$)", ",") + frac;
System.out.println(out); // -1,234,567.5

go deeper

for a junior

Can recognize the pattern inserts commas and that nothing is deleted because it is zero-width.

for a middle

Can explain each assertion and the role of the $ anchor, and apply the pattern to a plain integer string.

for a senior

Can trace match positions, explain why it is insertion not deletion, and note the digits-only/locale/performance caveats.

for a principal

Recommends DecimalFormat/Locale for production, treats the regex as a teaching/edge tool, and reasons about input normalization and performance trade-offs.

## The goal Turn `"1234567"` into `"1,234,567"` by inserting commas every three digits **from the right**, without altering any digit. ## Why naive replacement fails If you matched actual characters (e.g. groups of three digits and re-emitted them with commas), you'd have to reconstruct the string and handle the leftmost partial group carefully. Lookarounds let you instead target the **gaps between digits** directly. ## The pattern, piece by piece `(?<=\d)(?=(\d{3})+$)` This matches an **empty position** (zero characters) that satisfies two simultaneous conditions: 1. `(?<=\d)` - **positive lookbehind**: there is a digit immediately *before* this position. (Prevents inserting a comma before the very first digit.) 2. `(?=(\d{3})+$)` - **positive lookahead**: from this position to the end of the string, the text is one-or-more groups of exactly three digits, then end-of-string `$`. In other words, the number of digits remaining is a positive multiple of three. A comma belongs exactly where both are true: there is a digit behind, and the digits ahead come in clean groups of three down to the end. ## Trace on "1234567" Positions (between characters) indexed 0..7. Remaining-digit counts from each position: after index0 there are 7 digits left, index1->6, index2->5, index3->4, index4->3, etc. We need a digit behind AND remaining count a multiple of three: - index1 (after `1`): 6 remaining = 3x2, digit behind -> **match** -> insert `,` - index4 (after `4`): 3 remaining = 3x1, digit behind -> **match** -> insert `,` Result: `1,234,567`. No other position satisfies both. ## Why no digit is deleted Both lookbehind and lookahead are **zero-width**: the overall match consumes **zero characters**, so `group(0)` is the empty string at that position. `replaceAll` replaces the matched text (empty) with `","` - this is a pure **insertion**. Since digits are never part of the match, none can be removed. This is the headline reason zero-width assertions exist: they let you act *between* characters. ## Important caveats - **Digits-only input.** The trailing `$` and `\d{3}` assume the whole remaining string is digits. For `"1234567.89"` or `"-1234567"`, first split off the sign and fractional part, format the integer part, then reassemble. Or anchor the lookahead to stop before the decimal using a more complex sub-pattern. - **Locale.** This hardcodes `,`. For real formatting prefer `java.text.NumberFormat` / `DecimalFormat` with a `Locale`, which handles grouping, decimals, and locale-specific separators correctly. The regex trick is great for demonstrating lookarounds, less so as production money-formatting. - **Performance.** `(\d{3})+$` is re-evaluated at every position; for pathologically long numbers this is more work than a single linear pass. Fine for normal lengths. ## Variant: grouping from the left / fixed width The same idea generalizes: `(?<=\d)(?=(\d{4})+$)` groups by four, etc. The `+$` anchor is what makes grouping align to the right end.

  • What does the $ anchor contribute, and what happens without it?
    $ forces the (\d{3})+ to run all the way to the end, which is what makes grouping align from the right. Without it, the lookahead would succeed at many positions (any three-digit run ahead), inserting commas in the wrong places.
  • Why is DecimalFormat usually preferable in production?
    It handles locales (separator and decimal symbols differ by region), negative numbers, decimals, and currency correctly, and is clearer and less error-prone than a regex for the same job.
  • How would you adapt the pattern to a string that has a decimal part like "1234567.5"?
    Split on the dot, format only the integer part with the regex (or NumberFormat), then concatenate the fractional part back, since the (\d{3})+$ anchor assumes digits run to the end.

saying these in an interview costs you the question

  • Thinking the regex deletes/re-emits digits rather than inserting at a gap
  • Forgetting the (?<=\d) so a comma is placed before the first digit
  • Applying it to strings with decimals/signs without preprocessing
  • Recommending it as production money formatting instead of NumberFormat

context