How can lookahead/lookbehind be used with String.split or Pattern.split to keep delimiters, and what is the trade-off versus a normal split?
answer
- split discards the delimiter; zero-width split keeps it
- (?<=,) delimiter rides with the LEFT token; (?=,) with the RIGHT
- (?<=,)|(?=,) isolates the delimiter
- camelCase: (?<=[a-z])(?=[A-Z]) - no delimiter char needed
- Watch for empty pieces at the boundaries; bounded-lookbehind rule still applies
basics
~20 ssplit breaks on the matched delimiter and throws it away. If you split on a zero-width lookaround position instead, no characters are consumed, so the delimiter stays attached to a piece. For example split on (?<=,) keeps the comma at the end of each chunk.
solid answer
~40 sString.split/Pattern.split divide the input at each match of the pattern and discard the matched text. Because a lookaround is zero-width, splitting on a lookaround pattern cuts at a *position* without consuming any character, so the delimiter is preserved in the output. (?<=,) splits *after* each comma (comma stays at the end of the preceding token); (?=,) splits *before* each comma (comma starts the next token); combining (?<=,)|(?=,) isolates the comma as its own piece. The trade-off: it is concise and avoids a re-join step, but it is harder to read than splitting normally and re-attaching, can produce empty strings at boundaries, and won't benefit from split's trailing-empty-removal in obvious ways. For complex tokenization a Matcher loop or a real parser is clearer; for camelCase or unit-suffix splits the lookaround idiom is a neat fit.
code
java · 9 lines// keep the delimiter
String[] a = "a,b,c".split("(?<=,)"); // [a, , b, , c]
String[] b = "a,b,c".split("(?=,)"); // [a, ,b, ,c]
String[] c = "a,b".split("(?<=,)|(?=,)"); // [a, ',', b]
// split camelCase / number+unit without losing characters
String[] words = "getHTTPResponse".split("(?<=[a-z])(?=[A-Z])"); // [get, HTTPResponse]
String[] nu = "10px".split("(?<=\\d)(?=\\D)"); // [10, px]
System.out.println(java.util.Arrays.toString(words));go deeper
Knows that normal split removes the delimiter and that a lookaround split can keep it.
Can write (?<=,) / (?=,) / combined splits and the camelCase boundary split, and explain which side the delimiter attaches to.
Weighs the readability/empty-string/performance trade-offs and knows when a Matcher loop or parser is better.
Sets conventions: zero-width split for simple delimiter-keeping/boundary cases, real parsers for structured formats; documents the gotchas for the team.
## How split works `"a,b,c".split(",")` finds each `,`, **consumes** it, and returns the pieces around it: `["a","b","c"]`. The delimiter is gone because it was part of the match that split removes. ## Splitting on a zero-width position A lookaround matches a **position** and consumes **nothing**. So when `split` cuts at that position, there is no character to remove - the cut simply happens *between* two characters, and both characters stay in the output. This is the trick to **keep the delimiter**. ### Keep the delimiter at the end of each token `"a,b,c".split("(?<=,)")` -> `["a,", "b,", "c"]`. The lookbehind cuts right *after* each comma, so the comma rides along with the token before it. ### Keep the delimiter at the start of each token `"a,b,c".split("(?=,)")` -> `["a", ",b", ",c"]`. The lookahead cuts right *before* each comma. ### Isolate the delimiter `"a,b".split("(?<=,)|(?=,)")` -> `["a", ",", "b"]`. Cutting both before and after the comma puts it in a piece of its own. ## Real-world uses - **camelCase / PascalCase splitting:** `"getHTTPResponse".split("(?<=[a-z])(?=[A-Z])")` -> `["get", "HTTPResponse"]` (split between a lowercase and an uppercase letter without losing either). - **Number + unit:** `"10px".split("(?<=\\d)(?=\\D)")` -> `["10", "px"]`. - **Keeping sentence-ending punctuation** attached to each sentence. ## Trade-offs vs a normal split **Pros** - One step: no need to split then re-attach the delimiter. - Works for "split between two character classes" cases that have *no* delimiter character at all (camelCase) - a normal split has nothing to consume there, so a zero-width pattern is the natural tool. **Cons / gotchas** - **Readability:** `(?<=,)|(?=,)` is cryptic; a comment or a named helper helps. - **Empty strings:** splitting at boundaries can yield leading/trailing empty pieces; `split` removes *trailing* empties by default (limit 0) but not leading ones, so check the ends. With a positive limit, empties are kept. - **Performance:** marginally more work than a literal split, rarely significant. - **Lookbehind length rule still applies** if your lookbehind is variable length (must be bounded). ## When to prefer alternatives For anything beyond simple delimiter-keeping, a **Matcher loop** (`find()` + `group()`/`start()`/`end()`) gives full control and is often clearer; for structured input (CSV with quoting, code, JSON) use a **proper parser/library** rather than any regex.
- How do you split 'getHTTPResponse' into words without losing letters?Split on a zero-width boundary between case changes, e.g. split("(?<=[a-z])(?=[A-Z])") (and optionally also (?<=[A-Z])(?=[A-Z][a-z]) to handle acronym-to-word transitions).
- What happens to empty strings when splitting on a lookaround at the very start?You can get a leading empty string; split with limit 0 trims trailing empties but not leading ones, so handle or filter the first element if needed.
- Why does a normal split with a delimiter not work for camelCase?There is no delimiter character between 'get' and 'HTTP' to consume; the split point is purely a position between two classes, which only a zero-width assertion can express.
saying these in an interview costs you the question
- Thinking split keeps the delimiter by default
- Assuming the zero-width pattern consumes the comma
- Using it for quoted CSV instead of a real parser
- Forgetting leading empty strings are not auto-trimmed