How does String.split() behave with String/Char delimiters, the limit parameter, and ignoreCase — and how does it differ from the regex overload?
answer
- String/Char split is literal; only Regex overload is a pattern
- limit = n -> at most n parts, remainder in last element
- Empties from edge/adjacent delimiters are KEPT
- Multiple delimiters: first listed match wins
- splitToSequence for lazy large-input processing
basics
~20 ssplit breaks a string into a list of pieces around a separator. You can give one or more separators, cap how many pieces you get with limit, and ignore case. There is also a version that splits on a regex pattern.
solid answer
~40 ssplit(vararg delimiters: String, ignoreCase = false, limit = 0) and split(vararg delimiters: Char, ...) perform LITERAL splitting — delimiters are not regex. You can pass multiple delimiters; the first that matches at each position wins. limit caps the number of substrings: limit = n keeps at most n parts, putting the unsplit remainder into the last element; limit = 0 (default) means no limit. ignoreCase makes string delimiters case-insensitive. Adjacent delimiters and leading/trailing delimiters produce empty strings in the result (split does NOT drop empties). The separate overload split(regex: Regex, limit = 0) interprets a pattern. A common mistake is using the String overload with regex metacharacters expecting pattern behavior. Result type is List<String>; use splitToSequence for lazy processing of large inputs.
code
kotlin · 8 linesval header = "Set-Cookie: id=abc; Path=/"
val (name, value) = header.split(": ", limit = 2)
println(name) // Set-Cookie
println(value) // id=abc; Path=/ (rest kept because limit=2)
println("a,,b".split(",")) // [a, , b] empties kept
println("a1b2c".split(Regex("\\d"))) // [a, b, c]
println("x;y,z".split(";", ",")) // [x, y, z]go deeper
Knows split returns a List of pieces around a separator.
Knows literal vs regex overloads and that empties are kept; filters them when needed.
Uses limit for bounded key/value parsing, multiple delimiters, and chooses splitToSequence for large inputs.
Sets parsing guidelines (literal vs regex, empty handling, laziness) to avoid perf and correctness pitfalls in shared text-processing code.
## Signatures ```kotlin fun CharSequence.split(vararg delimiters: String, ignoreCase: Boolean = false, limit: Int = 0): List<String> fun CharSequence.split(vararg delimiters: Char, ignoreCase: Boolean = false, limit: Int = 0): List<String> fun CharSequence.split(regex: Regex, limit: Int = 0): List<String> ``` ## Literal vs regex The `String`/`Char` overloads are **literal** — a delimiter like `"."` matches a real dot, not "any char". Only the `Regex` overload treats the argument as a pattern. ```kotlin "a.b.c".split(".") // [a, b, c] (literal dot) "a1b2c".split(Regex("\\d")) // [a, b, c] (pattern) ``` ## Multiple delimiters You may pass several; at each position the **first listed** delimiter that matches is used: ```kotlin "a,b;c".split(",", ";") // [a, b, c] ``` ## The limit parameter - `limit = 0` (default): split on **every** occurrence, no cap. - `limit = n` (n > 0): produce **at most n** substrings; once n-1 splits happen, the **entire remainder** (including further delimiters) becomes the last element. ```kotlin "a=b=c".split("=", limit = 2) // [a, b=c] ``` This is ideal for key/value parsing where the value may itself contain the delimiter. ## Empty strings are kept Leading, trailing, and consecutive delimiters yield empty strings — `split` does **not** discard them: ```kotlin ",a,,b,".split(",") // ["", "a", "", "b", ""] ``` Filter with `.filter { it.isNotEmpty() }` if you want them gone. ## ignoreCase Applies to **String** delimiters: `"aXbxc".split("x", ignoreCase = true)` → `[a, b, c]`. ## Laziness `splitToSequence(...)` returns a `Sequence<String>` for lazy, allocation-friendly processing of large text without building the whole `List` up front. ## When to use what - Fixed separators → literal `split` (fast, no regex compile). - Variable patterns → `split(Regex(...))`. - Only the first side needed → prefer `substringBefore/After` for clarity. - Bounded fields (e.g. `key=value`) → `limit`. ## Common mistakes Expecting `"a.b".split(".")` to behave like regex (it is literal), and forgetting empties from edge delimiters.
- How do you split on whitespace runs of any length?Use the regex overload: text.split(Regex("\\s+")), since literal split would not collapse multiple spaces and would emit empties.
- Why use limit = 2 when parsing 'key=value'?It guarantees exactly two parts even if the value contains '='; the remainder after the first split goes entirely into the second element.
saying these in an interview costs you the question
- Assuming the String overload interprets regex metacharacters
- Expecting split to drop empty strings from edges/adjacent delimiters
- Not knowing limit puts the remainder in the last element
- Building a List with split for huge inputs instead of splitToSequence