skip to content

Text & Regex

String manipulation in Kotlin leans on a large extension API plus the Regex class and efficient text building. Interview coding tasks touch this constantly, so fluency here shows up immediately.

part ofKotlinoverview, primer and where to startread it →
on this pageshow

explore

questions

16

How do you create a Regex in Kotlin and check whether a string contains a match versus matches the whole input? Contrast matchEntire, find, and containsMatchIn.

level: juniorimportance: must knowfreq 70%

answer

  1. Regex(...) or "...".toRegex()
  2. containsMatchIn = anywhere (Boolean)
  3. find = first MatchResult?
  4. matchEntire/matches = whole string
  5. compile once, reuse

basics

~10 s

Make a Regex with Regex("...") or "...".toRegex(). Use containsMatchIn to see if a match appears anywhere, find to get the first match, and matchEntire when the whole string must match the pattern.

solid answer

~30 s

You build a pattern with the Regex(pattern) constructor or the String.toRegex() extension. To test it: containsMatchIn(input) returns a Boolean for 'is there a match anywhere'; find(input) returns the first MatchResult? (null if none) and supports a startIndex; matchEntire(input) returns a MatchResult? only if the ENTIRE input matches; and matches(input) is the Boolean form of matchEntire. Don't confuse partial vs full matching: "\\d+".toRegex().containsMatchIn("a1b") is true, but matchEntire("a1b") is null because letters surround the digit. Regex compilation is relatively expensive, so create the Regex once (e.g. a top-level val) and reuse it rather than recompiling per call.

code

kotlin · 5 lines
kotlin
val number = Regex("\\d+")

fun isAllDigits(s: String) = number.matches(s)        // whole string
fun hasDigit(s: String) = number.containsMatchIn(s)    // anywhere
fun firstNumber(s: String): String? = number.find(s)?.value

go deeper

for a junior

Knows how to construct a Regex and pick containsMatchIn vs matches for a basic 'contains' or 'validate' task.

for a middle

Articulates the partial-vs-full distinction, the nullable MatchResult return types, and the implicit anchoring of matches/matchEntire.

for a senior

Adds performance/thread-safety reasoning (compile once, immutable, reusable) and avoids the validation anti-pattern.

for a principal

Frames Regex choice in terms of API contracts and correctness risk, and reasons about when a parser/non-regex approach is safer than regex validation.

## Creating a Regex Kotlin's `kotlin.text.Regex` wraps the JVM's `java.util.regex.Pattern`. Two equivalent ways to build one: ```kotlin val r1 = Regex("\\d+") // constructor val r2 = "\\d+".toRegex() // String extension ``` Both take an optional `RegexOption` (e.g. `RegexOption.IGNORE_CASE`). Because compiling a pattern is comparatively costly, declare a `Regex` **once** (top-level `val`, companion, or property) and reuse it instead of constructing it inside a hot loop. ## Testing functions — partial vs full match - **`containsMatchIn(input): Boolean`** — true if the pattern matches *anywhere* in the input (a partial/unanchored search). This is what you want for "does this text contain a number?". - **`find(input, startIndex = 0): MatchResult?`** — returns the *first* match as a `MatchResult`, or `null`. Gives you position/value, not just a boolean. - **`matchEntire(input): MatchResult?`** — returns a `MatchResult` only if the **whole** input matches end to end; otherwise `null`. - **`matches(input): Boolean`** — boolean equivalent of `matchEntire`; in Kotlin you can also write `input matches regex` via the infix `String.matches`. ```kotlin val digits = Regex("\\d+") digits.containsMatchIn("a1b") // true (partial) digits.find("a1b")?.value // "1" digits.matchEntire("a1b") // null (letters around it) digits.matchEntire("123")?.value // "123" digits.matches("123") // true ``` ## Key gotcha: anchoring Unlike some languages, Kotlin/Java `matches`/`matchEntire` implicitly require the *entire* string to match — you do NOT need to add `^...$`. Conversely `find`/`containsMatchIn` are unanchored. Choosing the wrong one is the classic validation bug (e.g. accepting `"123abc"` as a number because you used `containsMatchIn`).

  • Why prefer a top-level val Regex over constructing it inside a function called in a loop?
    Compiling the pattern is expensive; reusing one immutable, thread-safe Regex avoids recompiling on every call and reduces allocation.
  • Does matches require ^ and $ anchors?
    No. matches/matchEntire already require the entire input to match, so explicit anchors are redundant.

containsMatchIn is asking 'is this word somewhere in the page?'; matchEntire is asking 'is the page exactly this one word?'.

saying these in an interview costs you the question

  • Using containsMatchIn for validation that should match the whole string
  • Thinking matchEntire returns a Boolean (it returns MatchResult?)
  • Recompiling Regex inside a loop or per request
  • Forgetting to escape backslashes (\\d) in a plain string literal
  • Believing find returns all matches

context

open as a page

Why is building a long string by concatenating with += inside a loop inefficient, and what does StringBuilder do differently?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Kotlin Strings are immutable, so each += makes a brand-new String and copies everything. In a loop that repeats over and over. StringBuilder keeps one growing buffer and just appends, avoiding all those copies.

open as a page

What is the difference between String.toInt() and String.toIntOrNull(), and when should you use each?

level: juniorimportance: must knowfreq 75%

basics

~10 s

toInt() converts text to a number but crashes with an error if the text is not a valid number. toIntOrNull() returns null instead of crashing, so you can handle bad input safely.

open as a page

Using findAll and MatchResult, how do you extract all matches and their captured groups? Explain what groupValues contains and the index convention.

level: middleimportance: must knowfreq 60%

basics

~10 s

findAll returns a lazy sequence of all matches. Each MatchResult has groupValues, a list where index 0 is the whole match and index 1+ are the captured groups in order.

open as a page

Explain isBlank/isEmpty, ifBlank/ifEmpty, and padStart/padEnd. How do they differ and how do they compose for input normalization?

level: middleimportance: must knowfreq 65%

basics

~10 s

isEmpty checks for zero length; isBlank also treats whitespace-only as empty. ifBlank/ifEmpty give a fallback value when the string is blank/empty. padStart/padEnd add filler characters to reach a target width.

open as a page

How do named capturing groups and MatchResult.destructured work in Kotlin? Show extracting fields by name and via destructuring (a, b).

level: middleimportance: should knowfreq 45%

basics

~10 s

Name a group with (?<name>...) and read it via match.groups["name"]?.value. You can also pull groups positionally with destructuring: val (a, b) = match.destructured.

open as a page

Walk through the core mutating operations of StringBuilder — append, insert, and deleteAt — including their index semantics and what happens with out-of-range indices.

level: middleimportance: should knowfreq 45%

basics

~10 s

append adds to the end. insert puts characters at a given position, shifting the rest right. deleteAt removes the single character at an index. Bad indices throw an IndexOutOfBoundsException.

open as a page

Explain how buildString works under the hood: what is the receiver inside its lambda, why is it inline, and how does it compare to constructing a StringBuilder manually?

level: middleimportance: should knowfreq 55%

basics

~10 s

buildString makes a StringBuilder for you, runs your code with that builder as this so you can just call append, and returns the finished string. It's basically StringBuilder().apply{...}.toString() wrapped up neatly.

open as a page

Explain substringBefore, substringAfter, substringBeforeLast, and substringAfterLast, including their delimiter and missingDelimiterValue behavior.

level: middleimportance: should knowfreq 60%

basics

~20 s

They slice a string around a chosen marker. 'Before' keeps the part to the left of the marker, 'After' keeps the part to the right. The 'Last' versions look for the last occurrence instead of the first.

open as a page

Compare trimIndent() and trimMargin(): how does each remove leading whitespace from multi-line strings, and what are the gotchas?

level: middleimportance: should knowfreq 55%

basics

~20 s

Both clean up the indentation of multi-line text you wrote inside indented code. trimIndent removes the common leading spaces automatically. trimMargin removes everything up to a marker character you put at the start of each line (default '|').

open as a page

Explain Regex.replace with a transform lambda. How does it differ from string-replacement, and what do $1/${name} mean in the string form?

level: seniorimportance: should knowfreq 45%

basics

~10 s

replace can take a lambda that receives each MatchResult and returns the replacement text, so you compute it dynamically. The string form instead uses $1 or ${name} to insert captured groups.

open as a page

Explain StringBuilder's internal capacity and growth strategy, and when pre-sizing capacity is worthwhile. How does this affect performance reasoning in hot paths?

level: seniorimportance: should knowfreq 35%

basics

~20 s

StringBuilder keeps a backing array bigger than the text it holds. When it fills, it makes a bigger array (usually about double) and copies. If you know roughly how long the result is, you can set the capacity up front to skip those copies.

open as a page

Contrast replace (String vs Char vs regex overloads), removePrefix/removeSuffix, and String.format("%.2f"). What are the common mistakes?

level: seniorimportance: should knowfreq 50%

basics

~10 s

replace swaps text for other text. removePrefix/removeSuffix strip a known beginning or ending only if it's there. format builds a string from a pattern, like showing a number with two decimals using %.2f.

open as a page

How does String.split() behave with String/Char delimiters, the limit parameter, and ignoreCase — and how does it differ from the regex overload?

level: seniorimportance: should knowfreq 48%

basics

~20 s

split breaks a string into a list of pieces around a separator. You can give one or more separators, cap how many pieces you get with limit, and ignore case. There is also a version that splits on a regex pattern.

open as a page

In Kotlin/JVM, how does StringBuilder relate to StringBuffer, and how should you handle text building that is shared across threads?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

StringBuilder is fast but not thread-safe. StringBuffer is the older, synchronized (thread-safe) version but slower. For shared mutable text, prefer giving each thread its own builder or using proper synchronization rather than relying on StringBuffer.

open as a page

What are RegexOption settings and common correctness/performance pitfalls (escaping, raw strings, catastrophic backtracking) when using Kotlin Regex on untrusted input?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

RegexOption tunes matching (ignore case, multiline, dot-matches-newline). Watch escaping in literals, use raw strings, and avoid patterns that can backtrack catastrophically on hostile input.

open as a page