skip to content

Text Processing

PHP strings are byte sequences built with quotes, heredoc and interpolation, handled by str* functions, mbstring for UTF-8 and PCRE for patterns. Interviewers check where bytes and characters differ.

part ofPHPoverview, primer and where to startread it →
on this pageshow

explore

questions

23

In PHP, how do explode(), implode() and trim() behave when you split a comma-separated list of invoice codes, and which edge cases bite?

level: juniorimportance: must knowfreq 62%

answer

  1. separator is the first argument
  2. limit keeps the rest in the last element
  3. empty input gives one empty element
  4. empty separator throws ValueError
  5. trim takes a character list

basics

~20 s

explode(',', $s) splits on a separator and returns an array, implode(',', $parts) joins it back, and trim() strips whitespace. Traps: an empty string explodes to [''], an empty separator throws ValueError, and trim's second argument is a set of characters.

solid answer

~40 s

`explode(string $separator, string $string, int $limit = PHP_INT_MAX): array` splits on every occurrence of the separator. A positive `$limit` caps the number of elements and leaves the remainder in the last one; a negative limit drops that many elements from the end. Exploding an empty string returns `['']`, one empty element, so `count()` is 1, and since PHP 8.0 an empty separator throws `ValueError`. `implode(', ', $parts)` joins with the separator first; the old reversed argument order was removed in 8.0. `trim()` strips spaces, tabs, newlines, carriage returns, vertical tabs and NUL bytes from both ends, and its optional second argument is a **list of characters**, not a substring, so `rtrim('1200', '0')` returns `'12'`. A robust split maps `trim` over the parts and drops the empty ones.

code

php · 15 lines
php
<?php
declare(strict_types=1);

$raw = " INV-001, INV-002 ,,INV-003\n";

$codes = array_map(trim(...), explode(',', $raw));
$codes = array_values(array_filter($codes, fn (string $c): bool => $c !== ''));

echo implode(';', $codes), "\n";      // INV-001;INV-002;INV-003

var_dump(explode(',', ''));            // array(1) { [0]=> string(0) "" }
var_dump(explode(',', 'a,b,c', 2));    // ['a', 'b,c']
var_dump(explode(',', 'a,b,c', -1));   // ['a', 'b']
var_dump(trim('INV-001-INV', 'INV'));  // string(5) "-001-"
var_dump(rtrim('1200', '0'));          // string(2) "12"

go deeper

for a junior

Recall the argument order of explode and implode, the default characters trim removes, and that trimming each part is part of a clean split.

for a middle

Explain the positive, negative and zero limit rules, the [''] result for an empty string, the ValueError for an empty separator and trim's character-set argument.

for a senior

Show you would harden an import against empty parts, pasted Unicode whitespace and '0' values that a callback-less array_filter drops, and switch to a CSV parser when fields can contain the separator.

for a principal

Argue when ad hoc explode parsing is acceptable versus when a documented format and a real parser should be mandated for data exchanged with other systems.

## Splitting with explode() `explode(string $separator, string $string, int $limit = PHP_INT_MAX): array` cuts `$string` at every occurrence of `$separator` and returns the pieces as a list. The separator comes **first** and the string second; swapping them is a common slip. The `$limit` argument changes the shape of the result: | Call | Result | Rule | |---|---|---| | `explode(',', 'a,b,c,d')` | `['a', 'b', 'c', 'd']` | split everywhere | | `explode(',', 'a,b,c,d', 2)` | `['a', 'b,c,d']` | at most 2 elements; the last holds the rest | | `explode(',', 'a,b,c,d', -1)` | `['a', 'b', 'c']` | drop the last element | | `explode(',', 'a,b,c,d', 0)` | `['a,b,c,d']` | 0 is treated as 1 | | `explode(',', 'abcd')` | `['abcd']` | no separator: one element | | `explode(',', '')` | `['']` | empty input: one empty element | A positive limit is useful for `key=value` lines where the value may itself contain the separator: `explode('=', $line, 2)` always yields at most two parts. ## Edge cases that bite 1. **Empty input is not an empty array.** `explode(',', '')` returns an array with one empty string, so `count()` is `1`. Code that asks "were any invoice codes supplied?" with `count(explode(...)) > 0` always says yes. 2. **Empty separator.** Since PHP 8.0 `explode('', $s)` throws a `ValueError`; before, it returned `false` with a warning. Splitting into single bytes is `str_split()`'s job. 3. **Adjacent separators.** `'a,,b'` produces an empty element in the middle, and a trailing comma produces one at the end. 4. **Whitespace survives.** `' INV-001, INV-002 '` explodes into elements that still carry their spaces. ## Cleaning the parts with trim() `trim(string $string, string $characters = " \n\r\t\v\0"): string` removes characters from both ends. The default set is: - the space `" "` - newline `\n` and carriage return `\r` - horizontal tab `\t` and vertical tab `\v` - the NUL byte `\0` `ltrim()` works only on the left, `rtrim()` only on the right, and `chop()` is an alias of `rtrim()`. The second argument is a **set of characters**, not a substring to remove. Any character in the set is stripped, in any order, until a character outside the set is reached: - `trim('INV-001-INV', 'INV')` returns `'-001-'`, because `I`, `N` and `V` are each stripped from both ends. - `rtrim('1200', '0')` returns `'12'`, which silently corrupts an amount if you meant to drop only formatting zeros. - A range can be written with `..`, as in `trim($binary, "\x00..\x1F")` to strip control bytes. `trim()` works on bytes and knows nothing about Unicode, so a UTF-8 non-breaking space pasted from a spreadsheet is **not** removed by default; the multibyte counterparts handle that. A robust split of a user-typed list therefore combines three steps: 1. `explode(',', $raw)` to cut the string; 2. `array_map(trim(...), ...)` to clean each part; 3. drop the empty strings with a callback that tests `$c !== ''`. Calling `array_filter()` **without** a callback is a trap here: it removes every falsy value, and the string `'0'` is falsy. ## Joining with implode() `implode(string $separator, array $array): string` glues the elements with the separator between them. `join()` is an alias. Details: - `implode(', ', [])` returns `''`. - `implode($array)` with a single argument uses an empty separator. - The legacy reversed order `implode($pieces, $glue)` was deprecated in PHP 7.4 and **removed in 8.0**. - Elements are converted to strings, so `implode(',', [1, 2.5, true])` gives `'1,2.5,1'`. `explode()` and `implode()` are not a quoting format: a code that itself contains the separator cannot round-trip. When fields may contain commas or quotes, use a real CSV parser instead. ## What a strong answer covers Name the argument order, the `$limit` semantics, the `['']` result for an empty string and the `ValueError` for an empty separator. Then show that you clean the pieces with `trim()`, know its default character list, and know that its second argument is a character set. That combination is what separates someone who has debugged an import from someone who has only read the signature.

  • Why does count(explode(',', $input)) > 0 fail as a check that at least one invoice code was supplied?
    Because, without a negative limit, `explode()` does not return an empty array: an empty `$input` yields `['']`, so the count is 1. Check the trimmed input for `''` first, or filter out empty strings after splitting and count what remains.
  • What does the second argument of trim() mean, and why does rtrim('1200', '0') return '12'?
    It is a list of characters to strip, not a substring. `rtrim()` removes any listed character from the right end until it meets one that is not listed, so both zeros go. To drop an exact suffix, test with `str_ends_with()` and cut with `substr()` instead.
  • What happens in PHP 8 when explode() receives an empty separator?
    It throws a `ValueError`. Before PHP 8.0 it emitted a warning and returned `false`. If the goal is to split a string into single bytes, `str_split()` is the function for that; for characters in UTF-8 text use the multibyte variant.

saying these in an interview costs you the question

  • explode returns an empty array when the input string is empty.
  • trim's second argument removes that exact substring from the ends.
  • trim also strips non-breaking spaces and other Unicode whitespace by default.
  • implode accepts the array and the separator in either order in PHP 8.
  • explode with an empty separator splits the string into characters.
  • A limit of 2 throws away everything after the second element.
open as a page

In PHP, why is if (strpos($code, 'INV')) a bug, and how should you test whether a string contains a substring?

level: juniorimportance: must knowfreq 80%

basics

~20 s

strpos() returns the byte offset of the first match or false, and a match at offset 0 is falsy, so a plain if treats it as not found. Compare with !== false, or call str_contains(), added in PHP 8.0.

open as a page

In PHP, why does strlen('日本語') return 9, and how do you count the characters of a UTF-8 string instead?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A PHP string is a sequence of bytes with no encoding attached, and strlen() counts bytes; each of those three kanji takes three bytes in UTF-8. mb_strlen($s, 'UTF-8') decodes the bytes and counts characters, returning 3.

open as a page

In PHP, what is the difference between single-quoted and double-quoted strings?

level: juniorimportance: must knowfreq 75%

basics

~10 s

Single-quoted strings are literal: only ' and \ are escapes and variables are not expanded. Double-quoted strings interpret escape sequences such as \n, \t and \u{...} and interpolate variables like $name.

open as a page

In PHP, why does a preg_* pattern need delimiters such as /…/ or #…#, and what do the i, u, x and s modifiers change?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Delimiters mark where a PHP pattern ends so modifiers can follow, as in '/ord-\d+/i'. i ignores case, u treats pattern and subject as UTF-8, x ignores pattern whitespace and allows comments, and s lets the dot match newlines.

open as a page

In a PHP marketplace, why does substr($name, 0, 20) corrupt Japanese product names, and how do you truncate them safely for a listing?

level: middleimportance: must knowfreq 55%

basics

~20 s

substr() cuts at a byte offset, so it can split a three-byte Japanese character and leave invalid UTF-8. mb_substr() cuts at character boundaries, and mb_strimwidth() truncates to a display width, counting full-width characters as two columns.

open as a page

In PHP, what can simple $var interpolation embed in a double-quoted string, and when do you need the {$expr} curly syntax?

level: middleimportance: must knowfreq 50%

basics

~10 s

Simple syntax embeds a variable, one array element with an unquoted key or number, or one property: "$name", "$arr[0]", "$obj->prop". Anything deeper, quoted keys or method calls need curly syntax: "{$obj->a->b}", "{$arr['key']}".

open as a page

In PHP, what do preg_match() and preg_match_all() return, and why should code extracting order references compare the result with === false?

level: middleimportance: must knowfreq 60%

basics

~20 s

preg_match() returns 1 for a match, 0 for none and false when the regex fails; preg_match_all() returns the number of matches or false. A truthy if merges 0 and false, so an engine error looks like 'no order reference found'.

open as a page

In PHP, how do $s[0] and $s[-1] read single characters of a string, and what happens with an offset past the end?

level: juniorimportance: should knowfreq 35%

basics

~20 s

$s[0] returns the first byte of a string as a one-character string and $s[-1] the last. Reading past the end gives an E_WARNING, Uninitialized string offset, and an empty string. Offsets count bytes, not characters.

open as a page

In PHP, how do sprintf() conversion specifications like %05d, %'*10s, %.2f and %1$s work when formatting invoice lines?

level: middleimportance: should knowfreq 45%

basics

~10 s

Each sprintf() specification is %[argnum$][flags][width][.precision]specifier: %05d zero-pads an integer to five characters, %'*10s pads a string with asterisks, %.2f prints two decimals and %1$s reuses the first argument. Too few values throw ArgumentCountError.

open as a page

In PHP, why can str_replace() with array arguments give a different result from strtr() given the same replacement pairs?

level: middleimportance: should knowfreq 38%

basics

~20 s

str_replace() applies each search/replace pair in turn over the whole string, so text inserted by one pair can be replaced by a later pair. strtr() with an array makes one pass, tries the longest key first and never re-scans replaced text.

open as a page

In PHP 8.4 and later, why do strtolower(), str_pad(), trim() and ucfirst() mishandle UTF-8 text, and which mb_* functions replace them?

level: middleimportance: should knowfreq 40%

basics

~10 s

The byte functions only understand ASCII: strtolower() leaves 'Ä' alone, str_pad() counts bytes, trim() ignores the ideographic space U+3000, and ucfirst() cannot uppercase 'é'. Use mb_strtolower(), mb_str_pad() (8.3), mb_trim() and mb_ucfirst() (8.4).

open as a page

In PHP, how do heredoc and nowdoc strings differ, and what does the flexible closing marker introduced in PHP 7.3 allow?

level: middleimportance: should knowfreq 40%

basics

~20 s

Heredoc (<<<EOT) is a multi-line double-quoted string that interpolates variables and escapes; nowdoc (<<<'EOT') is its literal, single-quoted twin. Since PHP 7.3 the closing marker may be indented, and that indentation is removed from every line.

open as a page

In PHP, how do named groups like (?<seq>\d+) appear in preg_match_all() results, and what do PREG_SET_ORDER and PREG_UNMATCHED_AS_NULL change?

level: middleimportance: should knowfreq 35%

basics

~20 s

A named group appears in the matches array twice, under its name and its number. preg_match_all() defaults to PREG_PATTERN_ORDER, one list per group; PREG_SET_ORDER gives one array per match; PREG_UNMATCHED_AS_NULL reports non-participating groups as null.

open as a page

In PHP, when do you use preg_replace_callback() instead of preg_replace(), and how do $1, ${1} and \1 references work in a replacement?

level: middleimportance: should knowfreq 45%

basics

~20 s

preg_replace() substitutes a template in which $1, \1 or ${1} insert captured groups; preg_replace_callback() calls a function with the matches array and uses its return value, for replacements that need logic such as zero-padding or a lookup.

open as a page

A PHP accounting export starts writing 1234,50 instead of 1234.50 after a library calls setlocale(); which formatting functions follow the locale, and how do you make the output locale-proof?

level: seniorimportance: should knowfreq 28%

basics

~20 s

sprintf() and printf() with %f or %g use the LC_NUMERIC decimal point, so a setlocale() call turns 1234.50 into 1234,50. Use %F, or number_format($x, 2, '.', ''), for machine output; echoing a float ignores the locale since PHP 8.0.

open as a page

A PHP marketplace truncates emoji-rich product names with mb_substr() and leaves broken flags and family emoji; why, and what does intl's grapheme API fix?

level: seniorimportance: should knowfreq 30%

basics

~20 s

mb_substr() counts code points, but a flag is two code points and a family emoji is several joined by zero-width joiners, so a cut can split one visible character. grapheme_substr() and grapheme_strlen() from intl count grapheme clusters instead.

open as a page

A PHP marketplace imports supplier product feeds in Shift_JIS and sometimes broken UTF-8; how do you validate and convert them with mb_check_encoding(), mb_convert_encoding() and iconv()?

level: seniorimportance: should knowfreq 32%

basics

~10 s

Convert each feed once, from its declared encoding, with mb_convert_encoding($s, 'UTF-8', 'SJIS'), then check the result with mb_check_encoding($s, 'UTF-8'). mbstring replaces invalid bytes with a substitute character; iconv() returns false unless //IGNORE is appended.

open as a page

After an upgrade to PHP 8.2 or later, logs fill with 'Using ${var} in strings is deprecated' notices; what triggers them and how do you fix each form?

level: seniorimportance: should knowfreq 22%

basics

~10 s

PHP 8.2 deprecated the ${...} interpolation syntax. "${name}" and "${name['key']}" become "{$name}" and "{$name['key']}"; any other "${expr}" is a variable variable and becomes "{${expr}}". PHP 8.5 still only deprecates it.

open as a page

A PHP worker extracting order references from large support emails with preg_match() silently finds none on some messages; how do pcre.backtrack_limit and preg_last_error() explain it?

level: seniorimportance: should knowfreq 30%

basics

~10 s

When a match needs more work than pcre.backtrack_limit (1000000 by default) allows, preg_match() returns false without a warning and preg_last_error() reports PREG_BACKTRACK_LIMIT_ERROR; code that tests the result for truthiness reads that as 'no reference'.

open as a page

In PHP, why does sorting invoice codes like INV-9 and INV-10 with strcmp() put INV-10 first, and what does strnatcmp() change?

level: middleimportance: nice to knowfreq 25%

basics

~20 s

strcmp() compares strings byte by byte, so at the first difference '1' sorts before '9' and INV-10 precedes INV-9. strnatcmp() compares runs of digits as numbers, giving the human order INV-9, INV-10; strnatcasecmp() also ignores ASCII case.

open as a page

In PHP, why can two identical-looking UTF-8 strings such as 'café' fail an === comparison, and how does intl's Normalizer fix it?

level: middleimportance: nice to knowfreq 22%

basics

~10 s

'é' can be one code point, U+00E9, or 'e' plus a combining accent U+0301; both render the same but differ in bytes, so === fails. Normalizer::normalize($s, Normalizer::FORM_C) converts both to the same composed form.

open as a page

In PHP, when does preg_split() beat explode(), and what do PREG_SPLIT_NO_EMPTY and PREG_SPLIT_DELIM_CAPTURE change in its result?

level: middleimportance: nice to knowfreq 25%

basics

~20 s

preg_split() splits on a pattern, so it handles separators that vary, such as mixed commas and semicolons or any line ending with \R. PREG_SPLIT_NO_EMPTY drops empty pieces, and PREG_SPLIT_DELIM_CAPTURE includes captured parts of the separator in the result.

open as a page