skip to content

In PHP, when does preg_split() beat explode(), and what do PREG_SPLIT_NO_EMPTY and PREG_SPLIT_DELIM_CAPTURE change in its result?

level: middleimportance: nice to knowfreq 25%

answer

  1. separator that varies
  2. \R for any line ending
  3. limit -1 or 0 means none
  4. NO_EMPTY drops empty pieces
  5. DELIM_CAPTURE keeps captured separators

basics

~20 s

preg_split() splits on a pattern, so it handles separators that vary, such as mixed commas and semicolons or any line ending with \R. PREG_SPLIT_NO_EMPTY drops empty pieces, and PREG_SPLIT_DELIM_CAPTURE includes captured parts of the separator in the result.

solid answer

~40 s

`preg_split(string $pattern, string $subject, int $limit = -1, int $flags = 0): array|false` cuts the subject wherever the pattern matches. Use it when the separator is not one fixed string: `preg_split('/\R/', $body)` splits an email on `\r\n`, `\n` or `\r` alike, and `'/[\s,;]+/'` splits a list of references typed with commas, semicolons or spaces. For a single literal separator, `explode()` is simpler and faster. A `$limit` of `-1` or `0` means no limit; a positive limit leaves the rest in the last piece. `PREG_SPLIT_NO_EMPTY` removes empty pieces, such as those from a trailing separator. `PREG_SPLIT_DELIM_CAPTURE` adds any **parenthesized** part of the separator to the result, so you can keep which separator was used. `PREG_SPLIT_OFFSET_CAPTURE` returns `[piece, byteOffset]` pairs. Flags combine with `|`, and the function returns `false` on failure.

code

php · 17 lines
php
<?php
declare(strict_types=1);

$body = "Order: ORD-1, ORD-2;ORD-3 \r\nReason: damaged\n\nThanks";

$lines = preg_split('/\R/', $body);
// ['Order: ORD-1, ORD-2;ORD-3 ', 'Reason: damaged', '', 'Thanks']

$list = substr($lines[0], strlen('Order: '));
$refs = preg_split('/[\s,;]+/', $list, -1, PREG_SPLIT_NO_EMPTY);
// ['ORD-1', 'ORD-2', 'ORD-3']

$withSeparators = preg_split('/\s*([,;])\s*/', 'ORD-1, ORD-2;ORD-3', -1, PREG_SPLIT_DELIM_CAPTURE);
// ['ORD-1', ',', 'ORD-2', ';', 'ORD-3']

$naive = explode(',', 'ORD-1, ORD-2;ORD-3');
// ['ORD-1', ' ORD-2;ORD-3']

go deeper

for a junior

Recall that preg_split splits on a pattern and that PREG_SPLIT_NO_EMPTY removes empty pieces.

for a middle

Explain \R, the limit semantics, and what DELIM_CAPTURE and OFFSET_CAPTURE add to the result.

for a senior

Choose explode for literal separators on hot paths, preg_split for variable ones, and check the false return on external input.

for a principal

Decide when ad hoc splitting of message text should give way to a structured format or a real parser for the data being extracted.

## What preg_split() does `preg_split(string $pattern, string $subject, int $limit = -1, int $flags = 0): array|false` returns the pieces of `$subject` between the matches of `$pattern`. It is the regex counterpart of `explode()`: where `explode()` needs one literal separator, `preg_split()` accepts any pattern. ## When it is the right tool | Task | Call | Why not explode() | |---|---|---| | split an email into lines | `preg_split('/\R/', $body)` | line endings may be `\r\n`, `\n` or `\r` | | split a typed list of references | `preg_split('/[\s,;]+/', $list, -1, PREG_SPLIT_NO_EMPTY)` | users mix commas, semicolons and spaces | | split on a word, any case | `preg_split('/\s+and\s+/i', $text)` | case and surrounding whitespace vary | | split on a fixed character | `explode(',', $list)` | no pattern needed, so no regex | `\R` is PCRE's escape for **any newline sequence**; it treats `\r\n` as one line break, so Windows-style email bodies do not produce empty lines between every line. ## The limit - `-1` (the default) or `0`: no limit. - A positive number `n`: at most `n` pieces; the last piece contains the unsplit rest of the subject. Unlike `explode()`, there is no negative-limit form that drops pieces from the end. ## The flags Flags are combined with `|`: 1. **`PREG_SPLIT_NO_EMPTY`** returns only non-empty pieces. Leading, trailing and doubled separators otherwise produce empty strings. 2. **`PREG_SPLIT_DELIM_CAPTURE`** includes the text of **capturing groups** in the separator pattern in the result, in order, between the pieces. A separator without parentheses is still discarded. 3. **`PREG_SPLIT_OFFSET_CAPTURE`** changes every element into `[piece, byteOffset]`, where the offset counts bytes into the subject. With `DELIM_CAPTURE`, splitting `'ORD-1, ORD-2;ORD-3'` on `'/\s*([,;])\s*/'` yields `['ORD-1', ',', 'ORD-2', ';', 'ORD-3']`: the separators survive, which helps when the separator carries meaning. ## UTF-8 and characters With the `u` modifier the pattern works on code points. `preg_split('//u', $s, -1, PREG_SPLIT_NO_EMPTY)` is a well-known idiom for splitting UTF-8 text into code points; an empty pattern matches between every character, and `NO_EMPTY` removes the empty pieces at the ends. A dedicated multibyte split function is clearer for that job today, but you will meet the idiom in existing code. ## Failure and cost `preg_split()` returns `false` when the pattern cannot be compiled or matching fails, for example on invalid UTF-8 under `u`. Code that passes the result straight to `foreach` then fails on a `bool`; check it first. Every call compiles or fetches a cached pattern and runs the regex engine, so for a fixed literal separator on a hot path, `explode()` is the cheaper choice. The two also differ in edge cases: `explode()` with an empty separator throws, while an empty pattern in `preg_split()` is legal and splits between characters. ## A worked email example To find reference lines in a support email: 1. Split the body into lines with `preg_split('/\R/', $body)`. 2. Keep the lines that start with `Order:` using a prefix check. 3. Split the rest of such a line on `'/[\s,;]+/'` with `PREG_SPLIT_NO_EMPTY` to get each reference. Each step uses the simplest tool that handles the variation in the input. ## Common mistakes - **Splitting on `\n` only.** Bodies written on Windows keep a trailing `\r` on every line, which then breaks equality checks and prefix tests. - **Expecting all separators from `DELIM_CAPTURE`.** Only the text of capturing groups is returned; an unparenthesized separator is still discarded. - **Forgetting `NO_EMPTY`.** A trailing comma or a double space produces empty pieces that later become empty references. - **Using a regex for a fixed character.** `explode()` does the same job without compiling a pattern and is easier to read. - **Ignoring `false`.** On invalid UTF-8 with `u`, the call fails, and a `foreach` over `false` raises a warning while processing nothing. Each of these shows up as a small data-quality defect in an import rather than as a crash, which is why reviewers look for them.

  • In PHP, why use \R rather than \n to split an email into lines?
    `\R` matches any newline sequence, including `\r\n` as a single unit, `\n` and `\r`. Splitting a Windows-style body on `\n` leaves a trailing `\r` on every line, and splitting on `[\r\n]` creates an empty piece between `\r` and `\n`.
  • What does PREG_SPLIT_DELIM_CAPTURE return for a separator pattern without parentheses?
    Nothing extra: only capturing groups in the separator are added to the result. Wrap the part you want to keep in parentheses, as in `'/\s*([,;])\s*/'`, to receive the comma or semicolon between the pieces.

saying these in an interview costs you the question

  • preg_split with PREG_SPLIT_DELIM_CAPTURE returns every separator, captured or not.
  • A limit of 0 in preg_split returns a single piece.
  • preg_split is always a better choice than explode.
  • PREG_SPLIT_OFFSET_CAPTURE reports character offsets for UTF-8 text.
  • preg_split never fails, so its result can go straight into foreach.