skip to content

PCRE Regular Expressions

The preg_* functions run PCRE patterns: preg_match, preg_replace_callback and preg_split, with delimiters, modifiers and named groups. Interviewers check false-versus-0 returns and backtrack limits.

part ofPHPoverview, primer and where to startread it →
on this pageshow

explore

questions

6

In PHP, why does a preg_* pattern need delimiters such as /…/ or #…#, and what do the i, u, x and s modifiers change?

level: juniorimportance: must knowfreq 55%

answer

  1. delimiter, pattern, delimiter, modifiers
  2. not alphanumeric, backslash or NUL
  3. escape the delimiter or pick another
  4. u means UTF-8 and Unicode classes
  5. preg_quote with the delimiter argument

basics

~20 s

Delimiters mark where a PHP pattern ends so modifiers can follow, as in '/ord-\d+/i'. i ignores case, u treats pattern and subject as UTF-8, x ignores pattern whitespace and allows comments, and s lets the dot match newlines.

solid answer

~40 s

The `preg_*` functions take a single string that contains the regex **between delimiters**, followed by **modifiers**: `'/ORD-\d+/iu'`. The delimiter can be any character except an alphanumeric, a backslash or a NUL byte; bracket pairs such as `{...}` also work, and `#` or `~` is common when the pattern contains slashes. A delimiter inside the pattern must be escaped, and user input is inserted with `preg_quote($input, '/')`, which escapes regex metacharacters and the delimiter. The modifiers: `i` matches case-insensitively; `u` treats pattern and subject as UTF-8, makes invalid UTF-8 an error, and lets `\d`, `\w` and `\s` match Unicode characters; `x` ignores unescaped whitespace and allows `#` comments; `s` makes `.` match newlines; `m` makes `^` and `$` match at line breaks. PHP 8.2 added `n`, which makes plain groups non-capturing.

code

php · 17 lines
php
<?php
declare(strict_types=1);

$body = "Order ord-2026-000123\nplaced Monday";

var_dump(preg_match('/ORD-\d{4}-\d{6}/', $body));   // int(0): case differs
var_dump(preg_match('/ORD-\d{4}-\d{6}/i', $body));  // int(1)
var_dump(preg_match('/000123.placed/', $body));      // int(0): . stops at \n
var_dump(preg_match('/000123.placed/s', $body));     // int(1)

var_dump(preg_match('/^\d+$/', '123'));            // int(0): bytes, ASCII \d
var_dump(preg_match('/^\d+$/u', '123'));           // int(1): Unicode digits
var_dump(preg_match('/^[0-9]+$/u', '123'));        // int(0): ASCII only

$prefix = 'EU/B2B';
$pattern = '/\b' . preg_quote($prefix, '/') . '-\d{6}\b/';
echo $pattern, "\n";                                 // /\bEU\/B2B-\d{6}\b/

go deeper

for a junior

Recall that a PHP pattern needs delimiters, that modifiers follow the closing delimiter, and what i, s and u do.

for a middle

Explain delimiter rules, escaping with preg_quote and its delimiter argument, the m, x and D modifiers, and what u changes for invalid input and \d.

for a senior

Show you would audit patterns built from user data for preg_quote, choose [0-9] over \d under u where it matters, and keep long patterns readable with x.

for a principal

Set conventions for regex in the codebase: delimiter style, mandatory u for user text, and when a pattern is complex enough to deserve a parser instead.

## Anatomy of a PHP pattern string Every `preg_*` function receives the regular expression as one string with three parts: ``` /ORD-(\d{4})-(\d{6})/iu ^ delimiter ^ delimiter, then modifiers ``` The delimiters mark where the expression ends so that **modifiers** can follow. Rules enforced by the PCRE extension: - Leading whitespace is skipped; the first other character is the delimiter. - The delimiter must **not** be alphanumeric, a backslash or a NUL byte; otherwise PHP warns `Delimiter must not be alphanumeric, backslash, or NUL byte` and the function returns `false`. - Bracket-style delimiters pair up: `(...)`, `[...]`, `{...}` and `<...>`. - If the closing delimiter is missing, PHP warns `No ending delimiter ... found`. A pattern written as `'ORD-\d+'` therefore fails: `O` is alphanumeric. ## Escaping the delimiter and user input A delimiter character inside the expression must be escaped with a backslash, or you choose a delimiter that does not occur: `#https?://#` reads better than `/https?:\/\//`. When part of the pattern comes from data, for example a customer's order prefix typed into a settings page, escape it: ```php $pattern = '/\b' . preg_quote($prefix, '/') . '-\d{6}\b/'; ``` `preg_quote(string $str, ?string $delimiter = null): string` escapes the characters `. \ + * ? [ ^ ] $ ( ) { } = ! < > | : - #` and, when given, the delimiter. `#` has been escaped since PHP 7.3. It is meant for the **pattern**, not for the replacement string of `preg_replace()`. Write patterns in **single-quoted** PHP strings, so that the regex backslashes reach PCRE unchanged instead of being interpreted as PHP escape sequences first. ## The modifiers you use most | Modifier | PCRE2 option | Effect | |---|---|---| | `i` | caseless | letters match regardless of case | | `m` | multiline | `^` and `$` also match at internal line breaks | | `s` | dotall | `.` also matches newline characters | | `x` | extended | unescaped whitespace is ignored and `#` starts a comment | | `u` | UTF plus Unicode properties | pattern and subject are UTF-8; `\d`, `\w`, `\s` become Unicode-aware | | `D` | dollar end only | `$` matches only at the very end, not before a final newline | | `U` | ungreedy | inverts the default greediness of quantifiers | | `A` | anchored | the match must start at the search offset | | `J` | duplicate names | allows several groups with the same name | | `n` | no auto capture (PHP 8.2) | only named groups capture | | `r` | caseless restrict (PHP 8.4) | with `i`, ASCII and non-ASCII letters do not match each other | ## The u modifier in detail Without `u`, PCRE works on **bytes**: `.` matches one byte of a multi-byte character and a character class like `[é]` holds two separate bytes. With `u`: 1. the pattern and the subject are treated as **UTF-8**, so `.` matches a whole code point; 2. an invalid UTF-8 subject makes the call fail with `PREG_BAD_UTF8_ERROR` instead of matching garbage; 3. PHP also enables Unicode properties, so `\d` matches any decimal digit, including full-width `123`, and `\w` matches letters of any script. Point 3 surprises people: an order-reference pattern with `\d{6}` and `u` accepts full-width digits, which later fail an integer conversion. Use `[0-9]` where only ASCII digits are valid. ## Removed and ignored modifiers - `e`, which evaluated the replacement as PHP code, was removed in PHP 7.0; today it produces the warning `Unknown modifier 'e'`. `preg_replace_callback()` replaces it. - `X` has been accepted but ignored since PHP 8.0, because invalid escape sequences are always errors now. - `S` is accepted and ignored; PCRE2 studies patterns automatically. ## Readable patterns with x For longer patterns the `x` modifier lets you split and comment the expression, which reviewers appreciate: ```php $ref = '/\b ORD - (?<year>[0-9]{4}) # year - (?<seq>[0-9]{6}) \b # sequence /x'; ``` Remember that literal spaces must then be written as `\s`, `[ ]` or `\ `.

  • Why does '/^\d{6}$/u' accept full-width digits in PHP?
    With the `u` modifier PHP compiles the pattern with Unicode properties enabled, so `\d` matches any decimal digit in Unicode, including full-width `0`-`9`. If only ASCII digits are valid, write `[0-9]`, or drop `u` when the subject is guaranteed ASCII.
  • What does preg_quote() escape, and why pass the delimiter as its second argument?
    It prefixes regex metacharacters such as `.`, `+`, `*`, `?`, brackets, `$`, `|` and `#` with a backslash. The delimiter is not special to PCRE, so it is escaped only when passed; forgetting it lets input containing `/` end the pattern early and trigger an unknown-modifier warning.
  • What happened to the e modifier?
    It made `preg_replace()` evaluate the replacement as PHP code, which turned regex input into code injection. It was removed in PHP 7.0 and now produces an `Unknown modifier 'e'` warning with a `false` or `null` result. `preg_replace_callback()` is the replacement.

saying these in an interview costs you the question

  • PHP patterns can be written without delimiters, like '\d+'.
  • A letter such as 'a' can be used as a pattern delimiter.
  • The u modifier only affects how the pattern is read, not the subject.
  • With u, \d still matches only the ASCII digits 0 to 9.
  • preg_quote escapes the delimiter even when it is not passed.
open as a page

In PHP, what do preg_match() and preg_match_all() return, and why should code extracting order references compare the result with === false?

level: middleimportance: must knowfreq 60%

basics

~20 s

preg_match() returns 1 for a match, 0 for none and false when the regex fails; preg_match_all() returns the number of matches or false. A truthy if merges 0 and false, so an engine error looks like 'no order reference found'.

open as a page

In PHP, how do named groups like (?<seq>\d+) appear in preg_match_all() results, and what do PREG_SET_ORDER and PREG_UNMATCHED_AS_NULL change?

level: middleimportance: should knowfreq 35%

basics

~20 s

A named group appears in the matches array twice, under its name and its number. preg_match_all() defaults to PREG_PATTERN_ORDER, one list per group; PREG_SET_ORDER gives one array per match; PREG_UNMATCHED_AS_NULL reports non-participating groups as null.

open as a page

In PHP, when do you use preg_replace_callback() instead of preg_replace(), and how do $1, ${1} and \1 references work in a replacement?

level: middleimportance: should knowfreq 45%

basics

~20 s

preg_replace() substitutes a template in which $1, \1 or ${1} insert captured groups; preg_replace_callback() calls a function with the matches array and uses its return value, for replacements that need logic such as zero-padding or a lookup.

open as a page

A PHP worker extracting order references from large support emails with preg_match() silently finds none on some messages; how do pcre.backtrack_limit and preg_last_error() explain it?

level: seniorimportance: should knowfreq 30%

basics

~10 s

When a match needs more work than pcre.backtrack_limit (1000000 by default) allows, preg_match() returns false without a warning and preg_last_error() reports PREG_BACKTRACK_LIMIT_ERROR; code that tests the result for truthiness reads that as 'no reference'.

open as a page

In PHP, when does preg_split() beat explode(), and what do PREG_SPLIT_NO_EMPTY and PREG_SPLIT_DELIM_CAPTURE change in its result?

level: middleimportance: nice to knowfreq 25%

basics

~20 s

preg_split() splits on a pattern, so it handles separators that vary, such as mixed commas and semicolons or any line ending with \R. PREG_SPLIT_NO_EMPTY drops empty pieces, and PREG_SPLIT_DELIM_CAPTURE includes captured parts of the separator in the result.

open as a page