skip to content

In PHP, how do named groups like (?<seq>\d+) appear in preg_match_all() results, and what do PREG_SET_ORDER and PREG_UNMATCHED_AS_NULL change?

level: middleimportance: should knowfreq 35%

answer

  1. name key plus numeric key
  2. PREG_PATTERN_ORDER is the default
  3. PREG_SET_ORDER groups each match
  4. trailing unmatched groups vanish
  5. n modifier: only named groups capture

basics

~20 s

A named group appears in the matches array twice, under its name and its number. preg_match_all() defaults to PREG_PATTERN_ORDER, one list per group; PREG_SET_ORDER gives one array per match; PREG_UNMATCHED_AS_NULL reports non-participating groups as null.

solid answer

~40 s

A named group, written `(?<seq>...)`, `(?'seq'...)` or `(?P<seq>...)`, is also numbered, so `$matches` contains both `'seq'` and its index. `preg_match_all()` fills `$matches` in `PREG_PATTERN_ORDER` by default: `$matches[0]` lists every full match and `$matches['seq']` lists every captured sequence, which suits collecting one column. `PREG_SET_ORDER` instead returns one array per match, each with its own name and number keys, which suits iterating over references. Optional groups are the trap: by default a group that did not take part becomes `''` when a later group matched, and a **trailing** unmatched group is omitted from that match's array, so `$m['note']` raises an undefined-key warning. `PREG_UNMATCHED_AS_NULL` always includes such groups as `null`. `PREG_OFFSET_CAPTURE` turns entries into `[text, byteOffset]` pairs, and PHP 8.2's `n` modifier makes plain parentheses non-capturing.

code

php · 19 lines
php
<?php
declare(strict_types=1);

$body = 'ORD-2026-000123 (gift) and ORD-2026-000456';
$pattern = '/ORD-(?<year>\d{4})-(?<seq>\d{6})(?:\s\((?<note>\w+)\))?/';

preg_match_all($pattern, $body, $cols);
print_r($cols['seq']);            // [0 => '000123', 1 => '000456']
print_r($cols['note']);           // [0 => 'gift', 1 => '']

preg_match_all($pattern, $body, $rows, PREG_SET_ORDER);
var_dump(isset($rows[1]['note'])); // bool(false): trailing group omitted

preg_match_all($pattern, $body, $rows, PREG_SET_ORDER | PREG_UNMATCHED_AS_NULL);
foreach ($rows as $r) {
    printf("%s %s %s\n", $r['year'], $r['seq'], $r['note'] ?? '-');
}
// 2026 000123 gift
// 2026 000456 -

go deeper

for a junior

Recall the (?<name>...) syntax and that $matches['name'] holds the named capture.

for a middle

Explain pattern order versus set order, the duplicate name and number keys, and how optional groups appear with and without PREG_UNMATCHED_AS_NULL.

for a senior

Show you would name every read group, add PREG_UNMATCHED_AS_NULL to patterns with optional parts, and keep byte offsets consistent across the code.

for a principal

Decide how extraction results are shaped for downstream code, for example mapping matches into typed value objects instead of passing raw arrays around.

## Named groups PCRE supports three spellings for a named capture group: - `(?<seq>\d{6})` - `(?'seq'\d{6})` - `(?P<seq>\d{6})` A named group is still a **numbered** group. In the matches array PHP stores it twice: once under its name, once under its number. For `'/ORD-(?<year>\d{4})-(?<seq>\d{6})/'` matching `ORD-2026-000123`, `preg_match()` fills: ```php [0 => 'ORD-2026-000123', 'year' => '2026', 1 => '2026', 'seq' => '000123', 2 => '000123'] ``` Names make extraction code self-documenting (`$m['seq']` instead of `$m[2]`) and survive the insertion of new groups earlier in the pattern. Since PHP 8.4, whose bundled PCRE2 is 10.44, a group name may be up to 128 characters long. The `J` modifier allows the same name on several alternatives. ## Result layouts of preg_match_all() `preg_match_all()` collects every match, and the `$flags` argument decides the layout: | Flag | `$matches` shape | Good for | |---|---|---| | `PREG_PATTERN_ORDER` (default) | one list per group: `$m[0]` all full matches, `$m['seq']` all sequences | collecting one field from every match | | `PREG_SET_ORDER` | one array per match, each keyed like a `preg_match()` result | looping over matches as records | | `PREG_OFFSET_CAPTURE` | each entry becomes `[text, byteOffset]` | locating matches in the subject | | `PREG_UNMATCHED_AS_NULL` | non-participating groups are `null` and always present | optional groups | `PREG_PATTERN_ORDER` and `PREG_SET_ORDER` are alternatives; the other two combine with either using `|`. ## Optional groups and missing keys Suppose a support email may add a note in brackets after the reference: `ORD-2026-000123 (gift)`. The pattern gets an optional group: `(?:\s\((?<note>\w+)\))?`. For a reference without a note, the `note` group does not participate. What you get depends on the flags: 1. **Default, group followed by a matched group**: reported as an empty string `''`. 2. **Default, trailing unmatched group** in `preg_match()` or `PREG_SET_ORDER`: omitted from that match's array. Reading `$m['note']` then raises an undefined array key warning. 3. **`PREG_PATTERN_ORDER`**: the per-group lists are padded with `''` so that every list has one entry per match. 4. **`PREG_UNMATCHED_AS_NULL`**: always present, value `null`, in every layout. With `PREG_UNMATCHED_AS_NULL` you can distinguish "the note group matched an empty string" from "there was no note", and use `$m['note'] ?? '-'` safely. ## Byte offsets `PREG_OFFSET_CAPTURE` reports offsets in **bytes**, even with the `u` modifier. To show a character position to a user, convert with a multibyte length function on the prefix. For an unmatched group the offset is `-1`. ## Fewer captures with n Patterns often use parentheses only for grouping, which clutters `$matches` with numbered entries. Options: - write non-capturing groups `(?:...)`; - or add the `n` modifier (PHP 8.2), which makes plain `(...)` groups non-capturing while named groups still capture. The array still contains numbered keys for the named groups. ## Practical guidance - Name every group you read, and read it by name. - Use `PREG_SET_ORDER` when each match is a record you process in a loop. - Add `PREG_UNMATCHED_AS_NULL` whenever a pattern has optional groups, so keys are always present. - Keep offsets in bytes throughout, converting only for display. ## A worked extraction A support-mail importer that records every reference with its optional note can be written in a few steps: 1. Define the pattern with named groups for `year`, `seq` and `note`, the note group optional. 2. Call `preg_match_all()` with `PREG_SET_ORDER | PREG_UNMATCHED_AS_NULL`, and treat a `false` result as an error. 3. Loop over the sets and build a small value object from `$m['year']`, `$m['seq']` and `$m['note']`, where `null` means no note. 4. Add `PREG_OFFSET_CAPTURE` only if the importer needs to highlight the reference in the original text; it changes every entry into a pair, so the rest of the loop must read index `0` of each entry. Keeping the flags in one place, next to the pattern, makes the expected array shape obvious to the next reader.

  • Why does $matches contain both 'seq' and a number for the same group?
    A named group is also a numbered group, and PHP exposes both keys so code can use either. Filtering out numeric keys is an array operation after the call; alternatively, read only the names and ignore the rest.
  • Are offsets from PREG_OFFSET_CAPTURE character positions when the u modifier is set?
    No. They are always byte offsets into the subject. For UTF-8 text with multi-byte characters, a byte offset must be converted, for example by measuring the prefix up to that offset with a multibyte length function, before showing it as a character position.

saying these in an interview costs you the question

  • Named groups replace the numeric keys in $matches.
  • preg_match_all returns one array per match by default.
  • An optional group that did not match is always present as an empty string.
  • PREG_OFFSET_CAPTURE reports character offsets when the u modifier is used.
  • The n modifier also stops named groups from capturing.