skip to content

In PHP, why does sorting invoice codes like INV-9 and INV-10 with strcmp() put INV-10 first, and what does strnatcmp() change?

level: middleimportance: nice to knowfreq 25%

answer

  1. byte-by-byte comparison
  2. '1' sorts before '9'
  3. only the sign is meaningful
  4. strcasecmp folds ASCII only
  5. digit runs compared as numbers

basics

~20 s

strcmp() compares strings byte by byte, so at the first difference '1' sorts before '9' and INV-10 precedes INV-9. strnatcmp() compares runs of digits as numbers, giving the human order INV-9, INV-10; strnatcasecmp() also ignores ASCII case.

solid answer

~40 s

`strcmp(string $string1, string $string2): int` is a binary-safe, case-sensitive comparison of bytes. It returns a negative number, zero or a positive number, and only the **sign** is meaningful; since PHP 8.2 it may return -1 or 1 instead of the length difference, so compare with `0`. Because it compares bytes, `'INV-10'` sorts before `'INV-9'`: the first differing bytes are `1` and `9`. `strnatcmp()` implements **natural order**: it reads runs of digits as numbers, so `INV-9` comes before `INV-10`. `strcasecmp()` and `strnatcasecmp()` ignore case, but only for ASCII letters. None of them is locale-aware; `strcoll()` or intl's `Collator` handle language rules. In practice you pass them to `usort()` as the comparator, for example `usort($codes, strnatcasecmp(...))`.

code

php · 16 lines
php
<?php
declare(strict_types=1);

$codes = ['INV-10', 'INV-9', 'inv-2', 'INV-100'];

usort($codes, strcmp(...));
echo implode(', ', $codes), "\n";  // INV-10, INV-100, INV-9, inv-2

usort($codes, strnatcasecmp(...));
echo implode(', ', $codes), "\n";  // inv-2, INV-9, INV-10, INV-100

var_dump(strcmp('INV-10', 'INV-9') < 0);     // bool(true): '1' < '9'
var_dump(strnatcmp('INV-10', 'INV-9') > 0);  // bool(true): 10 > 9
var_dump(strcasecmp('INV-9', 'inv-9') === 0); // bool(true): ASCII case folded
var_dump('1e3' == '1000');                    // bool(true): numeric strings
var_dump(strcmp('1e3', '1000') === 0);        // bool(false)

go deeper

for a junior

Recall that strcmp returns a negative, zero or positive int, and that strcasecmp is its case-insensitive version.

for a middle

Explain byte-wise comparison, why INV-10 sorts before INV-9, how strnatcmp reads digit runs, and why only the sign of the result counts.

for a senior

Choose between natural comparison, zero-padded identifiers and locale-aware collation for sortable codes, and catch code that depends on the exact strcmp value.

for a principal

Decide which system owns the canonical ordering of identifiers, so PHP sorting, database ORDER BY and exported files agree without per-caller comparators.

## Byte comparison with strcmp() `strcmp(string $string1, string $string2): int` compares two strings **byte by byte**. It stops at the first byte that differs and reports which string is smaller; if one string is a prefix of the other, the shorter one is smaller. The comparison is: - **binary-safe**: NUL bytes and arbitrary binary data are compared like any other byte; - **case-sensitive**: `'A'` (byte 65) sorts before `'a'` (byte 97); - **not locale-aware**: accented letters are just byte values. The return value is **less than 0**, **0** or **greater than 0**, and the manual says nothing beyond the sign can be relied on. Since PHP 8.2 the functions are no longer guaranteed to return the length difference `strlen($string1) - strlen($string2)` when lengths differ; they may return `-1` or `1`. Test `strcmp($a, $b) < 0`, never `=== -1`. ## Why INV-10 sorts before INV-9 Comparing `'INV-10'` with `'INV-9'`, the first four bytes are equal. The fifth bytes are `'1'` (byte 49) and `'9'` (byte 57), so `strcmp()` returns a negative value and `INV-10` comes first. The digits after that are never looked at. Sorting a batch of invoice codes with it gives: | Sorted with | Order | |---|---| | `strcmp()` | `INV-10`, `INV-100`, `INV-9`, `inv-2` | | `strcasecmp()` | `INV-10`, `INV-100`, `inv-2`, `INV-9` | | `strnatcmp()` | `INV-9`, `INV-10`, `INV-100`, `inv-2` | | `strnatcasecmp()` | `inv-2`, `INV-9`, `INV-10`, `INV-100` | The lowercase `inv-2` lands last under the case-sensitive functions because `i` (105) is greater than `I` (73). ## Natural order with strnatcmp() `strnatcmp(string $string1, string $string2): int` implements the "natural order" algorithm: it walks both strings and, where both have a **run of digits**, compares the runs by numeric value rather than byte by byte. That produces the order a person expects for file names, version-like labels and sequence numbers. `strnatcasecmp()` is the same with ASCII case folded. Natural order is a heuristic for mixed text and numbers, not a parser. When the numeric part is what matters, zero-padding it at creation time (`INV-00009`) makes every comparison, including plain `strcmp()` and database `ORDER BY`, agree. ## The rest of the family 1. **`strcasecmp()`**: like `strcmp()` but ASCII letters compare case-insensitively. Bytes outside ASCII are compared unchanged, so `strcasecmp('ÉCOLE', 'école')` is **not** 0 for UTF-8 input. 2. **`strncmp()` / `strncasecmp()`**: compare at most the first `$length` bytes, useful for prefix tests when you also need an ordering. 3. **`strcoll()`**: compares using the `LC_COLLATE` locale, which is process-wide state. 4. **`Collator::compare()`** from intl: locale-aware collation for a named locale. ## Using them to sort All of these return the negative/zero/positive contract that `usort()`, `uasort()` and `uksort()` expect, so they can be passed directly as the comparator: ```php usort($codes, strnatcasecmp(...)); ``` The `strnatcasecmp(...)` syntax creates a closure from the function (a first-class callable, PHP 8.1+); the string `'strnatcasecmp'` works too. Sort flags such as natural ordering built into `sort()` are an alternative with their own rules. ## strcmp() versus == and === For a plain equality test between two strings, `$a === $b` is clearer than `strcmp($a, $b) === 0` and gives the same answer. Avoid `==` for strings that may look numeric: two numeric strings compare as numbers under `==`, so `'1e3' == '1000'` is `true`, while `strcmp('1e3', '1000')` reports them as different. Reach for `strcmp()` when you need an **ordering**, not just equality. ## What interviewers listen for This question checks whether you can predict an order before running the code. A strong answer: - walks the bytes of `INV-10` and `INV-9` and names the first difference; - states the sign-only return contract and the PHP 8.2 note; - picks `strnatcmp()` or `strnatcasecmp()` for human-facing lists and explains its digit-run rule; - points out that ASCII-only case folding makes `strcasecmp()` unsuitable for non-English names; - suggests fixing the data, with zero-padded sequence numbers, when every consumer must agree on the order. Those points show you have debugged a real sorting complaint rather than memorised a function list.

  • Why should you not write strcmp($a, $b) === -1 in PHP 8?
    The documented contract is only the sign. Older versions often returned the length difference, and since PHP 8.2 the functions may return -1 or 1 instead, so the concrete value is an implementation detail. Compare with zero: `strcmp($a, $b) < 0`.
  • Is strcmp($a, $b) === 0 different from $a === $b when both are strings?
    For two strings they agree, and `===` is clearer. The difference that matters is with `==`: two numeric strings compare as numbers, so `'1e3' == '1000'` is true while `strcmp()` sees different bytes. Use `strcmp()` when you need an ordering rather than equality.
  • How do you compare strings according to a language's alphabet rules, such as accents?
    `strcmp()`, `strcasecmp()` and `strnatcmp()` only compare bytes, with ASCII case folding at most. For language rules use intl's `Collator::compare()` with a named locale, or `strcoll()`, which depends on the process-wide `LC_COLLATE` setting.

saying these in an interview costs you the question

  • strcmp returns true when the two strings are equal.
  • strcmp always returns exactly -1, 0 or 1, so === -1 is a safe test.
  • strcasecmp treats accented UTF-8 letters like É and é as equal.
  • strcmp compares the numbers inside strings by their numeric value.
  • strnatcmp ignores letter case.