In PHP, why does sorting invoice codes like INV-9 and INV-10 with strcmp() put INV-10 first, and what does strnatcmp() change?
answer
- byte-by-byte comparison
- '1' sorts before '9'
- only the sign is meaningful
- strcasecmp folds ASCII only
- digit runs compared as numbers
basics
~20 sstrcmp() compares strings byte by byte, so at the first difference '1' sorts before '9' and INV-10 precedes INV-9. strnatcmp() compares runs of digits as numbers, giving the human order INV-9, INV-10; strnatcasecmp() also ignores ASCII case.
solid answer
~40 s`strcmp(string $string1, string $string2): int` is a binary-safe, case-sensitive comparison of bytes. It returns a negative number, zero or a positive number, and only the **sign** is meaningful; since PHP 8.2 it may return -1 or 1 instead of the length difference, so compare with `0`. Because it compares bytes, `'INV-10'` sorts before `'INV-9'`: the first differing bytes are `1` and `9`. `strnatcmp()` implements **natural order**: it reads runs of digits as numbers, so `INV-9` comes before `INV-10`. `strcasecmp()` and `strnatcasecmp()` ignore case, but only for ASCII letters. None of them is locale-aware; `strcoll()` or intl's `Collator` handle language rules. In practice you pass them to `usort()` as the comparator, for example `usort($codes, strnatcasecmp(...))`.
code
php · 16 lines<?php
declare(strict_types=1);
$codes = ['INV-10', 'INV-9', 'inv-2', 'INV-100'];
usort($codes, strcmp(...));
echo implode(', ', $codes), "\n"; // INV-10, INV-100, INV-9, inv-2
usort($codes, strnatcasecmp(...));
echo implode(', ', $codes), "\n"; // inv-2, INV-9, INV-10, INV-100
var_dump(strcmp('INV-10', 'INV-9') < 0); // bool(true): '1' < '9'
var_dump(strnatcmp('INV-10', 'INV-9') > 0); // bool(true): 10 > 9
var_dump(strcasecmp('INV-9', 'inv-9') === 0); // bool(true): ASCII case folded
var_dump('1e3' == '1000'); // bool(true): numeric strings
var_dump(strcmp('1e3', '1000') === 0); // bool(false)go deeper
Recall that strcmp returns a negative, zero or positive int, and that strcasecmp is its case-insensitive version.
Explain byte-wise comparison, why INV-10 sorts before INV-9, how strnatcmp reads digit runs, and why only the sign of the result counts.
Choose between natural comparison, zero-padded identifiers and locale-aware collation for sortable codes, and catch code that depends on the exact strcmp value.
Decide which system owns the canonical ordering of identifiers, so PHP sorting, database ORDER BY and exported files agree without per-caller comparators.
## Byte comparison with strcmp() `strcmp(string $string1, string $string2): int` compares two strings **byte by byte**. It stops at the first byte that differs and reports which string is smaller; if one string is a prefix of the other, the shorter one is smaller. The comparison is: - **binary-safe**: NUL bytes and arbitrary binary data are compared like any other byte; - **case-sensitive**: `'A'` (byte 65) sorts before `'a'` (byte 97); - **not locale-aware**: accented letters are just byte values. The return value is **less than 0**, **0** or **greater than 0**, and the manual says nothing beyond the sign can be relied on. Since PHP 8.2 the functions are no longer guaranteed to return the length difference `strlen($string1) - strlen($string2)` when lengths differ; they may return `-1` or `1`. Test `strcmp($a, $b) < 0`, never `=== -1`. ## Why INV-10 sorts before INV-9 Comparing `'INV-10'` with `'INV-9'`, the first four bytes are equal. The fifth bytes are `'1'` (byte 49) and `'9'` (byte 57), so `strcmp()` returns a negative value and `INV-10` comes first. The digits after that are never looked at. Sorting a batch of invoice codes with it gives: | Sorted with | Order | |---|---| | `strcmp()` | `INV-10`, `INV-100`, `INV-9`, `inv-2` | | `strcasecmp()` | `INV-10`, `INV-100`, `inv-2`, `INV-9` | | `strnatcmp()` | `INV-9`, `INV-10`, `INV-100`, `inv-2` | | `strnatcasecmp()` | `inv-2`, `INV-9`, `INV-10`, `INV-100` | The lowercase `inv-2` lands last under the case-sensitive functions because `i` (105) is greater than `I` (73). ## Natural order with strnatcmp() `strnatcmp(string $string1, string $string2): int` implements the "natural order" algorithm: it walks both strings and, where both have a **run of digits**, compares the runs by numeric value rather than byte by byte. That produces the order a person expects for file names, version-like labels and sequence numbers. `strnatcasecmp()` is the same with ASCII case folded. Natural order is a heuristic for mixed text and numbers, not a parser. When the numeric part is what matters, zero-padding it at creation time (`INV-00009`) makes every comparison, including plain `strcmp()` and database `ORDER BY`, agree. ## The rest of the family 1. **`strcasecmp()`**: like `strcmp()` but ASCII letters compare case-insensitively. Bytes outside ASCII are compared unchanged, so `strcasecmp('ÉCOLE', 'école')` is **not** 0 for UTF-8 input. 2. **`strncmp()` / `strncasecmp()`**: compare at most the first `$length` bytes, useful for prefix tests when you also need an ordering. 3. **`strcoll()`**: compares using the `LC_COLLATE` locale, which is process-wide state. 4. **`Collator::compare()`** from intl: locale-aware collation for a named locale. ## Using them to sort All of these return the negative/zero/positive contract that `usort()`, `uasort()` and `uksort()` expect, so they can be passed directly as the comparator: ```php usort($codes, strnatcasecmp(...)); ``` The `strnatcasecmp(...)` syntax creates a closure from the function (a first-class callable, PHP 8.1+); the string `'strnatcasecmp'` works too. Sort flags such as natural ordering built into `sort()` are an alternative with their own rules. ## strcmp() versus == and === For a plain equality test between two strings, `$a === $b` is clearer than `strcmp($a, $b) === 0` and gives the same answer. Avoid `==` for strings that may look numeric: two numeric strings compare as numbers under `==`, so `'1e3' == '1000'` is `true`, while `strcmp('1e3', '1000')` reports them as different. Reach for `strcmp()` when you need an **ordering**, not just equality. ## What interviewers listen for This question checks whether you can predict an order before running the code. A strong answer: - walks the bytes of `INV-10` and `INV-9` and names the first difference; - states the sign-only return contract and the PHP 8.2 note; - picks `strnatcmp()` or `strnatcasecmp()` for human-facing lists and explains its digit-run rule; - points out that ASCII-only case folding makes `strcasecmp()` unsuitable for non-English names; - suggests fixing the data, with zero-padded sequence numbers, when every consumer must agree on the order. Those points show you have debugged a real sorting complaint rather than memorised a function list.
- Why should you not write strcmp($a, $b) === -1 in PHP 8?The documented contract is only the sign. Older versions often returned the length difference, and since PHP 8.2 the functions may return -1 or 1 instead, so the concrete value is an implementation detail. Compare with zero: `strcmp($a, $b) < 0`.
- Is strcmp($a, $b) === 0 different from $a === $b when both are strings?For two strings they agree, and `===` is clearer. The difference that matters is with `==`: two numeric strings compare as numbers, so `'1e3' == '1000'` is true while `strcmp()` sees different bytes. Use `strcmp()` when you need an ordering rather than equality.
- How do you compare strings according to a language's alphabet rules, such as accents?`strcmp()`, `strcasecmp()` and `strnatcmp()` only compare bytes, with ASCII case folding at most. For language rules use intl's `Collator::compare()` with a named locale, or `strcoll()`, which depends on the process-wide `LC_COLLATE` setting.
saying these in an interview costs you the question
- strcmp returns true when the two strings are equal.
- strcmp always returns exactly -1, 0 or 1, so === -1 is a safe test.
- strcasecmp treats accented UTF-8 letters like É and é as equal.
- strcmp compares the numbers inside strings by their numeric value.
- strnatcmp ignores letter case.