A warehouse CSV import stores (int) $row['qty'], so '1,200' becomes 1 and 'n/a' becomes 0. How do you parse quantities safely in PHP 8?
answer
- cast converts, never validates
- FILTER_VALIDATE_INT returns int or false
- min_range option, leading zeros rejected
- is_numeric accepts '1e3' and '1.5'
- reject and report the row
basics
~20 sValidate before converting: filter_var($raw, FILTER_VALIDATE_INT, ['options' => ['min_range' => 0]]) returns an int or false, so '1,200' and 'n/a' are rejected instead of becoming 1 and 0. Normalise known formats like thousands separators explicitly, and report bad rows.
solid answer
~40 sThe bug is using a **conversion** as if it were a **validation**: `(int)` never fails, so `'1,200'` silently becomes `1` and `'n/a'` becomes `0`. The fix is to validate first. `filter_var($raw, FILTER_VALIDATE_INT, ['options' => ['min_range' => 0]])` returns an `int` for a well-formed integer and `false` otherwise. It trims surrounding whitespace, accepts a sign, rejects leading zeros such as `'012'`, fractions, exponents, separators and out-of-range values, and enforces the range you pass. `is_numeric()` is too loose for quantities: it accepts `'1.5'`, `'-3'` and `'1e3'`. `ctype_digit()` accepts only digits, returns `false` for `''`, and must be given a string. Thousands separators should be removed deliberately when the file's format is known, not guessed by a cast. Rows that fail are **collected and reported**, never stored as `0`.
code
php · 19 lines<?php
declare(strict_types=1);
$rows = [['sku' => 'BX-1', 'qty' => ' 12 '], ['sku' => 'BX-2', 'qty' => '0'],
['sku' => 'BX-3', 'qty' => '1,200'], ['sku' => 'BX-4', 'qty' => 'n/a']];
$stock = [];
$errors = [];
foreach ($rows as $line => $row) {
$qty = filter_var($row['qty'], FILTER_VALIDATE_INT, ['options' => ['min_range' => 0]]);
if ($qty === false) {
$errors[] = sprintf('line %d: invalid qty %s', $line + 1, var_export($row['qty'], true));
continue;
}
$stock[$row['sku']] = $qty;
}
var_dump($stock); // ['BX-1' => 12, 'BX-2' => 0]
print_r($errors); // lines 3 and 4 reported, nothing stored as 0go deeper
Know that an (int) cast never fails and that filter_var with FILTER_VALIDATE_INT returns an int or false.
Compare FILTER_VALIDATE_INT, ctype_digit and is_numeric by the strings each accepts, and use the range options and an === false check.
Design the import to validate, normalise only documented formats, report failing rows and fail loudly on format changes instead of writing zeros.
Set a data-contract approach for supplier files, with documented number formats and alerting on validation failure rates, so silent corruption cannot recur.
## The failure A nightly job imports stock levels from a supplier's CSV. Each line is read with `fgetcsv()`, which returns every field as a **string** (or `null` for a blank line's single field), and the quantity is stored with: ```php $stock[$row['sku']] = (int) $row['qty']; ``` The supplier starts formatting large numbers as `'1,200'` and writing `'n/a'` for discontinued items. The job keeps running without a single error, and the warehouse now shows 1 unit of a product it has 1,200 of, and 0 of products whose stock is unknown. ## Why the cast hides it An `(int)` cast is a **conversion**, not a **validation**: - it takes the leading numeric part (`'1,200'` becomes `1`); - it returns `0` for anything without one (`'n/a'` becomes `0`); - it raises **no warning** for either case, by design. Even `declare(strict_types=1)` does not help: strict mode governs scalar type checks when calling functions, not what an explicit cast does. The fix has to be an explicit check before converting. ## The validation options | Tool | Accepts | Rejects | Returns | |---|---|---|---| | `filter_var($s, FILTER_VALIDATE_INT)` | `'12'`, `' 12 '`, `'+12'`, `'-3'`, `'0'` | `'012'`, `'1.5'`, `'1e3'`, `'1,200'`, `''`, overflow | `int` or `false` | | `ctype_digit($s)` | `'12'`, `'012'` | `'-3'`, `'+12'`, `' 12'`, `''` | `bool` | | `is_numeric($s)` | `'12'`, `' 12 '`, `'1.5'`, `'-3'`, `'1e3'`, `'.5'` | `'1,200'`, `'n/a'`, `''` | `bool` | For a stock quantity, **`FILTER_VALIDATE_INT` with a range** is the best fit: ```php $qty = filter_var($raw, FILTER_VALIDATE_INT, ['options' => ['min_range' => 0]]); if ($qty === false) { $errors[] = "line {$line}: invalid quantity " . var_export($raw, true); continue; } ``` Details that matter: 1. It **returns the converted `int`**, so validation and conversion happen in one step. 2. It returns **`false`** on failure. Compare with `=== false`, because `0` is a valid quantity and is falsy. The `FILTER_NULL_ON_FAILURE` flag makes it return `null` instead, if that suits the code better. 3. `min_range` and `max_range` options enforce business limits, such as no negative stock. 4. It trims surrounding whitespace, which CSV exports often add. `is_numeric()` is right for "any number" but wrong for quantities: `'1.5'` and `'1e3'` would pass, and a later cast would turn `'1e3'` into `1000`. `ctype_digit()` is a reasonable strict alternative, but since PHP 8.1 passing it a non-string is deprecated, it accepts leading zeros, and it returns `false` for the empty string. ## Normalising known formats A thousands separator is a formatting decision, not noise. If the supplier's file specification says quantities use `,` as a thousands separator, remove it explicitly **before** validating: `str_replace(',', '', $raw)`. Do not do this blindly: in many locales `,` is the decimal separator, and `'1,5'` would become `15`. When the format is unknown or mixed, reject the row and ask. ## Reporting instead of storing zero The most important design choice is what happens to a bad value: - **Collect errors per row** with the line number and raw value, and continue with the rest of the file. - **Fail the whole import** if the error count passes a threshold, since a supplier format change usually affects every row. - **Never write a default of `0`** for unparseable stock; zero is a meaningful quantity that triggers reorders and "out of stock" labels. ## What a senior answer adds The code change is small; the judgment is in naming the root cause (conversion used as validation), picking a validator whose accepted grammar matches the business rule, keeping `0` distinguishable from failure, and making format changes by a data supplier loud instead of silent. The same approach applies to every numeric field that arrives as a string: prices, weights, IDs.
- Why compare the result of filter_var() with === false rather than using if (!$qty)?Because `0` is a valid quantity and is falsy. `if (!$qty)` treats a real zero exactly like a validation failure. `filter_var()` returns the integer `0` for `'0'` and `false` only for invalid input, so `$qty === false` is the only check that keeps the two apart.
- Would declare(strict_types=1) have caught the '1,200' problem?No. `strict_types` makes scalar parameter and return types reject mismatched types on calls made from that file; it does not change what an explicit `(int)` cast does. The cast still silently returns `1`. Only an explicit validation step catches it.
- When is is_numeric() the right validator for a CSV field?When the field may legitimately hold any number: negative values, decimals or scientific notation, such as a measured weight or a temperature. It is too permissive for quantities, where `'1.5'` or `'1e3'` should be rejected, and it still needs a following conversion such as `(float)` or `+0`.
saying these in an interview costs you the question
- Believes declare(strict_types=1) makes (int) casts reject bad strings.
- Says is_numeric() is the right check for whole-number quantities.
- Checks filter_var() results with if (!$qty), rejecting a real zero.
- Strips every comma before casting without knowing the file's number format.
- Stores unparseable quantities as 0 so the import can finish.