In PHP, why shouldn't an upload handler trust $_FILES['cv']['type'] or the file's extension, and how do you check the real type?
answer
- type is the client's Content-Type header
- the extension comes from the client's name
- finfo reads magic bytes from tmp_name
- FILEINFO_MIME_TYPE, returns string|false
- allowlist, then a server-chosen name
basics
~20 sBoth type and the name's extension come from the client and can say anything. Detect the type from the file's bytes with finfo and FILEINFO_MIME_TYPE on tmp_name, compare it with an allowlist, and derive the stored extension from the detected type.
solid answer
~40 s`$_FILES['cv']['type']` is copied from the `Content-Type` header of the multipart part, and `name` is the client's file name; PHP checks neither, so `shell.php` can arrive as `cv.pdf` with type `application/pdf`. Detect the type from content instead, using the bundled fileinfo extension: `(new finfo(FILEINFO_MIME_TYPE))->file($_FILES['cv']['tmp_name'])` returns a MIME string such as `application/pdf`, or `false` on failure. Compare it strictly against an **allowlist**, pick the stored extension from your own type-to-extension map, generate the stored name on the server, and keep uploads outside the web root. Magic-byte detection proves what a file looks like, not that it is harmless: a file can be valid PDF and still carry other content, so never execute or include uploads. Since PHP 8.1 `finfo_open()` returns a `finfo` object, and PHP 8.5 deprecates `finfo_close()`.
code
php · 21 lines<?php
declare(strict_types=1);
const ALLOWED_CV_TYPES = [
'application/pdf' => 'pdf',
];
$cv = $_FILES['cv'] ?? null;
if (!is_array($cv) || $cv['error'] !== UPLOAD_ERR_OK || $cv['size'] > 5 * 1024 * 1024) {
exit('Please upload a PDF of at most 5 MB.');
}
$mime = (new finfo(FILEINFO_MIME_TYPE))->file($cv['tmp_name']);
if ($mime === false || !array_key_exists($mime, ALLOWED_CV_TYPES)) {
exit('Only PDF CVs are accepted.');
}
$stored = bin2hex(random_bytes(16)) . '.' . ALLOWED_CV_TYPES[$mime];
if (!move_uploaded_file($cv['tmp_name'], '/srv/jobs/storage/cvs/' . $stored)) {
exit('The CV could not be stored.');
}go deeper
Recall that type and name come from the client, and that finfo with FILEINFO_MIME_TYPE on tmp_name reads the real type.
Explain the full sequence: error, size, finfo detection, strict allowlist, server-chosen extension and name, storage outside the web root.
Explain what content detection cannot prove - polyglots, container formats, malicious but valid files - and why storage and naming make mislabelled files inert.
Decide how uploads are isolated across the platform: separate storage or domain for serving, scanning policy, and which formats the product accepts at all.
## What the client controls In a `multipart/form-data` upload, each file part carries headers written by the client: - a `filename`, which PHP stores in `$_FILES['cv']['name']` after stripping any directory part; - a `Content-Type`, which PHP copies into `$_FILES['cv']['type']` without looking at the bytes. A browser usually guesses the type from the extension, and any other client can send whatever it likes. So both fields are **claims**: `evil.php` renamed to `cv.pdf` and sent with `application/pdf` passes an extension check and a `type` check alike. Only `tmp_name`, `error` and `size` are PHP's own findings. ## Detecting the type from content The **fileinfo** extension, bundled with PHP, identifies a file from its leading bytes and structure (its "magic" signatures), the same approach as the Unix `file` command: | API | Signature (PHP 8.5) | On failure | |---|---|---| | `new finfo(int $flags = FILEINFO_NONE, ?string $magic_database = null)` | object API | - | | `finfo::file(string $filename, int $flags = FILEINFO_NONE, $context = null)` | returns `string\|false` | `false` | | `finfo_open()` + `finfo_file()` | procedural twins | `false` | | `mime_content_type($filename)` | one-shot helper | `false` | Pass `FILEINFO_MIME_TYPE` to get only the MIME type (`application/pdf`) rather than the type plus charset. Two version notes: since PHP 8.1 the procedural functions work with `finfo` **objects** rather than resources, and PHP 8.5 **deprecates `finfo_close()`**, since the object is freed automatically. ## A checking sequence for CV uploads 1. Require `$_FILES['cv']['error'] === UPLOAD_ERR_OK`. 2. Check `size` against your business limit. 3. Detect the MIME type from `tmp_name` with `finfo`; treat `false` as a rejection. 4. Compare the detected type with an **allowlist** using strict comparison, never a denylist of dangerous types. 5. Choose the stored extension from **your own** map (`'application/pdf' => 'pdf'`), not from the client's name. 6. Generate the stored file name on the server and move the file into a directory the web server does not execute or serve directly. For images, `getimagesize()` gives a second, format-aware check (dimensions and image type) on top of `finfo`. ## Common mistakes - **Checking `type` with `str_starts_with($type, 'application/pdf')`.** It is still the client's claim, however it is compared. - **Loose allowlist checks.** `in_array($mime, $allowed)` without `true` compares loosely; use strict comparison or array keys. - **Running `finfo` on `name`.** The client's name is not a path on the server; detection must read `tmp_name`. - **Keeping the client's extension after a successful check.** The extension decides how a web server treats the file, so it must come from the detected type. - **Treating `false` as "unknown but allowed".** A detection failure is a rejection. ## What content detection does not prove Magic-byte detection answers "what does this file look like?", not "is it safe?": - **Polyglots.** A file can be a valid PDF and also contain a script or HTML fragment elsewhere in its bytes. If that file is ever served from a directory where the server executes or renders it by extension, the claim of being a PDF does not help. - **Container formats.** Office documents are ZIP archives internally; depending on the magic database, detection may report the specific Office type or a generic archive type. Test each allowed format with real sample files on your deployment before writing the allowlist. - **Content risks.** A genuine PDF can still be malicious to the reader's viewer. Malware scanning, if required, is a separate step. That is why the last two steps of the sequence matter as much as the type check: a server-chosen name, a server-chosen extension and storage outside the executable web root make a mislabelled file inert. How to sanitise a client file name when you must keep it, and how to guard the storage path, are separate topics.
- Why pick the stored extension from the detected type instead of the client's file name?The client's name is only a claim. Keeping `cv.php` because the bytes looked like a PDF would still store a `.php` file, which some server setups execute. Mapping the detected, allowlisted MIME type to an extension you choose guarantees the stored name matches what you checked.
- In PHP 8.5, should you still call finfo_close() after finfo_file()?No. Since PHP 8.1 `finfo_open()` returns a `finfo` object, freed automatically when it goes out of scope, and PHP 8.5 deprecates `finfo_close()`, so calling it emits a deprecation. Use `new finfo(FILEINFO_MIME_TYPE)` and let it go out of scope.
saying these in an interview costs you the question
- PHP sets $_FILES['cv']['type'] by inspecting the uploaded bytes.
- Checking the extension of the uploaded name is enough to know the type.
- A file that finfo reports as application/pdf is safe to serve anywhere.
- A denylist of dangerous MIME types is as good as an allowlist.
- finfo_close() must be called to avoid leaking resources in PHP 8.5.