In Laravel, what does Storage::append() actually do on an S3 disk, and why are append() and allFiles() expensive on remote disks?
answer
- no native append on object storage
- fileExists, get, then put again
- separator defaults to PHP_EOL
- concurrent appends lose lines
- allFiles: full recursive listing, sorted
basics
~10 sStorage::append() downloads the whole file, adds the new data and uploads everything again, so each call costs the full file and concurrent calls can lose lines. allFiles() lists every object recursively into one array.
solid answer
~40 s`append($path, $data, $separator = PHP_EOL)` is implemented in Laravel's adapter, not in storage: if `fileExists($path)` it calls `put($path, $this->get($path).$separator.$data)`, otherwise `put($path, $data)`. On S3 that is a metadata request, a download of the whole object and an upload of the whole object for every line, so an audit log that grows to 50 MB costs 100 MB of transfer per append. It is also a read-modify-write, so two requests appending at once can each read the old content and one line is lost. `prepend()` has the same shape. `allFiles($dir)` asks Flysystem for a recursive listing, which on S3 means paging through every key under the prefix, then sorts and returns them as one array. Write one object per event instead of appending, and iterate `listContents()` lazily or query your database instead of listing buckets.
code
php · 20 lines<?php
use Illuminate\Support\Facades\Storage;
use Illuminate\Support\Str;
// Costly and racy on S3: downloads and re-uploads the whole log each time
Storage::disk('s3')->append('audit/downloads.log', $line);
// One object per event: no read, no overwrite, no lost updates
Storage::disk('s3')->put(
'audit/'.now()->format('Y/m/d').'/'.Str::uuid().'.json',
json_encode(['user' => $userId, 'document' => $documentId, 'at' => now()->toIso8601String()]),
);
// Lazy iteration instead of allFiles() for a cleanup job
foreach (Storage::disk('s3')->listContents('tmp', true) as $item) {
if ($item->isFile() && $item->lastModified() < now()->subDay()->getTimestamp()) {
Storage::disk('s3')->delete($item->path());
}
}go deeper
Recall that append() rewrites the whole file and that allFiles() lists everything recursively into one array.
Explain append() as fileExists, get and put, the lost-update race it creates, and why allFiles() must collect the full listing to sort it.
Replace append-to-object with one object per event or a database table, iterate large prefixes lazily, and keep an index instead of listing buckets.
Decide which data belongs in object storage versus a database or log pipeline, weighing cost per call against durability and audit needs.
## append() is read-modify-write A clinic portal must record who downloaded which medical document. A tempting first version writes each event to a log file on the documents disk: ```php Storage::disk('s3')->append('audit/downloads.log', $line); ``` Laravel's `FilesystemAdapter::append($path, $data, $separator = PHP_EOL)` does this: 1. `fileExists($path)` — on S3, a metadata request. 2. If it exists: `get($path)` — download **the whole object** into a PHP string. 3. `put($path, $existing.$separator.$data)` — upload **the whole object** again. 4. If it does not exist: `put($path, $data)`. `prepend()` is the mirror image. Object storage has no "append these bytes" operation in Flysystem's API, so the adapter emulates it, and the emulation scales with the file, not with the line. ## What that costs | Log size | Transfer per append | Memory per append | |---|---|---| | 10 KB | ~20 KB | ~20 KB | | 50 MB | ~100 MB | ~100 MB | | 500 MB | ~1 GB | beyond typical `memory_limit` | Latency grows with the file too, so a download link that once took milliseconds to log starts taking seconds. ## Lost updates Because append is read, modify, write, two requests arriving together behave like this: 1. Request A reads the log (N lines). 2. Request B reads the log (N lines). 3. A writes N+1 lines. 4. B writes a different N+1 lines, overwriting A's. A's event is gone, with no error. The local driver writes files with an exclusive lock, but the lock covers only the final write, not the read before it, so the race exists there too. For an audit trail of medical-record access, a silently lost entry is a compliance problem, not a cosmetic one. ## allFiles() materializes everything `allFiles($directory)` is `files($directory, true)`: - Flysystem lists contents **recursively**; on S3 that means paging through every key under the prefix. - Laravel filters to files, **sorts by path** (which requires collecting the whole listing first) and returns a plain array. For a prefix holding hundreds of thousands of scans, that is many listing requests and one very large array, often inside a web request. `files()` without recursion and `directories()`/`allDirectories()` have the same shape at a smaller scale. Related: `exists()` means "file or directory", so on S3 a miss costs a metadata request plus a listing request; `fileExists()` makes only the first. ## Better patterns - **One object per event.** Write `audit/2026/09/29/{uuid}.json` with `put()`; nothing is ever re-read or overwritten, and concurrent writes cannot collide. - **Use the right store.** Access events belong in a database table or a log channel; object storage is for files. - **Batch, then write.** Collect lines in a queue job and write a new object per batch. - **Iterate lazily.** Calling `Storage::disk('s3')->listContents($dir, true)` is forwarded to Flysystem and returns a `DirectoryListing` you can `foreach` without Laravel sorting it into one array, which suits cleanup and backfill jobs. - **Keep an index.** If the app needs to show "all scans for this patient", store paths in the database when writing and query there, instead of listing the bucket. ## Measuring the cost before it bites Remote-disk costs hide well in development, where every disk is local and fast. Ways to surface them early: - Count storage calls per request in a staging environment pointed at the real bucket; an endpoint that makes dozens of metadata and listing calls will be slow in production. - Look for loops that call `exists()`, `size()` or `lastModified()` per file; each is a network round trip on S3. - Search the codebase for `append(`, `prepend(` and `allFiles(` on disks that may be remote, and check how large those files and prefixes can grow. - Prefer data already in your database (sizes, types, owners) over asking the storage service again. ## When append() is fine - small files on the local disk written by a single process, such as a script's own progress notes; - development and tests, where size and concurrency do not matter. The API looks identical on every disk; the cost does not. That is the real lesson interviewers probe: an abstraction that makes remote storage look local hides the price of each call.
- Does the local driver's exclusive file lock make Storage::append() safe under concurrency?No. The lock applies to the final write only. `append()` first reads the file with `get()` and then writes the combined contents, so two requests can both read the old version before either writes, and the second write discards the first request's line.
- Why can Storage::exists() cost two requests on S3 while fileExists() costs one?`exists()` is true for a file or a directory, so Flysystem checks the object first and, on a miss, runs a listing to see whether the path is a directory prefix. `fileExists()` only checks the object, which is one request.
saying these in an interview costs you the question
- Storage::append() sends only the new bytes to S3
- append() is atomic, so concurrent requests can never lose lines
- allFiles() streams results lazily, so memory stays flat for any prefix
- The local disk's file lock makes append() safe under concurrency
- Storage methods cost the same on every disk because the API is identical