A landing bucket gets thousands of 20 KB JSON files an hour and the hourly load keeps slowing. Why?
answer
- the bytes are not the problem
- cost per object, not per megabyte
- runtime tracks file count, not volume
- fix the writer or add a hop before the loader
basics
~20 sPer-file fixed cost is dominating: paginated listing, a request and open per object, and one scheduled task per file. The bytes are trivial; the overhead is not. Fix it by batching at the producer or adding a compaction hop before the load.
solid answer
~50 sEvery object costs the same fixed overhead regardless of size — a LIST page entry, a GET with its round trip and authentication, a file open and close, one split or task in the loader, and one row of tracking state. At 20 KB that overhead is many times the payload, so runtime scales with file *count*, not volume, and it degrades as the prefix grows. Fixes in order of preference: buffer at the producer and flush on size **or** elapsed time, whichever comes first; if the producer cannot change, add a compaction step that reads a closed window from the landing prefix, writes consolidated files to a separate prefix, and loads from there; coarsen the partition grain so you are not creating a partition per file. Compact only windows that are closed, and make compaction itself idempotent — deterministic output names, overwrite rather than append.
code
text · 9 lines# landing prefix, one hour
raw/events/ingest_hour=2026-08-20T14/ev-0000001.json 19 KB
raw/events/ingest_hour=2026-08-20T14/ev-0000002.json 21 KB
... 6,412 more objects ... total 128 MB in 6,414 objects
# after compacting the closed hour
compacted/events/ingest_hour=2026-08-20T14/part-0000.parquet 41 MB
compacted/events/ingest_hour=2026-08-20T14/part-0001.parquet 38 MB
# same rows, 2 objects; the loader now issues 2 GETs, not 6,414go deeper
Know that many tiny files cost far more than one large file of the same total size, because each object carries its own request and open cost.
Enumerate the per-object costs — listing pages, a request per file, a task per file, per-file tracking state — and explain why runtime scales with count rather than bytes.
Diagnose it from the symptom, then choose between producer batching and a compaction hop, handling the closed-window race, idempotent compaction output, and retention of the raw drop.
Name the tradeoff: compaction costs a rewrite and adds latency equal to its window, so if consumers need freshness the file interface itself is the wrong shape and a streaming transport is the real answer.
## Where the time actually goes The instinct is to blame the data volume. Thousands of 20 KB files an hour is tens of megabytes — nothing. The cost is entirely per-object and it comes from several places at once: **Listing.** LIST is paginated, typically a thousand keys per call, and the loader must enumerate before it reads. As the landing prefix accumulates, this grows even if the hourly arrival rate is flat. If the loader lists a parent prefix rather than a dated one, it re-walks history every hour. **Per-object requests.** Each file is at least one GET: a connection or connection-pool slot, TLS and auth overhead, a round trip whose latency is measured in tens of milliseconds and is largely independent of whether you asked for 20 KB or 200 MB. Ten thousand of those serialised is minutes of pure waiting. **Per-file work in the engine.** Distributed loaders split work by file. One tiny file becomes one task: scheduled, dispatched, started, its output collected. Task overhead is often larger than the work, and thousands of trivial tasks add scheduler pressure and a long tail. For self-describing formats there is a metadata read per file on top. **Per-file state.** If you track processed objects for idempotency, that ledger now grows by thousands of rows an hour, and every run queries it. **Downstream.** Small inputs tend to produce small outputs, so the target accumulates fragmentation too, and the cost shows up again at query time. The signature is diagnostic: runtime tracks file count rather than bytes, most of the wall clock elapses before meaningful rows are read, and the job gets slower month over month with flat volume. ## Fix it at the producer first The cheapest fix is upstream. Producers usually emit per-file because they emit per-event or per-request without thinking about it. Buffer instead, and flush on **whichever comes first** of a size threshold and a time threshold — size alone stalls a low-traffic tenant forever, time alone still produces tiny files when traffic is thin. Size the target to what your loader likes, which is orders of magnitude larger than 20 KB, not a marginal improvement to 200 KB. If the producer is genuinely per-event and cannot buffer, that is a signal the interface is wrong: this is a stream, and dropping a file per event into a bucket is an expensive way to imitate a queue. Say so, because the right answer is sometimes to change the transport rather than optimise around it. ## When you cannot change the producer Add a compaction hop. Read a **closed** window from the landing prefix, concatenate into a small number of larger files in a separate `compacted/` prefix, and point the loader at that. Keep the raw drop for replay and expire it by lifecycle rule once the compacted copy is durable. Three things to get right: **Only compact closed windows.** Compacting the current hour while files are still landing is a race: files arriving between your listing and your write are either missed or, if you re-compact later, counted twice. Wait for the window boundary plus a grace period, or for a completeness marker. **Make compaction idempotent.** Derive the output file names deterministically from the window and overwrite rather than append, so a failed and retried compaction converges instead of duplicating. Otherwise you have moved the duplicate-rows problem one hop upstream. **Keep the loader honest about which prefix is authoritative.** A loader that reads both raw and compacted prefixes double-counts. One prefix is the source of truth for loading; the other is an archive. ## Adjacent levers **Partition grain.** If the layout has minute-level or per-tenant-per-minute prefixes, you may have manufactured a partition per file. Coarsening the prefix reduces both listing cost and metadata pressure, and it is usually a smaller change than the compaction pipeline. **Parallel listing and reading.** Parallelism raises throughput but not efficiency — you are still paying the per-object cost, just concurrently, and you will hit request-rate limits and cost. Use it as relief, not as the fix. **Format.** Converting tiny files to a columnar format without consolidating them makes things slightly worse, because each file then also carries schema and footer metadata. Consolidate first, then convert. ## The tradeoff to name out loud Compaction costs a full read and rewrite of the data, plus a pipeline to operate, and it introduces latency equal to the window you wait for. If the consumer needs minute-level freshness, aggressive compaction fights that requirement directly — and that is the honest signal that a batch-file interface is the wrong shape for the workload. Do not compact so far ahead of the loader that you lose the raw archive or destroy the natural replay unit.
- Why should compaction only run on closed windows?Because files arriving between the listing and the write are silently missed, and a later re-compaction of the same window can pick them up a second time. Waiting for the window boundary plus a grace period, or for a completeness marker, makes the input set fixed before you read it, which is what lets the output be deterministic.
- Why is raising loader parallelism a weak fix here?It hides the cost rather than removing it. You still pay a request, a round trip and a task per object; you just pay them concurrently, until you hit request-rate limits or cost ceilings. It buys headroom while the file count keeps growing, so it postpones the incident instead of resolving it.
- How should a producer decide when to flush its buffer?On whichever comes first of a size threshold and a time threshold. Size alone starves a low-traffic tenant, whose file never fills and never lands; time alone still emits tiny files when traffic is thin but at least bounds latency. The pair bounds both file size and staleness.
- When is the right answer to stop using files entirely?When the producer emits per event and consumers want low latency. A file per event in a bucket is an expensive imitation of a queue: you pay object overhead for every message and then build compaction to undo it. At that point a streaming transport is the cheaper and simpler interface, and saying so is part of the answer.
It is the difference between one truck carrying ten thousand parcels and ten thousand couriers each carrying one — the parcels weigh the same, but you are paying for the couriers.
saying these in an interview costs you the question
- Blames data volume when runtime tracks file count
- Compacts the window that is still receiving files
- Adds parallelism and calls the problem solved
- Deletes raw files before the compacted copy is durable
- Loads from both the raw and compacted prefixes