What does bin-packing compaction do to a table's data files, and how is the target size chosen?
answer
- fill bins up to a capacity, then write the bin
- fewer, larger files with identical rows
- no sort, no shuffle, just a repack
- both too small and too large hurt
- the old files linger until retention cleanup
basics
~20 sBin-packing compaction groups many small files into bins that add up to a target size, rewrites each bin as one new file, and atomically swaps the table's file list. It preserves row content and order-agnostic semantics; the target is a band, commonly a few hundred megabytes.
solid answer
~50 sBin-packing treats each output file as a bin with a capacity — the target file size — and greedily fills bins with input files until they are full. It reads the small files, writes fewer larger ones, and commits a metadata change that removes the old files from the table and adds the new ones. Rows are unchanged and unmoved between files in any meaningful sense; it is purely a physical repack, so it is the cheapest form of compaction: no sorting, no shuffle, often a per-partition local operation. The target size trades read parallelism against per-file overhead. Too small and overhead dominates; too large and you get fewer splits than workers, plus more data rewritten per row-level change. A few hundred megabytes is the usual band, tuned down for tables read with high concurrency and up for large scan-heavy tables.
go deeper
Recall that compaction rewrites many small files into fewer larger ones and swaps them in atomically, and that rows and query results are unchanged by it.
Explain the bin-packing mechanic — fill bins to a target capacity, rewrite each bin, commit the swap — and argue both directions of the size tradeoff: overhead below the band, lost parallelism and rewrite cost above it.
Be ready to defend a specific target size against a described workload, to explain why compaction is recurring rather than one-shot, and to reason about running it safely alongside live ingestion.
Own the maintenance policy across a platform: which tables get compacted, on what cadence, at what target, who pays for the compute, and how the schedule is derived from ingest cadence rather than guessed per team.
## The mechanic "Bin-packing" is the classic algorithm applied to file layout. Each **bin** is a future output file with a capacity equal to the configured **target file size**. The compactor lists the table's current data files, filters to those meaningfully below target, and greedily assigns them to bins until each bin's summed size approaches capacity. Each bin is then read and written out as a single new file. Crucially, this is a **rewrite, not an edit**. Data files in these formats are immutable — you cannot open a 3 MB file and append 200 MB into it. The compactor produces entirely new files and then performs a metadata commit that says: these old files are no longer part of the table, these new files are. Because that commit is atomic, readers see either the old file set or the new one, never a mixture, and never a moment where rows are missing or doubled. ## What bin-packing deliberately does *not* do It does not reorder rows across the whole table. It does not cluster by a column. It does not repartition. It is the minimal operation that fixes file count, and that minimality is the point: it requires no shuffle, so it can run per partition, in parallel, reading and writing locally with no cross-node data movement. That makes it dramatically cheaper than sort-based rewriting, which is why it is the default flavor of compaction and the one you schedule frequently. Sorting or clustering during a rewrite is a strictly more expensive variant, used when pruning — not file count — is the problem. The two are often conflated in interviews; keep them separate. ## Choosing the target size The target is a **band**, and both ends of it hurt. **Too small** and you are back in the small-files regime: per-file planning entries, per-file object-store requests, weak encoding, coarse statistics. **Too large** hurts in several distinct ways: - **Read parallelism collapses.** Engines assign work by split, and splits usually derive from files (or row groups within them). If a partition is one enormous file and you have 200 idle workers, most of them do nothing. - **Row-level changes get expensive.** Under a copy-on-write style update, changing one row rewrites the whole file containing it. Bigger files mean more bytes rewritten per row touched — the write-amplification tax. - **Failure and retry cost rises.** A task that dies at 90% of a 5 GB file redoes all of it. - **Memory pressure on writers.** Building large row groups needs buffer. The practical band most platforms land in is a few hundred megabytes per file — large enough that per-file overhead is a rounding error, small enough to keep splits plentiful and rewrite cost bounded. Skew the choice by workload: - **Scan-heavy, few concurrent queries, huge tables** → toward the upper end. Sequential throughput wins. - **Many small concurrent queries, or heavy row-level mutation** → toward the lower end. Parallelism and rewrite cost win. - **Highly selective point-lookup workloads** → smaller files with tight statistics can prune better than a few giant ones, because a big file's min/max necessarily covers a wider range. ## Why compaction never fully "finishes" A freshly compacted table drifts back the moment ingestion resumes. If a streaming job keeps writing every minute, the table accrues new small files continuously, and compaction is a **recurring maintenance job**, not a one-time repair. This is the single most common operational surprise: teams run compaction once, see the improvement, and are confused when the problem returns a week later. The steady state is a pipeline where new data arrives in small files, and a scheduled compactor sweeps them into target-sized files behind the ingest edge. Tuning that schedule is a tradeoff between freshness of layout and cost of rewriting. ## The commit and the leftovers After compaction commits, the *old* small files are still physically present in storage. They must remain readable for a while — running queries planned against the old file list are still reading them, and time-travel or rollback to a pre-compaction state needs them. A separate, retention-aware cleanup step removes them once they are outside the retention window. Deleting them immediately is a classic self-inflicted outage: in-flight readers get file-not-found errors. ## A note on idempotence and concurrency Compaction competes with writers. If an ingest commit lands between the moment the compactor read the file list and the moment it tries to commit, the compaction commit may conflict and have to be retried against the new state. Well-behaved compactors detect this and either retry or drop the affected group, because blindly committing would resurrect deleted rows or lose newly added ones. This is why compaction is safe to run concurrently with ingestion *only* because the table format's commit protocol validates the assumption the compactor made. ```sql -- Conceptually, compaction is: read a set of files, write fewer, -- and swap them in one atomic metadata commit. -- Row-level content and query results are identical before and after. ```
- Why is bin-packing cheaper than a sort-based rewrite of the same data?Bin-packing needs no data movement between nodes — each bin's inputs can be read and rewritten locally, usually within one partition. A sort-based rewrite must shuffle rows globally so that the sort key is ordered across output files, which costs a full network exchange plus spill. Bin-packing fixes file count; sorting fixes pruning, and you pay for it.
- How does compaction interact with a concurrent ingest job writing to the same table?The compactor reads a file list, does its work, then commits a swap of exactly those files. If ingestion committed in the meantime, the format's concurrency check detects that the compactor's assumption is stale and rejects or forces a retry. Newly added files are simply not part of the compacted group and get swept in the next run.
- Why can't compaction delete the old small files as part of the same commit?It removes them from the *table*, but not from storage. Queries that already planned against the old file list are still reading those objects, and rollback or time travel to earlier states needs them. Physical deletion is a separate step gated by a retention window longer than the longest expected query.
- Does compaction change query results?No. It is a pure physical reorganization: same rows, same values, same schema. Only file boundaries, file count, statistics and possibly row order within files change. If results differ before and after, something is wrong — most likely a concurrent write was lost or a delete was resurrected.
Repacking a shelf of half-empty boxes into full ones: the contents are unchanged, you just stop paying to carry air.
saying these in an interview costs you the question
- Saying compaction appends small files into an existing large file
- Confusing bin-packing with sorting or clustering the data
- Claiming there is one universally correct target file size
- Assuming compaction is a one-time fix rather than a recurring job
- Expecting the old files to disappear from storage at commit time