How would you set ingest batching and compaction policy for a continuously loaded, dashboard-queried columnar table?
answer
- freshness, read cost, write amplification
- ask what stale actually costs the business
- stop merging once files hit the target
- recent data is hot, old partitions are settled
- compaction bytes divided by ingested bytes
basics
~20 sStart from the freshness the dashboards actually need, size ingest batches to that interval, and let compaction absorb whatever fragmentation remains. Then bound write amplification by choosing how many merge passes data goes through and how large the final files get.
solid answer
~50 sTreat it as a three-way budget between **freshness, read cost and write amplification**. Establish the real freshness requirement first — most dashboards tolerate minutes, not seconds — and set the ingest flush interval to it, so files land as close to the target size as the arrival rate allows. Whatever fragmentation remains is compaction's job: merge small recent files into progressively larger ones, stopping once files reach the target, because merging beyond that buys nothing and burns compute. Bound the write amplification explicitly: each merge level rewrites the data once, so a shallow tiered policy is cheap on writes but leaves more files for readers, while an aggressive policy keeps reads fast and pays more compute. Isolate compaction from query and ingest compute so it cannot starve either. Coarsen partitioning if thin partitions cap file sizes. Finally, decide who pays: compaction cost is real and should be attributed, not hidden.
code
text · 5 linesfreshness SLA : 5 min -> writer flushes at 5 min or 256 MB
target file size : 256 MB -> files at target excluded from merging
merge scope : partitions younger than 7 days
write amplification : compaction_bytes / ingest_bytes, alert above 3x
compute isolation : dedicated compaction pool, off the query pathgo deeper
Understand the shape of the trade: writing more often means fresher but smaller files, and someone has to merge them later. Know that compaction is a real background job that costs compute.
Explain how batch interval, target file size and merge policy interact, and why merging files that already meet the target is wasted work.
Set and operate the policy: derive the batch interval from a real freshness requirement, bound compaction scope to recent partitions, isolate its compute, and monitor file size, planning share and compaction lag.
Own the budget across the platform. Quantify write amplification, attribute compaction cost to the workload that causes it, challenge unexamined real-time requirements, and define the metrics that trigger revisiting the policy as volume grows.
## Frame it as a budget, not a setting There is no correct compaction interval in the abstract. Three quantities trade against each other and the workload picks the point: - **Freshness** — how long after an event lands it may become visible. - **Read cost** — how much per-file overhead and dead data queries carry. - **Write amplification** — how many times the same bytes are written before they settle. Move one and at least one other moves. Longer ingest batches produce larger files, cutting both read cost and merge work, at the price of freshness. Aggressive compaction cuts read cost and pays with compute. Doing nothing is cheapest today and worst in six months. ## Start with the freshness requirement, honestly Ask what a stale dashboard actually costs. Most executive dashboards refresh hourly and are consumed daily; "real time" is usually an aspiration inherited from the requirements document rather than a decision anyone made. Operational alerting genuinely needs seconds, but it is often a different, much smaller table. Once the number is real, set the writer's flush trigger to it — flush at whichever comes first, the target file size or the freshness deadline. If arrivals are heavy enough that the size trigger fires first, you have won: files land well-sized and compaction has little to do. If the deadline fires first, you know in advance that compaction must clean up, and roughly how much. ## Sizing the merge policy Compaction reads small files and writes larger ones. The two classic shapes: - **Tiered/size-based**: merge files of similar size into one bigger file, repeatedly. Write amplification is roughly the number of levels the data passes through, which is low. Readers may face several overlapping files covering the same key range. - **Levelled/aggressive**: maintain a tighter, more ordered file layout by rewriting more often. Reads are cleaner and skipping is more effective; writes are rewritten more times. Two policy points matter more than the label. First, a **stop condition**: once a file reaches the target size, exclude it from further merging. Merging two 400 MB files into 800 MB rewrites 800 MB and buys nothing, and this omission is the most common way a compaction bill runs away. Second, **recency awareness**: recent data is small, fragmented, and hot; old partitions are settled and should be compacted once and then left alone. A policy that periodically re-examines every partition of a multi-year table wastes almost all of its work. ## Isolate the compute Compaction competes with ingestion for write bandwidth and with queries for CPU and I/O. Give it its own compute pool or its own window. This is also what makes the cost visible: when compaction runs on the same resources as dashboards, its cost hides inside "queries got slower" and nobody can reason about it. A separate pool turns it into a line item you can size, schedule and defend. ## Check partitioning before tuning anything else If the table is partitioned finely enough that each partition receives only a trickle per interval, no batching produces large files and compaction has nothing to merge across. Very fine partitioning is the most common root cause of a compaction policy that appears not to work. Coarsening the partition granularity — hour to day, or dropping a high-cardinality dimension from the partition key — fixes more small-file problems than any writer tuning. ## The hot/cold split when freshness genuinely is tight When a consumer truly needs sub-minute visibility and the table is also queried heavily over history, split the concerns: a small recent region accepting frequent small writes, and the historical body kept in well-sized files, with a scheduled promotion that folds the hot region into the cold one. Queries union the two. This costs a little query complexity and buys both freshness and read efficiency, and it localises the fragmentation to a region small enough that its overhead is bounded. ## Instrument the policy Make the budget observable or it will drift: - average file size and file count per partition, with an alert threshold; - planning time as a share of query time; - bytes written by compaction versus bytes ingested — this **is** the write-amplification factor, and it is the number to defend in a cost review; - compaction lag: how far behind the merge queue is running. ## What a strong answer sounds like It refuses to name an interval without first asking about freshness and arrival rate. It states the three-way trade explicitly, gives a stop condition for merging, separates hot recent data from settled history, insists on isolated compute so the cost is attributable, and names partition granularity as the thing to check before tuning anything. It also concedes that the right answer changes as volume grows, and specifies which metric would trigger revisiting it.
- How do you decide whether compaction is costing too much?Track bytes written by compaction divided by bytes ingested — the write-amplification factor — against the read-side benefit. A small multiple is normal; a large one usually means files at target size are still being merged, or settled historical partitions are being re-examined. Both are policy bugs, not capacity problems.
- A team insists on second-level freshness for a table queried mostly over months of history. What do you propose?Split the table's physical treatment: a small hot region accepting frequent small writes for recent data, and the historical body kept in well-sized files, with scheduled promotion folding one into the other. Queries read the union. Fragmentation stays bounded to a region small enough that its per-file overhead does not matter.
- Why isolate compaction onto its own compute rather than running it alongside queries?Two reasons. Operationally it stops merges from starving dashboards or ingest during peaks. Economically it makes the cost visible: shared compute hides compaction inside 'queries got slower', while a dedicated pool turns it into a number you can size, schedule and defend in a cost review.
- What would make you revisit a policy that is working today?A material change in arrival rate, partition cardinality or query pattern. Concretely: average file size per partition drifting below target, planning time growing as a share of query time, write amplification crossing its threshold, or compaction lag growing. Each is an alert, and each maps to a specific knob rather than a general retune.
It is warehouse restocking: unpack pallets as they arrive and the floor fills with cartons; consolidate constantly and you pay staff to move the same boxes twice. You set a shelf size and a restock cadence, not a rule against either.
saying these in an interview costs you the question
- Names a compaction interval before asking about freshness
- Merges files that already meet the target size
- Re-scans settled historical partitions every cycle
- Runs compaction on the query compute and hides its cost
- Treats real-time freshness as a given rather than a cost