skip to content

Which Apache Hudi table type is the default, and what does it write on an update?

level: juniorimportance: should knowfreq 52%

answer

  1. the default favours readers, not writers
  2. one changed row, one rewritten file
  3. nothing left to merge at query time
  4. the type name has three hyphenated words
  5. COPY_ON_WRITE is set once at creation

basics

~20 s

Copy-on-Write is the default Hudi table type. Updating a record rewrites the whole Parquet base file that holds it as a new file version, so readers just scan the newest base file per file group with no merging.

solid answer

~40 s

A Hudi table's type is fixed when the table is created, through `hoodie.datasource.write.table.type`, and the default is `COPY_ON_WRITE`. On an upsert, Hudi's index locates which file group already holds each incoming record key, reads that group's current Parquet base file, merges the new records in, and writes a **brand-new base file version**. File groups that no incoming record touched are neither read nor rewritten. The write becomes visible only when the completed `commit` instant lands on the `.hoodie` timeline. Readers then take the latest base file of each file group and read plain Parquet — no merge work at query time. The price is write amplification: changing one row rewrites an entire file. The other type, `MERGE_ON_READ`, appends updates to log files instead.

code

properties · 4 lines
properties
hoodie.datasource.write.table.type=COPY_ON_WRITE
hoodie.datasource.write.recordkey.field=order_id
hoodie.datasource.write.precombine.field=updated_at
hoodie.datasource.write.operation=upsert

go deeper

for a junior

Recall the two type names and the one-line difference: Copy-on-Write rewrites the Parquet file that holds the record, and it is the default. Be able to say why that makes reads fast.

for a middle

Explain the three steps of a Copy-on-Write upsert — index lookup, merge, new base file — and that only touched file groups are rewritten. Name write amplification as the tradeoff you accept.

for a senior

Be ready to judge when the default stops working: sparse updates scattered across many files, or commit intervals short enough that rewrite cost dominates. Tie base-file sizing settings to both scan speed and rewrite cost.

for a principal

Own the policy question of which tables get which type across a platform, and the fact that the choice is effectively irreversible without a rewrite. Frame it as buying read latency with write cost.

## Table type is a table-level decision Every Apache Hudi table is created as one of exactly two storage types: `COPY_ON_WRITE` (CoW) or `MERGE_ON_READ` (MoR). With the Spark datasource the option is `hoodie.datasource.write.table.type`; if you never set it you get `COPY_ON_WRITE`. The value is recorded in the table's properties, and it is not a knob you flip later on a loaded table — moving an existing table to the other type means rewriting it into a new table. Because it is the default, most Hudi tables you will meet in the wild are Copy-on-Write. ## How Hudi lays data out Under the table's base path there are partition directories holding data files, and a `.hoodie` directory holding metadata — chiefly the *timeline*, an ordered set of instants describing every action ever taken on the table. Within a partition, records are not scattered arbitrarily. Hudi groups them into **file groups**, each identified by a stable file ID. A file group owns a set of record keys over time; every version of that group is a **file slice**. In a Copy-on-Write table a file slice is simply one Parquet **base file**, whose name carries the file ID and the instant time of the commit that produced it. ## What an update actually does An upsert into a CoW table runs in three phases. 1. **Tagging.** Hudi asks its index, for each incoming record key, which file group already contains that key. Keys with no match are treated as inserts and are packed into new or under-sized existing base files. 2. **Merging.** For every file group that received at least one update, the writer streams the current base file, merges the incoming records into it, and uses the precombine field to decide which of two records for the same key wins. 3. **Writing.** It writes a completely new Parquet base file for that file group containing the untouched rows plus the updated ones. File groups nothing landed in are never read or rewritten. When every file is written, Hudi adds a completed `commit` instant to the timeline. Until that instant exists, no reader sees any of the new files — that is what makes the write atomic. The previous base file is *not* deleted at commit time; it stays as an older file slice so that time-travel readers and any in-flight query still work, and Hudi's cleaner removes it later under the configured retention policy. ## Why readers like it The newest base file of each file group is a complete, self-contained Parquet file. A reader lists the file groups in the surviving partitions, takes the latest slice of each, and hands plain Parquet to the engine. There is no log replay, no on-the-fly merge, no delete file to apply. That is why Spark, Hive, Presto and Trino read a Copy-on-Write table at essentially raw-Parquet speed, and why the "snapshot" and "read-optimized" views of a CoW table are identical — there is nothing extra to merge. ## The cost: write amplification Updating a single row inside a large base file rewrites that entire file. If your updates are sparse and scattered — the classic late-arriving-updates-across-many-partitions pattern — a small batch of changed rows can rewrite a large fraction of the table. Commit latency therefore scales with the *bytes touched*, not with the number of rows changed. That makes Copy-on-Write a poor fit for minute-level ingestion of random-key updates, and an excellent fit for batch-loaded tables that are read far more often than written. ## When to move off the default Choose `MERGE_ON_READ` when write latency and write amplification matter more than raw scan speed: high-frequency CDC ingestion, streaming writers committing every few minutes, or update patterns that touch a little of a lot of files. MoR appends updates as log blocks inside the file group and merges them either at read time or during a later compaction, trading query work for cheap writes. ## Configuration in practice ```properties hoodie.datasource.write.table.type=COPY_ON_WRITE hoodie.datasource.write.recordkey.field=order_id hoodie.datasource.write.precombine.field=updated_at hoodie.datasource.write.operation=upsert ``` Two more settings shape the result: `hoodie.parquet.max.file.size` caps how large a base file grows, and `hoodie.parquet.small.file.limit` tells the writer which existing base files are still small enough to absorb new inserts. Together they keep a Copy-on-Write table from degenerating into a directory of tiny files, which would erase the read advantage that justified choosing it.

  • If Copy-on-Write rewrites the base file, why is the old file still on storage after the commit?
    Because the old file is a previous file slice, not garbage. Time-travel and incremental queries need it, and any query that started before the commit is still reading it. Hudi's cleaner removes superseded slices later according to the retention policy, which is why an aggressive cleaner is what usually breaks time travel, not the commit itself.
  • Can you switch a Copy-on-Write table to Merge-on-Read after it is loaded?
    Not as an in-place property change. The table type is baked into the table's properties and into how its file slices are structured, so the practical path is to create a new table with the desired type and rewrite the data into it, then repoint the catalog. Plan the type per table before the first load.
  • Where does a Copy-on-Write table put brand-new records that match no existing key?
    They are treated as inserts. Hudi packs them into existing base files that are still under the small-file limit, and opens new file groups once those are full. That bin-packing is what keeps a frequently written Copy-on-Write table from producing a swarm of tiny Parquet files.

It is the difference between editing a document by printing a fresh clean copy each time versus scribbling changes on sticky notes the reader must reconcile. Copy-on-Write always prints the fresh copy.

saying these in an interview costs you the question

  • Says Copy-on-Write appends the new row and leaves the old file
  • Claims the whole partition is rewritten, not just affected files
  • Thinks Merge-on-Read is the default because writes are cheaper
  • Says readers must merge base and log files in Copy-on-Write
  • Believes the table type can be flipped with a config change later

context