skip to content

In Apache Hudi, how does the bulk_insert operation differ from upsert?

level: middleimportance: should knowfreq 42%

answer

  1. one path checks whether the row already exists
  2. the other assumes the table is empty
  3. skipping the lookup is where the speed comes from
  4. great for the first load, dangerous on live data
  5. a third option sits between the two

basics

~20 s

upsert looks every incoming record up in the index and merges it into the file group already holding that key. bulk_insert skips the index lookup and small-file sizing entirely and writes new files fast, so it cannot deduplicate against existing rows.

solid answer

~50 s

`hoodie.datasource.write.operation` picks the write path. **upsert** — the default — tags every incoming record against the index to find which file group holds its record key, routes matches there as updates, and places the rest as inserts while trying to fill undersized files toward the target size. That tagging step is the expensive part. **bulk_insert** skips it: no index lookup, no small-file handling, no merge against existing data. It sorts the input according to `hoodie.bulkinsert.sort.mode` (global sort, partition-level sort, or none) and writes fresh files, which makes it the right way to load a large table for the first time or to rewrite one. The cost is that it cannot detect a key that already exists, so running it over live data creates duplicates. **insert** sits between them: no index lookup against existing data, but it does size files and can deduplicate within the incoming batch.

code

properties · 6 lines
properties
# initial load
hoodie.datasource.write.operation=bulk_insert
hoodie.bulkinsert.sort.mode=GLOBAL_SORT

# steady-state ingest afterwards
hoodie.datasource.write.operation=upsert

go deeper

for a junior

Know that Hudi has more than one write operation and that upsert is the default. Be able to say bulk_insert is for loading a table quickly and does not check whether a row already exists.

for a middle

Explain the tagging step that upsert performs and what bulk_insert skips: index lookup, merge against existing files, and small-file handling. Know that insert is a distinct middle option.

for a senior

Show the operational pattern — bulk_insert for the backfill, upsert for steady state — and be able to diagnose duplicates or a tiny-file explosion back to the operation and sort mode that produced them.

for a principal

Frame it as a cost and correctness contract for the platform: which sources are allowed the cheap path, how initial loads are performed safely, and what guardrails stop someone re-running a bulk load against live data.

## The three write paths `hoodie.datasource.write.operation` selects how a batch is written. Three values matter for ingest, and they differ mainly in how much work Hudi does *before* it writes a byte. ### upsert (the default) The full path. For every incoming record Hudi: 1. **Tags** it — asks the index which file group already holds this record key. This is a lookup against the whole table's index structure and is usually the dominant cost of the job. 2. **Routes** tagged records to their existing file group. On a Copy-on-Write table that means reading the current base file, merging the updates in, and writing a new version of it; on Merge-on-Read it means appending to a log file in that file group. 3. **Places** untagged records as inserts, preferring file groups that are below the target file size so the table does not accumulate small files. 4. **Deduplicates** within the batch using the precombine field, so one commit never writes the same key twice. You get correctness — one row per key — and self-maintaining file sizes, and you pay for both. ### bulk_insert The loading path. It does none of steps 1–3. There is no index lookup, so Hudi has no idea whether a key already exists; there is no read-modify-write of existing files, so nothing is merged; and there is no small-file handling in the upsert sense — file sizing comes from how the input is partitioned and sorted plus the target file size, not from probing existing files. What it does do is arrange the input before writing, via `hoodie.bulkinsert.sort.mode`: a global sort across the whole input, a sort within each output partition, or none. Global sorting produces well-bounded record-key ranges per file, which directly helps a bloom index prune candidate files later; it also costs a full shuffle. bulk_insert can also use a row-writing path that avoids converting each record into Hudi's internal record representation, which is a large part of why it is fast. Use it for: the initial load of a table, migrating an existing dataset into Hudi, or rewriting a table under a new key or partitioning scheme. Do not use it as a steady-state ingest path over live data — it will happily write a second copy of a key that already exists, and nothing will tell you. ### insert The middle path. No index lookup against existing data, so it also cannot update or deduplicate against what is stored, but it *does* apply small-file handling like upsert, and it can deduplicate within the incoming batch when combine-before-insert is enabled. It suits append-only fact streams where the source guarantees no re-delivery of a key and you still want reasonable file sizes. ## Choosing between them The decision is driven by one question: can this batch contain a key that already exists in the table? - **Yes, routinely** (CDC, mutable dimensions, late-arriving events) → upsert. The index cost is the price of correctness. - **Never, by construction** (immutable event log with a unique key from the source) → insert, and you get file sizing without paying for tagging. - **The table is empty or is being rebuilt** → bulk_insert, with global sort if a bloom index will serve later reads. A very common production shape is bulk_insert once for the backfill, then upsert forever after. Another is bulk_insert into a staging table followed by a single upsert into the target, when the incoming volume is far larger than the mutable slice it touches. ## Failure modes to name - **Duplicates after "a quick reload".** Someone re-ran the initial load with bulk_insert against a table that already had data. There is no index check, so every key exists twice and only a rewrite or a deduplicating job fixes it. - **A bulk_insert that leaves thousands of tiny files.** Sort mode `none` with a highly parallel input writes one file per task per partition. Either sort, or reduce write parallelism, or follow with clustering. - **An upsert job whose runtime is nearly all tagging.** That is not the write path's fault; it is the index choice, and the fix is a different index or a key layout that lets ranges prune. - **Assuming bulk_insert is "upsert but faster".** It is a different operation with different guarantees. Faster is a consequence of doing less, and the thing it skips is exactly what makes upserts correct. ## Interview framing Interviewers use this question to check whether you understand *why* Hudi upserts cost more than a plain Parquet write. The answer is the tagging step and the merge it enables. Candidates who describe bulk_insert as an optimisation flag rather than a semantically different operation get caught by the follow-up: what happens if you bulk_insert a key that already exists?

  • What happens if you run bulk_insert over a Hudi table that already contains the same record keys?
    You get duplicates. bulk_insert performs no index lookup, so it cannot know the key exists; it simply writes new files containing a second copy. Queries then return both rows, and later upserts may update only one of them. Recovering means rewriting the table or running a deduplicating job — there is no flag that retroactively merges them.
  • Why does the bulk_insert sort mode matter for later query performance?
    Sorting decides how record keys and partition values are distributed across output files. A global sort gives each file a narrow record-key range, which lets a bloom index prune most files during later upsert tagging, and gives min/max statistics that help query-time file skipping. Sort mode `none` interleaves keys everywhere, so almost every file becomes a candidate.
  • When would you choose insert rather than upsert for an ongoing pipeline?
    When the source guarantees each record key arrives once — an immutable event log with a producer-generated id, for example. `insert` skips the index lookup, so the job avoids the tagging cost, while still doing small-file handling. You lose the ability to correct or deduplicate against stored data, so it is only safe when re-delivery is genuinely impossible.

saying these in an interview costs you the question

  • Describes bulk_insert as just a faster upsert with the same guarantees
  • Thinks bulk_insert still deduplicates against existing records
  • Says upsert and insert differ only in performance, not semantics
  • Assumes bulk_insert handles small files the way upsert does
  • Cannot say which step makes upserts expensive

context