skip to content

In Apache Hudi, what is a file group and how does a file slice relate to it?

level: middleimportance: nice to knowfreq 32%

answer

  1. records do not move between drawers
  2. the identifier in the file name never changes
  3. one group, many versions over time
  4. the version is what compaction rewrites
  5. the index maps a key to this identifier

basics

~20 s

A Hudi file group is a set of records inside a partition, identified by a stable file ID and owning those record keys over time. A file slice is one version of that group: a base Parquet file plus, in Merge-on-Read, its log files.

solid answer

~40 s

Inside a partition, Hudi does not scatter records arbitrarily. It assigns each record key to a **file group**, named by a file ID that never changes. Every write that touches the group produces a new **file slice** — a version of the group tagged with the instant time that created it. In `COPY_ON_WRITE`, a slice is a single Parquet base file. In `MERGE_ON_READ`, a slice is a base file plus the log files appended after it. The file ID is what makes upserts possible: the index maps a record key to a file group, so the writer knows exactly which files to merge into. It is also the unit of work for compaction, of versioning for the cleaner, and of file sizing for the small-file problem.

code

text · 5 lines
text
country=US/
  a1b2c3d4-0_0-24-25_20240115103000123.parquet    group a1b2c3d4-0, slice @103000123
  a1b2c3d4-0_0-31-44_20240115110000789.parquet    same group, newer slice @110000789
  9f8e7d6c-1_0-24-26_20240115103000123.parquet    a different file group
  .9f8e7d6c-1_20240115103000123.log.1_0-31-45     log attached to that group's slice

go deeper

for a junior

Recall that Hudi organizes records inside a partition into groups with stable identifiers, and that each version of a group is a slice. That is enough to read a directory listing.

for a middle

Explain what a slice contains under each table type, how the identifier in the file name ties logs to their base file, and why the index depends on the grouping being stable.

for a senior

Connect the structure to operations: compaction works per slice, retention is expressed in slices, and file sizing settings decide when a new group is opened. Diagnose a table by grouping its listing by file ID.

for a principal

Own retention and layout policy in these terms — how many slices of history you keep, what that costs, and which operations are allowed to rewrite records into new groups and thus disturb index locality.

## Why the grouping exists at all A plain directory of Parquet files has no notion of *where a given row lives*. Apache Hudi needs that notion, because its core operation is an upsert by record key: given an incoming key, the writer must find the existing copy in order to merge or replace it. The structure that provides it is the **file group**. A file group is a set of records inside one partition, identified by a **file ID** — a stable identifier baked into the names of every file belonging to the group. Once a record key is assigned to a file group, it stays there for the life of the table (barring operations that explicitly replace file groups, such as clustering). Hudi's index is essentially a map from record key to file group, and the file ID is the value it returns. ## The slice is the version A file group changes over time, and each version of it is a **file slice**. A slice is identified by the instant time that created its base file. What a slice physically contains depends on the table type: - **Copy-on-Write.** A slice is exactly one Parquet **base file**. Every write that touches the group merges the group's contents and emits a new base file, so a new instant time means a new slice. - **Merge-on-Read.** A slice is a base file *plus* any **log files** appended after it. Ingestion appends log blocks rather than rewriting, so the slice grows in place until compaction reads base plus logs and emits a new slice with a fresh base file and no logs. The naming makes this readable straight off a directory listing: base files carry the file ID and the instant that wrote them, and log files carry the same file ID together with the base instant of the slice they belong to. Group the listing by file ID and you can see every version of every group. ## What the structure buys you Four things in Hudi are defined in terms of file groups and slices. **Upserts.** Because the index resolves a key to a file ID, an update touches exactly one file group, not the whole partition. This is what makes record-level updates on immutable files tractable at all. **Compaction.** In Merge-on-Read, compaction operates per file slice: read this slice's base file and its logs, write a new base file. That makes compaction incremental and parallelizable — slices with the most log data can be prioritized without rewriting the table. **Versioning and cleaning.** Older slices of a group are not deleted at commit time. They remain so that time-travel and incremental readers, and any query already in flight, still have consistent files to read. Hudi's cleaner removes superseded slices later according to the configured retention. The relationship is direct: retention policy is expressed in slices per group and instants of history, and cleaning too aggressively is what breaks time travel. **File sizing.** The small-file problem is managed at group granularity. The writer bin-packs new inserts into existing base files still under `hoodie.parquet.small.file.limit`, opening new file groups only when existing ones approach `hoodie.parquet.max.file.size`. That is why a well-configured Hudi table does not spray a new tiny file per commit the way a naive append-only writer does. ## Common confusions A file group is **not** a partition. A partition typically contains many file groups; the partition path is a directory, the file group is a logical set of records inside it identified by a file ID. A file slice is **not** a snapshot of the table. Slices are per file group. A table-level read assembles the current slice of every relevant file group; different groups can have been last written at different instants, and that is normal. A log file does **not** belong to a partition in general — it belongs to a specific file group and a specific base instant. That association is what lets a reader know which base file to merge it with, and it is why a stray log file with no matching base instant is an orphan rather than data. ## Reading a listing Given a partition directory containing files for two file IDs, one with a single Parquet file and one with a Parquet file plus two logs, you can state immediately: this is a Merge-on-Read table; the first group has been compacted or has received no updates since its base file was written; the second has unmerged updates that a snapshot query is paying to merge and a read-optimized query is ignoring. That inference — from listing to behaviour — is the reason the vocabulary is worth knowing.

  • Why does Hudi keep old file slices around after a newer one is written?
    So that readers stay consistent. A query that started before the commit keeps reading the slice it planned against, and time-travel and incremental queries need historical slices to reconstruct earlier states. The cleaner deletes superseded slices later under a retention policy, which is the knob that trades storage cost against how far back you can look.
  • How does the file group structure relate to the small-files problem?
    Sizing happens per group. Rather than writing a fresh file per commit, Hudi steers new inserts into existing base files that are still under the small-file limit and only opens a new file group when files approach the maximum size. That bin-packing at write time is why Hudi tables usually need less remedial compaction than a naive append-only writer's output.
  • Can a record key move from one file group to another?
    Not during ordinary upserts — that stability is what makes the index useful. Operations that explicitly replace file groups, such as clustering or an insert overwrite, do rewrite records into new groups and record a `replacecommit` on the timeline, after which the index reflects the new placement.

A file group is a numbered drawer that always holds the same customers' folders; a file slice is the contents of that drawer as of a given day, with Merge-on-Read leaving loose notes clipped to the front until someone refiles them.

saying these in an interview costs you the question

  • Says a file group is just another name for a partition
  • Thinks a file slice is a table-wide snapshot
  • Claims records are rehomed to new file groups on every upsert
  • Believes log files belong to the partition, not to a specific group
  • Assumes old slices are deleted the moment a new one commits

context