In Apache Hudi, what is an instant on the .hoodie timeline?
answer
- a small directory beside the data decides what is real
- every action leaves a stamped entry there
- three attributes: what, when, how far along
- planned, running, done — as separate files
- readers trust only the last of the three
basics
~20 sAn instant is one entry on the timeline in a Hudi table's .hoodie directory: an action such as commit, deltacommit, compaction or clean, stamped with a monotonic time and a state of requested, inflight or completed.
solid answer
~40 sEvery action against a Hudi table is recorded in the `.hoodie` directory as an *instant*, described by three things: an action type, an instant time (a monotonically increasing timestamp string), and a state. The action tells you what happened — `commit` for a Copy-on-Write write, `deltacommit` for a Merge-on-Read write into log files, `compaction`, `clean`, `rollback`, `savepoint`, `replacecommit` for clustering and insert-overwrite. The state moves `REQUESTED` → `INFLIGHT` → `COMPLETED`, and in 0.x each state is a separate file, with the completed one carrying no state suffix. This is Hudi's atomicity mechanism: readers build a view from *completed* instants only, so partially written base and log files are invisible until the completed instant appears, and a crashed writer's inflight instant is later rolled back.
code
text · 8 lines.hoodie/
hoodie.properties
20240418093012345.deltacommit.requested
20240418093012345.deltacommit.inflight
20240418093012345.deltacommit
20240418094500123.compaction.requested
20240418095000456.clean
metadata/go deeper
Know that a Hudi table has a .hoodie directory beside the data and that queries are planned from it, not from listing files. Be able to say an entry there records an action, a time and how far along it is.
Explain the instant triple and the requested/inflight/completed progression, and why only completed instants are visible. Name the common actions and say which one a Merge-on-Read write produces.
Demonstrate operational reasoning: how a crashed writer is rolled back, how retention on the timeline plus the cleaner bounds time travel, and why an async compactor needs the requested state to carry a plan.
Own retention as policy. Decide how many commits the active timeline keeps given your slowest incremental consumer, your recovery expectations and your storage costs, and make that a documented platform default rather than a per-table accident.
## What the timeline is A Hudi table is a directory of partitions containing base files and log files, plus a `.hoodie` directory that holds the table's *timeline*. The timeline is the ordered log of everything that has ever been done to the table, and it is the only thing that makes the pile of files a table: a reader never lists a partition and trusts what it finds, it reads the timeline and derives which file slices are valid. The unit of the timeline is the **instant**. An instant has three attributes. **Action** — what is being done. The ones you meet constantly: - `commit` — a write that produced new base files (Copy-on-Write tables, and compaction output). - `deltacommit` — a write into log files on a Merge-on-Read table. - `compaction` — merging a file group's log files into a new base file. - `clean` — deleting old file slices that no longer serve any retained query. - `rollback` — undoing a failed write's partial files. - `savepoint` / `restore` — pinning a point so cleaning skips it, and rewinding to it. - `replacecommit` — an action that replaces one set of file groups with another, which is how clustering and insert-overwrite are recorded. **Instant time** — a monotonically increasing timestamp string (a `yyyyMMddHHmmssSSS`-shaped value). It orders the timeline and is stamped into every row written by that action, in the `_hoodie_commit_time` metadata column. That is what time travel and incremental queries position on. **State** — `REQUESTED` when the action is planned, `INFLIGHT` while it runs, `COMPLETED` when it has succeeded. In 0.x each state is its own small file in `.hoodie`, named `<instantTime>.<action>.<state>`, with the completed instant written as plain `<instantTime>.<action>` — no suffix. The completed file carries the metadata describing what the action did. ## Why three states The state machine is what gives Hudi atomicity on object storage without a transaction manager. Writing data files is not a commit; only the appearance of the completed instant is. A reader constructs its view from completed instants, so a half-written batch of base files sitting in a partition is simply not part of any file slice and is invisible to queries. If the writer dies mid-batch, its inflight instant stays on the timeline as evidence, and a later writer or cleaner issues a `rollback` instant that deletes the orphaned files and clears the inflight marker. The `REQUESTED` state matters for actions that are planned separately from being executed. A compaction is *scheduled* as a requested instant containing the compaction plan — which file groups, which log files — and an asynchronous compactor later picks that plan up, moves it to inflight, and completes it. The same split lets clustering be planned by one process and executed by another. ## Active and archived timeline The timeline cannot grow forever. Hudi keeps a bounded *active* timeline of recent instants and archives older ones into a compacted archived timeline, governed by retention configs of the `hoodie.keep.min.commits` / `hoodie.keep.max.commits` family. The archived instants are still a history record but are not used for normal planning. This matters operationally: time travel and incremental reads work over the active timeline, so archival retention, and the cleaner's file retention, jointly set how far back you can actually query. An incremental consumer that lags past the retention window can no longer resolve its start point. ## Version differences The `.hoodie` layout described above is the long-standing 0.x arrangement. Hudi 1.x reorganised the timeline — it lives under a dedicated timeline directory, completed instants record a completion time alongside the action's start time, and the archived timeline is stored in an LSM-tree structure. The completion time is the significant change: it lets incremental readers position on when a commit became visible rather than when it started, which closes a real gap with concurrent long-running writers. Say which line you are describing when the mechanics matter. ## Do not cross-attribute The timeline is Hudi's own machinery. It is not Delta Lake's `_delta_log`, which is a numbered sequence of JSON commit files containing `add`/`remove` actions with periodic Parquet checkpoints, and it is not Iceberg's metadata-file/manifest-list/manifest chain. All three achieve atomic commits; only Hudi expresses them as instants with requested/inflight/completed states, and confusing the three is the fastest way to fail this question. ## What interviewers probe Typical follow-ons: which action a Merge-on-Read write produces (`deltacommit`, not `commit`); what happens to files left by a crashed writer (rollback); why a snapshot query does not see an inflight instant; and how retention settings on the timeline interact with the cleaner to bound time travel.
- A Hudi writer crashes after writing several base files but before the instant completes. What happens to those files?They are orphaned: no completed instant references them, so no query sees them. The inflight instant stays on the timeline as a marker, and a subsequent writer or the cleaner issues a `rollback` instant that deletes the partial files, using the marker files Hudi wrote during the attempt to know what to remove. Nothing is exposed at any point.
- Which timeline action does clustering or an insert-overwrite record?`replacecommit`. Unlike a normal commit, which adds file slices, a replacecommit records that one set of file groups replaces another, so readers stop seeing the old groups atomically at that instant. Clustering, insert_overwrite and insert_overwrite_table all use it, which is why they show up differently from ordinary writes on the timeline.
- How long do instants stay on the active timeline?Until archival moves them out, bounded by the `hoodie.keep.min.commits` / `hoodie.keep.max.commits` retention family. Archived instants remain as history but are not used for normal planning, so together with the cleaner's file retention they set how far back time travel and incremental reads can go. Consumers that lag past that window cannot resolve their start instant.
saying these in an interview costs you the question
- Thinks data files become visible as soon as they land in a partition
- Says queries read requested and inflight instants too
- Calls Hudi's timeline a transaction log of add and remove actions like Delta's
- Assumes every write produces a commit instant, even on Merge-on-Read
- Believes the timeline retains every instant forever