skip to content

In an ORC data file, what are stripes, the file footer and row indexes for?

level: middleimportance: nice to knowfreq 30%

answer

  1. one chunk that a task can read on its own
  2. the file's map lives at the very end
  3. how a reader jumps past 10,000 rows at a time
  4. three structures: chunk, table-of-contents, fine-grained skip

basics

~20 s

An ORC file splits rows into stripes of roughly 64 MB that can be read independently, a footer at the end lists those stripes with the file's schema and column statistics, and row indexes inside each stripe let a reader skip groups of rows that cannot match a filter.

solid answer

~50 s

ORC is a columnar data-file format, an alternative to Parquet for the data layer of a lakehouse table. A file is a sequence of **stripes** — self-contained chunks of rows, around 64 MB by default — each holding index data, the column data itself, and a stripe footer describing its streams. At the end of the file sits the **file footer**, which records the schema, the row count, the location and size of every stripe, and file-level column statistics; after it a small **postscript** stores the compression codec and the footer's length, which is why a reader seeks to the end of the file first. Within a stripe, **row indexes** carry min/max and stream positions for every group of 10,000 rows by default, so a reader can jump straight past groups that cannot satisfy a predicate. Optional bloom filters add equality-match skipping.

code

text · 6 lines
text
ORC file layout (read from the end backwards)
  Stripe 1: [index data][row data streams][stripe footer]   ~64 MB
  Stripe 2: [index data][row data streams][stripe footer]
  ...
  File footer : schema, row count, stripe offsets, column stats
  Postscript  : compression codec, footer length   <- read first

go deeper

for a junior

Recall the three names and one purpose each: stripes are independently readable row chunks, the footer at the end maps the file, and row indexes allow skipping groups of rows.

for a middle

Explain the read path — postscript, then footer, then selected stripes and streams — and why writing statistics last forces the map to the end of the file.

for a senior

Judge when the defaults need changing: stripe size versus the engine's split behaviour, and whether bloom filters on specific high-cardinality filter columns earn their write cost.

for a principal

Own the estate-level question: whether a large existing ORC lake is adopted in place under a table format that supports it, or re-encoded to standardize on one data-file format across engines.

## Where ORC sits ORC (Optimized Row Columnar) is a **data-file format**, the same layer as Parquet — it describes the inside of one immutable file and knows nothing about the table it belongs to. It matters in a lakehouse conversation because a table format can hold data files in more than one format: Apache Iceberg supports Parquet, ORC and Avro data files, while Delta Lake's data files are Parquet. Teams arriving from a Hive estate often already have petabytes of ORC, so the question "can we adopt a table format without re-encoding?" turns on whether ORC is supported. ## Stripes An ORC file is a sequence of **stripes**, each an independently readable chunk of rows, sized around 64 MB by default. Independence matters twice: it lets a distributed engine hand different stripes of a large file to different tasks, and it means a reader can fetch one stripe without touching the rest of the file. Each stripe has three parts in order: index data, row data (the actual column streams), and a stripe footer that describes where each column's streams live within the stripe and how they are encoded. Because columns are stored as separate streams, a query touching three of two hundred columns reads only those three streams from each stripe it needs. ## The file footer and postscript At the end of the file sits the **file footer**. It holds the table-of-contents for the file: the schema, the total row count, the offset and length of every stripe, and file-level statistics per column such as min, max and null counts. After the footer comes a short **postscript**, the very last bytes of the file, which records the compression codec, the compression block size and the footer's length. That ordering explains the read sequence: a reader seeks to the end of the file, reads the postscript to learn how long the footer is and how it is compressed, reads the footer to learn the schema and stripe layout, then fetches only the stripes and streams it needs. This is why an ORC file cannot be read from a pure forward-only stream without buffering — its map is at the back, a consequence of writing statistics that are only known once all the rows have been seen. ## Row indexes Inside each stripe's index data are **row indexes**, one entry per row group. The default row index stride is 10,000 rows, so a 1,000,000-row stripe carries roughly 100 entries per column. Each entry stores statistics for its 10,000 rows — min, max, null presence — plus the positions in each compressed stream where that row group begins. Those positions are the useful part. A predicate such as `WHERE order_ts >= '2026-03-01 12:00'` is compared against each entry's range; groups that cannot match are skipped by seeking directly to the next candidate group's recorded position rather than decoding forward through everything. The result is skipping at a much finer grain than the stripe. Optionally, a writer can add **bloom filters** for chosen columns. Min/max ranges answer "could this range contain values ≥ X", which is weak for equality filters on high-cardinality columns scattered across the range; a bloom filter answers "is this specific value definitely absent", which is exactly the equality case, at the cost of extra bytes in the index. ## How this relates to the table layer ORC's internal machinery skips *within* a file. It cannot decide which files belong to the table, and it cannot avoid opening a file in the first place — that decision belongs to the table layer, which keeps per-file statistics in its own metadata and prunes the file list during planning. The two compose: the table layer drops most files without opening them, and ORC's footer and row indexes then avoid decoding most of what remains in the files that survive. ## What to say and what to avoid A good answer names the three structures and what each is *for*: stripes for parallel, independent reads; the footer for the file's schema, stripe map and column statistics; row indexes for fine-grained skipping within a stripe. A weak answer describes ORC as a table format, or claims its statistics give the table ACID properties or time travel — neither is a property any single file can provide.

  • Why does an ORC reader have to seek to the end of the file before it can read anything useful?
    The file footer — schema, stripe locations, column statistics — is written last, because statistics are only known once every row has been written. The postscript at the very end records the footer's length and compression, so the reader fetches the tail first, decodes the postscript, then the footer, and only then knows where to fetch stripes from.
  • When does adding bloom filters to an ORC column pay off?
    When queries do equality lookups on a high-cardinality column whose values are scattered so widely that min/max ranges in every row group overlap the search value. The bloom filter can then rule out row groups that ranges cannot. It costs index size and write time, so it is worth enabling only for the specific columns real queries filter on.
  • If a lakehouse table's data files are ORC, does that change what the table format does?
    Barely. The table layer still tracks the file list, schema, versions and per-file statistics; the file format only determines how the bytes inside are encoded and how a reader skips within a file. The practical constraints are which formats the table format supports for reading and writing, and whether every engine in your stack reads that combination well.

saying these in an interview costs you the question

  • Describing ORC as a table format that provides transactions
  • Placing the file footer at the start of the file
  • Confusing stripe-level skipping with pruning whole files
  • Assuming row indexes guarantee matching rows in a group
  • Claiming ORC and Parquet cannot coexist under one table format

context