skip to content

Why is the default HDFS block size 128 MB rather than a few kilobytes?

level: middleimportance: should knowfreq 52%

answer

  1. think about what a disk is slow at
  2. one number ties seek time to transfer time
  3. the master pays per object, not per byte
  4. a task needs enough work to be worth starting
  5. short files do not pad the last block

basics

~20 s

A large HDFS block amortises disk seek time over a long sequential read and keeps the NameNode's per-byte metadata cost tiny. Blocks are logical, so a 5 MB file consumes 5 MB of disk, not 128 MB.

solid answer

~50 s

HDFS is built for streaming scans of huge files, and the block size is tuned for that. With `dfs.blocksize` at 128 MB the seek to the start of a block is negligible next to the time spent reading it, so transfer runs near raw disk bandwidth. A large block also means far fewer block objects for a given volume of data, and every block is an entry in the NameNode's in-memory map — halving the block size doubles that pressure. Finally, a block is the unit of work handed to a task, so 128 MB gives each task enough to do to cover its scheduling overhead. Crucially, blocks are **logical**: a 5 MB file occupies one block that takes 5 MB of DataNode disk, not a padded 128 MB. Clusters with very large files often raise the setting to 256 MB or more per directory or per file.

code

xml · 4 lines
xml
<property>
  <name>dfs.blocksize</name>
  <value>268435456</value>
</property>

go deeper

for a junior

Recall that HDFS chops files into large blocks, that the default is 128 MB, and that a file smaller than a block does not waste the remainder on disk. Do not confuse it with an OS filesystem block.

for a middle

Explain the three drivers: seek time amortised over a long sequential transfer, fewer block objects in NameNode memory per byte stored, and one block being the unit of scheduled work. Mention that the setting applies at write time only.

for a senior

Show judgment about tuning it per dataset — raising it for archival scans, and the parallelism and re-read costs that follow. Be ready to relate block count to task count and to NameNode heap headroom during capacity planning.

for a principal

Own block sizing as a cluster-wide policy question: it sets the ceiling on how much data a single NameNode's heap can front, and it fixes the granularity every downstream engine inherits. Weigh it against erasure coding and against moving cold data off HDFS entirely.

## What a block actually is HDFS splits every file into fixed-size **blocks** and replicates each block independently across DataNodes. The size is controlled by `dfs.blocksize` (in `hdfs-site.xml`), whose default in Hadoop 2.x and 3.x is 128 MB — written as `134217728` bytes. It can be overridden per directory or per file at creation time, so one cluster can hold 128 MB blocks for general data and 512 MB blocks for a bulk archive. Two properties matter and are frequently conflated: - The block size is a **maximum**, not an allocation unit. Blocks are logical. A 5 MB file has exactly one block, and that block occupies 5 MB on each DataNode that holds a replica. There is no internal fragmentation the way there is with an operating-system filesystem block. - A block is the unit of **replication, placement and recovery**. Three-way replication means three copies of each block, potentially on three different sets of machines for different blocks of the same file. ## Reason one: amortise the seek A spinning disk costs a few milliseconds to position the head and then delivers on the order of a hundred megabytes per second sequentially. If the block were 4 MB, a scan would pay a seek for every fraction of a second of reading, and a meaningful share of wall-clock time would be head movement. At 128 MB, the seek is a rounding error against a full second or more of sequential transfer, so a scan runs close to the drive's streaming bandwidth. The classic sizing rule is to pick a block big enough that the seek is a small percentage of the transfer time. That rule is what produced 64 MB in early Hadoop and 128 MB later as disks got faster and files got bigger. On flash-backed clusters the seek argument weakens, but the other two reasons stand. ## Reason two: NameNode metadata per byte Every block is an object in the NameNode's in-memory block map, and every file is another object holding the ordered list of its block IDs. Storing one petabyte in 128 MB blocks costs roughly eight million block objects; the same petabyte in 8 MB blocks costs sixteen times that, with no change in the bytes stored. Since the NameNode holds the entire namespace in heap, block size is directly a **capacity knob for the master**. Bigger blocks mean the cluster can hold more data before the NameNode becomes the limit. ## Reason three: the unit of parallel work Processing frameworks derive their unit of parallelism from block boundaries. A MapReduce input split, or a Spark partition reading a splittable file, typically corresponds to one block, and the task is scheduled — where possible — on a node that already holds that block, so the read is local. Launching a task costs real time; a task that reads 128 MB has enough work to hide that cost, whereas a task that reads 2 MB spends most of its life starting and stopping. This is also why block size and task count are linked: raising `dfs.blocksize` for a dataset reduces the number of tasks that scan it. ## Choosing a different size Raise it when files are consistently much larger than the block size and scans are the dominant access pattern — a 256 MB or 512 MB block cuts block count and task count proportionally. Be aware of the consequences: fewer, larger tasks means coarser parallelism, so a small cluster can end up with fewer tasks than it has slots, and a failed task re-reads more data. Lowering the block size is rarely the right answer; if a job is under-parallelised, the usual fix is more files or an explicit repartition, not smaller blocks. The setting applies at write time. Changing `dfs.blocksize` does not re-block existing files — they keep the size they were written with, which is visible per file via `hdfs fsck` or `hdfs dfs -stat %o`. ## The last block, and what it means for small files A 300 MB file at a 128 MB block size becomes three blocks: two full and one of 44 MB. The short final block wastes no disk. But note what a very small file costs: one file object plus one block object in NameNode memory, regardless of how few bytes it holds. That is why "blocks are logical, so small files waste no disk" and "small files are an HDFS problem" are both true at once — the waste is in the master's heap and in per-file overhead on every reader, not on the DataNode's platters. ## Interaction with erasure coding Under erasure coding the block concept is striped rather than contiguous: a logical block group is split into cells spread across many DataNodes with parity cells alongside. The nominal block size still governs the size of the group, but the "one block sits whole on one DataNode" mental model no longer holds, and data locality for reads largely disappears.

  • Does a 5 MB file stored in HDFS with a 128 MB block size consume 128 MB of disk?
    No. HDFS blocks are logical, so the file occupies one block holding 5 MB, and each replica takes 5 MB of DataNode disk. The real cost of small files is elsewhere: one file object plus one block object in the NameNode's heap, and per-file RPC and open overhead for every reader, regardless of how few bytes the file contains.
  • If you change dfs.blocksize on a running cluster, what happens to files that already exist?
    Nothing. Block size is fixed at write time and recorded per file, so existing files keep the size they were created with. Only new writes pick up the new value. To re-block old data you have to rewrite it, and you can check any file's actual block size with `hdfs dfs -stat %o` or `hdfs fsck -files -blocks`.
  • What goes wrong if you raise the block size to 1 GB on a small cluster?
    Parallelism collapses. A 4 GB dataset becomes four blocks and therefore roughly four tasks, so most of the cluster idles even though there is work to do. Recovery also gets more expensive: a failed task re-reads a gigabyte. Large blocks pay off when files are far larger than the block and the cluster has enough blocks to keep every slot busy.

saying these in an interview costs you the question

  • Claiming a small file wastes a whole 128 MB block on disk
  • Saying the block size is still 64 MB in current Hadoop
  • Thinking changing dfs.blocksize re-blocks existing files
  • Suggesting a tiny block size to increase parallelism
  • Ignoring that each block costs NameNode heap regardless of fill

context