skip to content

Memory-Mapped Files & File Locking

FileChannel.map exposes a file region as a MappedByteBuffer so the OS handles paging, which is why memory mapping is fast for large files. FileLock adds shared or exclusive advisory locks for coordinating between processes.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a memory-mapped file in Java, and how do you create one with FileChannel.map()?

level: middleimportance: must knowfreq 58%

answer

  1. FileChannel.map(mode, pos, size) -> MappedByteBuffer
  2. virtual memory + demand paging, no explicit read/write
  3. ~2 GB per mapping (int capacity)
  4. force() to flush, unmap is GC-bound

basics

~20 s

A memory-mapped file links a region of a file directly into memory so you read and write it like an array instead of calling read/write. In Java you get one by calling channel.map(mode, position, size), which returns a MappedByteBuffer.

solid answer

~40 s

Memory mapping uses the OS virtual-memory system to map a file region into the process's address space. FileChannel.map(MapMode, position, size) returns a MappedByteBuffer; accessing its bytes triggers the OS to page file contents in on demand and to write dirty pages back, so I never explicitly call read()/write(). It is attractive for large files and random access because it avoids copying data through user-space buffers and lets multiple processes share the same physical pages. The buffer is limited to Integer.MAX_VALUE bytes (~2 GB) per mapping, so huge files need several mappings. MapMode controls access: READ_ONLY, READ_WRITE, or PRIVATE (copy-on-write). I call force() to flush READ_WRITE changes durably. A caveat is that unmapping is tied to GC, so the file may stay locked until the buffer is collected.

code

java · 12 lines
java
import java.nio.MappedByteBuffer;
import java.nio.channels.FileChannel;
import java.nio.file.*;

try (FileChannel ch = FileChannel.open(Path.of("data.bin"),
        StandardOpenOption.READ, StandardOpenOption.WRITE)) {
    MappedByteBuffer buf =
        ch.map(FileChannel.MapMode.READ_WRITE, 0, ch.size());
    int header = buf.getInt(0);   // read 4 bytes as if from RAM
    buf.putInt(0, header + 1);     // mutate in place
    buf.force();                   // flush dirty pages to disk
}

go deeper

for a junior

Knows a memory-mapped file lets you treat a file region like an in-memory array and is created via FileChannel.map() returning a MappedByteBuffer.

for a middle

Can explain demand paging / page cache, the read/write-without-syscalls benefit, the ~2 GB-per-mapping limit, and that force() flushes writes.

for a senior

Discusses when mapping pays off vs. buffered I/O, the GC-bound unmap problem and its Windows file-lock consequences, and multi-window mapping for large files.

for a principal

Reasons about page-cache pressure, TLB/page-fault costs, durability semantics of force() vs fsync, cross-process shared mappings, and migration to the FFM API / MemorySegment for >2 GB and deterministic cleanup.

## The problem memory mapping solves Normally, to read a file you call something like `channel.read(buffer)`. The operating system copies bytes from the disk into the kernel's page cache, then copies them again from the kernel into your program's buffer. To write, the reverse happens. For large files or random access, all that copying and all those system calls are expensive. **Memory mapping** removes the explicit copy and the per-operation system call. Instead of *asking* the OS for bytes, you tell the OS: "make this file region appear in my address space as if it were ordinary memory." After that, reading byte 5,000,000 of the file is just reading memory location `base + 5,000,000`. No `read()` call. ## How it works under the hood Modern CPUs use **virtual memory**: programs see virtual addresses, and the OS maps each page (typically 4 KB) of virtual address space to a page of physical RAM (or to disk). Memory mapping reuses this machinery. When you map a file, the OS sets up page-table entries that point at the file. The pages start out **not loaded**. The first time your code touches a page, the CPU raises a *page fault*; the OS catches it, reads that page from the file into the page cache, fixes the page table, and resumes your code — transparently. This is **demand paging**: only the parts you actually touch are loaded. Dirty (modified) pages are written back to disk later by the OS, or immediately when you call `force()`. ## The Java API ```java try (FileChannel ch = FileChannel.open(path, StandardOpenOption.READ, StandardOpenOption.WRITE)) { MappedByteBuffer buf = ch.map(FileChannel.MapMode.READ_WRITE, 0, ch.size()); byte first = buf.get(0); // read like an array buf.put(0, (byte) 42); // write like an array } ``` `FileChannel.map(MapMode mode, long position, long size)` returns a `MappedByteBuffer` (a subclass of `ByteBuffer`). Terms: - **FileChannel** — Java's NIO object representing an open file you can read/write/map. You get one from `FileChannel.open(...)` or from a stream's `getChannel()`. - **MappedByteBuffer** — a `ByteBuffer` whose backing storage is the mapped file region rather than ordinary heap/native memory. - **position / size** — the byte offset in the file where the mapping starts and how many bytes it covers. ### The 2 GB limit `size` is a `long`, but a `ByteBuffer`'s capacity is an `int`. So one mapping can cover at most `Integer.MAX_VALUE` (~2.1 GB) bytes. Larger files require multiple mappings. (Java 21+ adds the Foreign Function & Memory API / `MemorySegment` to map larger regions without this cap.) ## When mapping helps — and when it doesn't Good fit: large files, random access, repeated access to the same regions, sharing data between processes. Poor fit: small files (mapping overhead isn't worth it), purely sequential one-pass reads (a buffered stream is often as fast and simpler), or when you need deterministic resource release (see below). ## The unmapping gotcha There is **no public, portable `unmap()`** in standard Java before the FFM API. A mapping is released when its `MappedByteBuffer` is garbage-collected. Until then the OS may keep the file open/locked — on Windows this can block deleting or renaming the file. This non-deterministic cleanup is the single most surprising part of the API, and a frequent interview point. ## Summary Memory mapping turns file I/O into memory access by leveraging virtual memory and demand paging, eliminating user-space copies and per-operation syscalls. You create it with `FileChannel.map(mode, position, size)` → `MappedByteBuffer`, mind the ~2 GB-per-mapping limit, call `force()` to flush writes durably, and remember cleanup is GC-bound.

  • Why can't a single MappedByteBuffer map a 4 GB file?
    Because ByteBuffer capacity is an int, one mapping is capped at Integer.MAX_VALUE (~2.1 GB). You map the file in multiple windows, or use the Java 21+ Foreign Function & Memory API (MemorySegment) which uses long-sized addressing.
  • How do you make sure mapped writes survive a crash?
    Call MappedByteBuffer.force(), which asks the OS to flush the dirty pages of that mapping to the storage device, analogous to fsync. Without it, the OS may delay write-back and the data could be lost on a power failure.

It's like putting a document on a shared whiteboard instead of photocopying pages back and forth: everyone reads and writes the same surface directly, and only the parts someone looks at are actually drawn in.

saying these in an interview costs you the question

  • Claiming map() copies the whole file into the heap — it uses demand paging and the OS page cache, not the Java heap
  • Thinking a single mapping can cover any file size — capacity is an int (~2 GB)
  • Assuming the mapping is released immediately when you close the channel — it's freed only when the buffer is GC'd
  • Confusing MappedByteBuffer with a normal heap ByteBuffer allocated by ByteBuffer.allocate()

context

open as a page

What is a FileLock in Java NIO, and how do shared vs. exclusive locks and lock() vs. tryLock() work?

level: middleimportance: should knowfreq 48%

basics

~20 s

A FileLock lets a program lock a file (or part of it) so processes coordinate access. A shared lock allows multiple readers; an exclusive lock allows only one writer. lock() waits until it can get the lock; tryLock() returns immediately, giving null if the lock isn't available.

open as a page

What are the MapMode options for FileChannel.map() (READ_ONLY, READ_WRITE, PRIVATE) and how do they differ?

level: middleimportance: should knowfreq 42%

basics

~20 s

READ_ONLY lets you only read the file. READ_WRITE lets you read and write, and changes go to the file. PRIVATE lets you write, but your changes stay in memory (copy-on-write) and are never saved back to the file.

open as a page

What does MappedByteBuffer.force() do, and why is it important for durability of memory-mapped writes?

level: seniorimportance: should knowfreq 40%

basics

~20 s

When you write through a memory-mapped file, the changes sit in memory and the OS decides when to save them to disk. force() tells the OS to flush those changes to disk now, so they aren't lost if the program or machine crashes.

open as a page

What are the main pitfalls of memory-mapped files in Java — resource cleanup, the 2 GB limit, and crash safety — and how do you address them?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Mapped files can't be reliably unmapped before the buffer is garbage-collected, so files may stay locked (notably on Windows). One mapping can't exceed about 2 GB. And writes aren't safe on disk until you call force(). Plan for all three.

open as a page