What is a memory-mapped file in Java, and how do you create one with FileChannel.map()?
answer
- FileChannel.map(mode, pos, size) -> MappedByteBuffer
- virtual memory + demand paging, no explicit read/write
- ~2 GB per mapping (int capacity)
- force() to flush, unmap is GC-bound
basics
~20 sA memory-mapped file links a region of a file directly into memory so you read and write it like an array instead of calling read/write. In Java you get one by calling channel.map(mode, position, size), which returns a MappedByteBuffer.
solid answer
~40 sMemory mapping uses the OS virtual-memory system to map a file region into the process's address space. FileChannel.map(MapMode, position, size) returns a MappedByteBuffer; accessing its bytes triggers the OS to page file contents in on demand and to write dirty pages back, so I never explicitly call read()/write(). It is attractive for large files and random access because it avoids copying data through user-space buffers and lets multiple processes share the same physical pages. The buffer is limited to Integer.MAX_VALUE bytes (~2 GB) per mapping, so huge files need several mappings. MapMode controls access: READ_ONLY, READ_WRITE, or PRIVATE (copy-on-write). I call force() to flush READ_WRITE changes durably. A caveat is that unmapping is tied to GC, so the file may stay locked until the buffer is collected.
code
java · 12 linesimport java.nio.MappedByteBuffer;
import java.nio.channels.FileChannel;
import java.nio.file.*;
try (FileChannel ch = FileChannel.open(Path.of("data.bin"),
StandardOpenOption.READ, StandardOpenOption.WRITE)) {
MappedByteBuffer buf =
ch.map(FileChannel.MapMode.READ_WRITE, 0, ch.size());
int header = buf.getInt(0); // read 4 bytes as if from RAM
buf.putInt(0, header + 1); // mutate in place
buf.force(); // flush dirty pages to disk
}go deeper
Knows a memory-mapped file lets you treat a file region like an in-memory array and is created via FileChannel.map() returning a MappedByteBuffer.
Can explain demand paging / page cache, the read/write-without-syscalls benefit, the ~2 GB-per-mapping limit, and that force() flushes writes.
Discusses when mapping pays off vs. buffered I/O, the GC-bound unmap problem and its Windows file-lock consequences, and multi-window mapping for large files.
Reasons about page-cache pressure, TLB/page-fault costs, durability semantics of force() vs fsync, cross-process shared mappings, and migration to the FFM API / MemorySegment for >2 GB and deterministic cleanup.
## The problem memory mapping solves Normally, to read a file you call something like `channel.read(buffer)`. The operating system copies bytes from the disk into the kernel's page cache, then copies them again from the kernel into your program's buffer. To write, the reverse happens. For large files or random access, all that copying and all those system calls are expensive. **Memory mapping** removes the explicit copy and the per-operation system call. Instead of *asking* the OS for bytes, you tell the OS: "make this file region appear in my address space as if it were ordinary memory." After that, reading byte 5,000,000 of the file is just reading memory location `base + 5,000,000`. No `read()` call. ## How it works under the hood Modern CPUs use **virtual memory**: programs see virtual addresses, and the OS maps each page (typically 4 KB) of virtual address space to a page of physical RAM (or to disk). Memory mapping reuses this machinery. When you map a file, the OS sets up page-table entries that point at the file. The pages start out **not loaded**. The first time your code touches a page, the CPU raises a *page fault*; the OS catches it, reads that page from the file into the page cache, fixes the page table, and resumes your code — transparently. This is **demand paging**: only the parts you actually touch are loaded. Dirty (modified) pages are written back to disk later by the OS, or immediately when you call `force()`. ## The Java API ```java try (FileChannel ch = FileChannel.open(path, StandardOpenOption.READ, StandardOpenOption.WRITE)) { MappedByteBuffer buf = ch.map(FileChannel.MapMode.READ_WRITE, 0, ch.size()); byte first = buf.get(0); // read like an array buf.put(0, (byte) 42); // write like an array } ``` `FileChannel.map(MapMode mode, long position, long size)` returns a `MappedByteBuffer` (a subclass of `ByteBuffer`). Terms: - **FileChannel** — Java's NIO object representing an open file you can read/write/map. You get one from `FileChannel.open(...)` or from a stream's `getChannel()`. - **MappedByteBuffer** — a `ByteBuffer` whose backing storage is the mapped file region rather than ordinary heap/native memory. - **position / size** — the byte offset in the file where the mapping starts and how many bytes it covers. ### The 2 GB limit `size` is a `long`, but a `ByteBuffer`'s capacity is an `int`. So one mapping can cover at most `Integer.MAX_VALUE` (~2.1 GB) bytes. Larger files require multiple mappings. (Java 21+ adds the Foreign Function & Memory API / `MemorySegment` to map larger regions without this cap.) ## When mapping helps — and when it doesn't Good fit: large files, random access, repeated access to the same regions, sharing data between processes. Poor fit: small files (mapping overhead isn't worth it), purely sequential one-pass reads (a buffered stream is often as fast and simpler), or when you need deterministic resource release (see below). ## The unmapping gotcha There is **no public, portable `unmap()`** in standard Java before the FFM API. A mapping is released when its `MappedByteBuffer` is garbage-collected. Until then the OS may keep the file open/locked — on Windows this can block deleting or renaming the file. This non-deterministic cleanup is the single most surprising part of the API, and a frequent interview point. ## Summary Memory mapping turns file I/O into memory access by leveraging virtual memory and demand paging, eliminating user-space copies and per-operation syscalls. You create it with `FileChannel.map(mode, position, size)` → `MappedByteBuffer`, mind the ~2 GB-per-mapping limit, call `force()` to flush writes durably, and remember cleanup is GC-bound.
- Why can't a single MappedByteBuffer map a 4 GB file?Because ByteBuffer capacity is an int, one mapping is capped at Integer.MAX_VALUE (~2.1 GB). You map the file in multiple windows, or use the Java 21+ Foreign Function & Memory API (MemorySegment) which uses long-sized addressing.
- How do you make sure mapped writes survive a crash?Call MappedByteBuffer.force(), which asks the OS to flush the dirty pages of that mapping to the storage device, analogous to fsync. Without it, the OS may delay write-back and the data could be lost on a power failure.
It's like putting a document on a shared whiteboard instead of photocopying pages back and forth: everyone reads and writes the same surface directly, and only the parts someone looks at are actually drawn in.
saying these in an interview costs you the question
- Claiming map() copies the whole file into the heap — it uses demand paging and the OS page cache, not the Java heap
- Thinking a single mapping can cover any file size — capacity is an int (~2 GB)
- Assuming the mapping is released immediately when you close the channel — it's freed only when the buffer is GC'd
- Confusing MappedByteBuffer with a normal heap ByteBuffer allocated by ByteBuffer.allocate()