skip to content

Why does storing many int values in an ArrayList<Integer> cost more memory and CPU than an int[], and how big is the difference?

level: middleimportance: should knowfreq 60%

answer

  1. int[] = packed 4-byte values, no per-element overhead
  2. ArrayList<Integer> = each int boxed into ~16-byte object + reference
  3. ~4–6x memory, worse cache locality, more GC, boxing CPU
  4. Integer cache only covers -128..127
  5. Use primitive-collection libraries (fastutil/Eclipse) to avoid boxing

basics

~20 s

An int[] stores raw numbers packed together. An ArrayList<Integer> must wrap each int in an Integer object, which adds object overhead and an extra pointer per element, and costs CPU for boxing and unboxing. So it uses several times more memory and is slower.

solid answer

~50 s

A raw int[] stores 4-byte ints contiguously with no per-element overhead, so a million ints is about 4 MB plus a small array header. An ArrayList<Integer> can only hold objects, so every int is autoboxed into an Integer object. Each Integer carries a 12–16 byte object header plus the 4-byte value (often padded to 16 bytes total), and the backing Object[] stores a reference (4 or 8 bytes) to each one. That's roughly 16 + 8 = 20+ bytes per element versus 4 — about 4–5x the memory, plus the boxed objects are scattered on the heap, hurting cache locality and adding GC pressure. CPU-wise, every add() boxes and every get() may unbox, and the JVM caches only small Integers (-128..127), so larger values allocate fresh objects. For large numeric datasets prefer int[] or a specialized primitive-collection library.

code

java · 12 lines
java
// Packed primitives: ~4 MB, cache-friendly
int[] packed = new int[1_000_000];

// Boxed: each element is an Integer object + a reference
List<Integer> boxed = new ArrayList<>();
for (int i = 0; i < 1_000_000; i++) {
    boxed.add(i); // autoboxing on every add
}

// Integer cache trap
System.out.println(Integer.valueOf(100) == Integer.valueOf(100)); // true  (cached)
System.out.println(Integer.valueOf(1000) == Integer.valueOf(1000)); // false (new objects)

go deeper

for a junior

Understands that ArrayList<Integer> wraps each int in an object and therefore uses more memory than int[], even if the exact byte counts aren't memorized.

for a middle

Explains autoboxing, the per-element object header + reference overhead, the rough 4–6x figure, and the boxing/unboxing CPU cost; knows the Integer cache range.

for a senior

Adds cache-locality and GC-pressure reasoning, recommends primitive-collection libraries or int[] for hot paths, and warns against == on wrappers; weighs readability vs perf by scale.

for a principal

Reasons about memory layout (Project Valhalla / value types as the future fix), benchmarks claims rather than asserting, and sets codebase guidance on when primitive collections are worth the dependency and API friction.

## The starting point: what a primitive and an object cost In Java, an `int` is a **primitive**: a raw 32-bit (4-byte) value with no identity, no header, nothing extra. When you make an array of them, the values sit **contiguously** in memory: ```java int[] a = new int[1_000_000]; // ~4 MB of packed 4-byte ints + a small header ``` An **object**, by contrast, is never "just its value." Every Java object on the heap carries an **object header** — bookkeeping the JVM needs for things like its class pointer and its identity/lock/GC state. On the common HotSpot JVM this header is about **12 bytes** (16 with some settings), and objects are **8-byte aligned**, so they get padded up to the next multiple of 8. ## Why ArrayList forces objects: autoboxing An `ArrayList` is generic and generics only accept **reference types**, so you cannot have `ArrayList<int>`. You use `ArrayList<Integer>`, where `Integer` is the **wrapper class** around an `int`. Converting `int` → `Integer` automatically is **autoboxing**; the reverse is **unboxing**: ```java List<Integer> list = new ArrayList<>(); list.add(42); // autoboxing: makes/looks up an Integer object int v = list.get(0); // unboxing: pulls the int back out ``` So each number becomes a separate heap object: header (~12B) + the 4-byte int, padded to **16 bytes**. On top of that, the ArrayList's internal `Object[]` holds a **reference** (a pointer) to each Integer — 4 bytes with compressed pointers, 8 without. So per element you pay roughly **16 (the Integer) + 4–8 (the reference) ≈ 20–24 bytes**, versus **4 bytes** for an `int[]` slot. That's about **5x–6x** the memory for a large list. ## The Integer cache Java keeps a small **cache of Integer objects for the range −128 to 127** (the values most programs use most). Autoboxing a value in that range returns a shared cached object instead of allocating; outside that range, every box allocates a brand-new Integer. So `Integer.valueOf(100) == Integer.valueOf(100)` is `true` (same cached object) but `Integer.valueOf(1000) == Integer.valueOf(1000)` is `false` — a classic interview trap, and a reminder that you must compare wrappers with `.equals()`, never `==`. ## Two performance costs, not one 1. **Memory & GC.** Millions of tiny boxed Integers are scattered across the heap. They take several times more space and they are individual objects the **garbage collector** must track and scan, increasing GC work. 2. **CPU & cache locality.** Boxing on insert and unboxing on read both burn cycles. Worse, because the Integers are scattered (the array holds pointers, not the values), iterating chases pointers all over the heap — terrible for **CPU cache locality** compared to walking a packed `int[]`, where the next value is physically adjacent. ## What to do about it - For large numeric datasets in hot paths, prefer **`int[]`** (or `long[]`, `double[]`). - If you also need list-like growth for primitives, use a **specialized primitive collection** from a library (e.g. Eclipse Collections `IntArrayList`, fastutil `IntArrayList`, or Trove), which stores raw ints internally and avoids boxing entirely. - For ordinary small lists, none of this matters — `ArrayList<Integer>` is perfectly fine; readability wins. The boxing tax only becomes important at scale or in tight loops. ## Quick numbers to remember - `int` slot in an array: **4 bytes**, packed, cache-friendly. - `Integer` in a list: **~16 bytes object + ~4–8 byte reference**, scattered. - Net: roughly **4–6x memory**, plus boxing CPU and extra GC. Integer cache: **−128..127** only.

  • Why does Integer.valueOf(100) == Integer.valueOf(100) return true but the same with 1000 returns false?
    The JVM caches Integer objects for -128..127, so boxing 100 returns the same shared object both times (== is true). 1000 is outside the cache, so each call allocates a new object with a different identity, making == false. Always use .equals() for value comparison.
  • How would you store a large list of growing ints without boxing?
    Use a primitive-specialized collection such as fastutil's IntArrayList or Eclipse Collections' IntArrayList. They store raw int values in an internal int[] and grow like ArrayList, giving resizability without per-element Integer objects.

saying these in an interview costs you the question

  • Saying boxing is free or negligible at any scale
  • Believing Integer caching covers all values (it's only -128..127)
  • Comparing boxed Integers with == instead of .equals()
  • Claiming an int[] and ArrayList<Integer> use the same memory

context