skip to content

Vector API

The incubating Vector API lets you express data-parallel computations that compile to real SIMD instructions, with lanes, masks and a scalar fallback. It matters for numeric, cryptographic and ML workloads, and its incubator status is worth naming.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

Explain VectorSpecies, lanes, and SPECIES_PREFERRED in the Vector API. How do you size a vector loop correctly?

level: middleimportance: should knowfreq 30%

answer

  1. lane = one element slot in a vector
  2. species = element type + lane count (bit width)
  3. SPECIES_PREFERRED = widest the CPU supports, portable
  4. step by S.length(), bound by S.loopBound(n)
  5. always handle the tail (scalar loop or mask)

basics

~20 s

A VectorSpecies says what element type and how many slots (lanes) a vector holds. SPECIES_PREFERRED picks the best size for the current CPU. You step the loop by species.length() and use a tail loop for leftover elements.

solid answer

~50 s

A VectorSpecies<E> is the shape of a vector: the element type (Float, Int, Double...) plus the lane count — the 'lanes' being the individual element slots that hold one value each. The total bit width is lanes x element-bits (e.g. 8 float lanes = 256 bits = AVX2). You usually pick FloatVector.SPECIES_PREFERRED, which the runtime resolves to the widest species the actual CPU supports, so the same code runs optimally on AVX2, AVX-512, or NEON. To size a loop, step by species.length() and bound the vector portion with species.loopBound(n) (the largest multiple of the lane count that fits); then a scalar tail loop handles the remaining elements that don't fill a full vector, or you use a mask for that final partial step. Choosing a fixed species (e.g. SPECIES_256) sacrifices portability for predictability; SPECIES_PREFERRED is the usual default.

code

java · 12 lines
java
static final VectorSpecies<Double> S = DoubleVector.SPECIES_PREFERRED;

static double sum(double[] a) {
    DoubleVector acc = DoubleVector.zero(S);
    int i = 0, bound = S.loopBound(a.length);
    for (; i < bound; i += S.length()) {
        acc = acc.add(DoubleVector.fromArray(S, a, i));
    }
    double total = acc.reduceLanes(VectorOperators.ADD); // horizontal add
    for (; i < a.length; i++) total += a[i];             // scalar tail
    return total;
}

go deeper

for a junior

Knows a vector holds several values (lanes) and you step the loop by the lane count, handling leftovers separately.

for a middle

Defines species (type + lane count = bit width), uses SPECIES_PREFERRED, and writes a correct main+tail loop with loopBound.

for a senior

Explains PREFERRED vs fixed-width trade-offs, masked vs scalar tails, and why species should be a static final constant for JIT folding.

for a principal

Reasons about portability vs reproducibility, AVX-512 downclocking, and how species choice interacts with overall performance and numerical determinism guarantees.

## Lanes: the fundamental unit A **vector** in this API is a fixed-size pack of values of one primitive type. Each value slot is a **lane**. If a vector has 8 lanes of `float`, it holds 8 independent floats, and a lane-wise operation (like `add`) applies to each lane in parallel and independently — lane 0 with lane 0, lane 1 with lane 1, and so on. The number of lanes times the bits per element gives the vector's **bit width**: 8 floats x 32 bits = **256 bits**, which matches Intel **AVX2** registers; 16 floats = 512 bits matches **AVX-512**; 4 floats = 128 bits matches **ARM NEON** or SSE. ## VectorSpecies: the shape descriptor A **`VectorSpecies<E>`** is an immutable object describing *what shape* of vector you are working with: the **element type** `E` (boxed marker like `Float`, `Integer`, `Double`) and the **lane count**. Almost every factory and operation takes a species so it knows how many elements to load/store/operate on. You typically declare one `static final` species and reuse it: ```java static final VectorSpecies<Float> S = FloatVector.SPECIES_PREFERRED; int lanes = S.length(); // e.g. 8 on an AVX2 machine ``` ### Named species vs PREFERRED - **Fixed-width species** — `FloatVector.SPECIES_128`, `SPECIES_256`, `SPECIES_512` — force a specific bit width. Predictable, but if the hardware is narrower the runtime must emulate it (slower) and if it's wider you leave performance on the table. - **`SPECIES_PREFERRED`** — resolved at runtime to the **widest species the current CPU implements efficiently**. This is the portable default: write once, run optimally on whatever vector unit exists, and fall back to scalar if there is none. ## Sizing a loop correctly The array length is rarely an exact multiple of the lane count, so a vector loop has **two parts**: 1. **Main vector loop** — steps by `S.length()`, bounded by **`S.loopBound(n)`**, a helper that returns the largest multiple of the lane count that is `<= n`. This guarantees every iteration reads/writes a *full* vector and never runs off the end of the array. 2. **Tail (remainder) handling** — the leftover `n - loopBound(n)` elements. Either a plain **scalar tail loop**, or a single masked vector step (see the masks topic) that processes only the valid lanes. ```java int i = 0, n = a.length; int bound = S.loopBound(n); for (; i < bound; i += S.length()) { /* full vectors */ } for (; i < n; i++) { /* scalar tail */ } ``` ## Common pitfalls - **Don't hardcode `8`** — the lane count is hardware-dependent; always use `S.length()`. - **Don't forget the tail** — looping by `length()` up to `n` directly will throw `IndexOutOfBoundsException` (or silently skip) on the partial final block. - **Match the element type** — a `FloatVector` species can't load an `int[]` correctly; choose the species matching your array's primitive type. - **Reuse the species** as a `static final` constant so the JIT can constant-fold the lane count.

  • When would you deliberately choose SPECIES_256 over SPECIES_PREFERRED?
    When you need deterministic, reproducible behavior across machines (e.g. bit-exact numerical results or benchmarking a specific width), or when AVX-512 downclocking on a particular CPU makes the wider species slower in practice.
  • Why declare the species as static final?
    It lets the JIT treat the lane count as a constant, enabling better loop unrolling and bound computation, and avoids recomputing the species each call.

saying these in an interview costs you the question

  • Hardcoding the lane count (e.g. i += 8) instead of S.length()
  • Forgetting the remainder/tail, causing out-of-bounds or skipped elements
  • Thinking SPECIES_PREFERRED is a compile-time constant fixed per JVM build rather than resolved for the running CPU
  • Mixing element types — loading an int[] with a FloatVector species

context

open as a page

What is the Java Vector API, and what problem does it solve compared to ordinary scalar arithmetic loops?

level: middleimportance: should knowfreq 35%

basics

~20 s

The Vector API lets you write a loop that processes several array elements at once instead of one at a time. It maps your code to special CPU instructions (SIMD) so number-heavy work runs much faster.

open as a page

What are masks in the Vector API, and how do they let you handle loop tails and conditional (branchless) computation?

level: seniorimportance: should knowfreq 25%

basics

~20 s

A mask is a per-lane on/off switch. It tells an operation which lanes to actually apply to. You use it to safely process a partial final chunk of an array and to do if-style logic without branches by selecting values per lane.

open as a page

Why is the Vector API still an incubating module, what does that mean for using it, and how does it relate to Project Valhalla?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Incubating means it is an unfinished, opt-in module whose API can change between Java releases. You must explicitly enable it. It is waiting on Project Valhalla so vectors can become lightweight value types with no heap overhead before it is finalized.

open as a page

As an architect, when should you adopt the Vector API versus relying on JIT auto-vectorization or a native library? What are the costs and risks?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Use it only for hot, data-parallel numeric kernels where you have measured a real bottleneck and the JIT is not vectorizing reliably. For small or branchy code, let the JIT handle it; for heavy, well-tuned math, a native BLAS library may beat it. Weigh the cost of an unstable, incubating API.

open as a page