skip to content

Shared Array Loading

Loading a data file once for the whole instance instead of once per virtual user, and the ordinary array habits that quietly undo that saving on the first line of an iteration.

on this pageshow

explore

questions

5

What does k6's SharedArray from k6/data change about how a data file is held in memory?

level: juniorimportance: must knowfreq 76%

answer

  1. one copy per process, not per VU
  2. built only in the init context
  3. the name identifies it across runtimes
  4. elements kept as JSON strings
  5. index access parses a fresh copy

basics

~20 s

k6's SharedArray holds one copy of a data set per k6 process, not one per VU. Its loader runs once, k6 stores each element as a JSON string, and a VU gets a copy only on index access.

solid answer

~30 s

Every k6 VU is a separate JavaScript runtime, so a 50,000-row user list parsed at module scope is parsed and held once **per VU**. `new SharedArray(name, fn)` from `k6/data` breaks that multiplication: k6 calls `fn` at most once per name for the whole process, stores each returned element as a JSON string outside any VU's JS heap, and hands every VU a thin array-like wrapper over that one copy. Reading `users[i]` parses just that element's string and returns a fresh, deep-frozen copy, so you trade a small amount of CPU per access for memory that no longer scales with VU count.

code

javascript · 16 lines
javascript
import { SharedArray } from 'k6/data';

const users = new SharedArray('users', function () {
  const rows = new Array(50000);
  for (let i = 0; i < rows.length; i++) {
    rows[i] = { id: i, name: `user${i}`, token: `t-${i}` };
  }
  return rows;
});

export const options = { vus: 50, iterations: 200 };

export default function () {
  const user = users[Math.floor(Math.random() * users.length)];
  console.log(`picked ${user.name} of ${users.length} shared rows`);
}

go deeper

for a junior

Learn the three moving parts: import SharedArray from k6/data, construct it at module scope with a name and a loader function, and index it inside the default function.

for a middle

Explain the mechanism rather than the recipe: the loader runs once per name for the whole process, k6 stores each element as a JSON string, and every index read parses one of those strings back.

for a senior

Show that you have measured it. Say when the per-access parse outweighs the memory saved, and how you would confirm from a run that the sharing is actually happening rather than assumed.

for a principal

Frame it as a trade rather than a rule: a shared reference set costs CPU on every read and buys back heap that would otherwise scale with VU count, so the sensible default depends on data size and reads per iteration.

## The problem `SharedArray` solves Every k6 virtual user is a **separate JavaScript runtime**. Whatever the init code allocates, it allocates once per VU. Parse a 50,000-row user list with `JSON.parse(open('./users-50k.json'))` at module scope and a 200-VU run holds 200 independent parsed copies of those 50,000 rows on 200 separate JS heaps. The rows are identical in all of them, and no VU can see any other VU's copy. `SharedArray`, the single export of the built-in `k6/data` module, removes that multiplication. The data is produced once, marshalled once and stored once for the whole k6 process; each VU is handed a lightweight array-like wrapper over that one copy. ## The constructor ```javascript import { SharedArray } from 'k6/data'; const users = new SharedArray('users', function () { return JSON.parse(open('./users-50k.json')); // must return an array }); ``` Three parts of that call are load-bearing: 1. **It must run in the init context** — at module scope, before `setup()` or the default function. Constructing one later throws `new SharedArray must be called in the init context`. 2. **The first argument is a non-empty name.** VUs are separate runtimes, so the name — not the variable — is the identity. k6 keeps one process-wide map from name to data. 3. **The second argument is a plain, non-`async` function that returns an array.** k6 calls it at most once per name, no matter how many VUs reach that line. ## Where the data actually lives k6 does not keep the JavaScript array the loader returned. It walks it, runs `JSON.stringify()` on **each element**, and keeps the resulting strings. The shared copy therefore lives as a list of JSON texts held by the k6 process itself, outside every VU's JavaScript heap. That single implementation choice explains every other behaviour of the type. Two things follow immediately. The shared set occupies no JavaScript heap in any VU, so adding VUs does not add copies of it; and whatever the loader returned has already been flattened to text, so the array k6 keeps is emphatically not the array your loader built. ## What an element access costs Reading `users[i]` is not a pointer dereference. k6 takes the stored string for element `i`, runs `JSON.parse()` on it, deep-freezes the result and hands the VU that fresh value. The consequences: - Each read produces a **new object**, so `users[0] === users[0]` is `false`. - The copy is per-iteration garbage; it never grows the shared set. - An out-of-range index returns `undefined` rather than throwing. - `length`, index access, `for...of` and `forEach()` all work — the wrapper really is an array, and `Array.isArray(users)` is `true`. - Nothing is cached between reads, so an iteration needing the same row twice should hold it in a local variable. - The price is **CPU per access**, not memory: you pay a small parse in exchange for copies you no longer hold. | holding 50,000 rows | `const rows = JSON.parse(open(...))` | `new SharedArray('rows', ...)` | |---|---|---| | copies in memory | one per VU | one per k6 process | | when the parsing happens | once per VU, during init | once per name, during init | | cost of `rows[i]` | a plain property read | one `JSON.parse` of that element | | writable? | yes | no — throws `TypeError: SharedArray is immutable` | | usable to pass data between VUs? | no | no — read-only by design | ## When the trade pays off The saving scales with the size of the set and the number of VUs; the parse cost scales with how often each iteration reads. k6's own documentation reports that below roughly a thousand elements the memory benefit is small enough that the per-access parse can dominate, while from around ten thousand elements upward the saving becomes decisive. A 50,000-row user list read once or twice per iteration sits squarely in the range `SharedArray` was built for; a five-row lookup table hammered in a tight inner loop does not. ## What `SharedArray` is not It is **not** a communication channel. Once built it is read-only for the life of the test and there is no write path at all, so one VU cannot leave anything in it for another. It is also the wrong home for a single large object: wrapping one big configuration blob in a one-element array leaves you with a single stored JSON string that must be parsed in full on every access — the parse cost without the per-element granularity that makes the trade worth taking. And it is not a lazy view of a file: the loader runs to completion during init, so whatever it returned is fixed for the whole run, and a later change on disk is invisible to the test.

  • Does k6 keep the exact JavaScript array that the SharedArray loader returned?
    No. k6 walks the returned array and stores `JSON.stringify()` of each element as a separate string; the original array is discarded once the init context finishes. Every later read reconstructs one element from its stored string, which is why the shared data lives outside any VU's JavaScript heap.
  • In k6, does a SharedArray help at all in a run with a single VU?
    Barely. With one VU there is only one copy either way, and you still pay a `JSON.parse` on every element access, so a SharedArray can be marginally slower. The saving comes from the VU count multiplying the copies you would otherwise hold.
  • Why must a k6 SharedArray loader return an array rather than an object?
    k6 stores one JSON string per element and addresses them by integer index, so it needs a real array to walk. A loader returning anything else fails during init with `only arrays can be made into SharedArray`.

It is the reference copy of a directory kept behind the library desk: one book for the whole building, and each reader walks up and photocopies only the entry they need instead of taking a whole book home.

saying these in an interview costs you the question

  • Thinks a SharedArray lets one VU write data for other VUs
  • Says each VU still parses the file and k6 deduplicates afterwards
  • Believes indexing a SharedArray is free because the data is shared
  • Constructs the SharedArray inside the default function
  • Assumes a SharedArray beats a plain array for a handful of rows
open as a page

What does k6 require of the two arguments to new SharedArray(name, fn), and of where it is called?

level: middleimportance: should knowfreq 48%

basics

~20 s

k6 requires a non-empty name, a non-async function as the second argument that returns an array, and the call itself to sit in the init context. Each violation throws during init, before any VU iterates.

open as a page

What happens in k6 when a script assigns into a SharedArray or calls push() on it?

level: middleimportance: should knowfreq 57%

basics

~20 s

A k6 SharedArray is read-only: assigning to an index, calling push, or calling sort all raise TypeError: SharedArray is immutable. Elements read out of it come back deep-frozen, so a SharedArray cannot carry data between VUs.

open as a page

Why does calling .filter() on a k6 SharedArray inside the default function undo its memory saving?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Array methods such as filter and map read every element out of a k6 SharedArray and build an ordinary array in the calling VU's own heap, so each VU rebuilds the whole data set on every iteration that runs them.

open as a page

In k6, what exactly does a VU receive when it reads an element from a SharedArray?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

A fresh copy, parsed from that element's stored JSON text on every read and then deep-frozen. Two reads of the same index give two different objects, and anything JSON cannot express does not survive the round trip.

open as a page