Why does calling .filter() on a k6 SharedArray inside the default function undo its memory saving?
answer
- array methods build a new array
- the copy lives in the VU's heap
- every element is parsed again
- filter in the loader, not the iteration
basics
~20 sArray methods such as filter and map read every element out of a k6 SharedArray and build an ordinary array in the calling VU's own heap, so each VU rebuilds the whole data set on every iteration that runs them.
solid answer
~40 sThe wrapper k6 hands you really is an array — `Array.isArray()` is `true` and it inherits `Array.prototype` — so `filter`, `map`, `slice`, spread and `JSON.stringify` all work, and none of them throws. That is the trap. Each of them walks every index, and every index read parses that element's stored JSON string into a fresh object. One `.filter()` over 50,000 shared rows therefore performs 50,000 parses and leaves a plain, per-VU array of the matches on that VU's heap. Put it on the first line of the default function with 200 VUs and you have re-created the per-VU duplication the `SharedArray` existed to remove. Filter inside the loader instead, where k6 runs it once.
code
javascript · 18 linesimport { SharedArray } from 'k6/data';
import { sleep } from 'k6';
const users = new SharedArray('users', function () {
const rows = new Array(50000);
for (let i = 0; i < rows.length; i++) rows[i] = { id: i, active: i % 2 === 0 };
return rows;
});
export const options = { vus: 200, duration: '30s' };
export default function () {
// Parses 50000 elements and retains ~25000 objects, per VU, per iteration.
const active = users.filter((u) => u.active);
const user = active[Math.floor(Math.random() * active.length)];
if (!user) throw new Error('empty subset');
sleep(1);
}go deeper
Reach into the array by index — rows[i] — rather than calling filter or map inside the default function. Anything that returns a new array has just copied the data set into your own VU.
Explain the mechanism: array methods walk every index, each index read parses that element's stored JSON, and the result is an ordinary per-VU array that k6 no longer shares.
Look for it whenever memory grows with VU count although a SharedArray is in the script. The usual culprit is one filter, map or spread near the top of the default function.
Make it a convention that all shaping of shared data happens in the loader, where k6 runs it once. Per-iteration reshaping is the failure mode that survives review precisely because nothing throws.
## What `.filter()` actually does to a SharedArray The object `new SharedArray(...)` returns is a genuine array as far as JavaScript is concerned: `Array.isArray(rows)` is `true`, it has a `length`, and it inherits every method on `Array.prototype`. Nothing about it signals that reads are expensive. The read-only methods therefore all work, and all work the same way: they walk index `0` to `length - 1`, and each index read makes k6 fetch that element's stored JSON string, run `JSON.parse()` on it and deep-freeze the result. `.filter()` then collects the matches into a **new, ordinary JavaScript array** in the calling VU's own heap — mutable, unshared, and retained for as long as the variable that holds it lives. ## The cost, with a 50,000-row list Take a 50,000-row user list held once per k6 process and this on the first line of the default function: ```javascript const active = users.filter((u) => u.active); ``` Every iteration, in every VU, now: 1. Parses 50,000 JSON strings. 2. Allocates up to 50,000 fresh objects and freezes each one. 3. Keeps the matching subset alive in that VU's heap for the rest of the iteration. With 200 VUs that is 200 private half-copies of the data set live at any moment, plus ten million parses per round of iterations. The `SharedArray` is still doing its job — the shared strings are still stored once — but the script has bolted the old per-VU duplication back on top of it. None of this raises an error, fails a check or appears anywhere in the end-of-test summary; the run completes and reports normally. ## The same trap in five other disguises - `rows.map((r) => r.id)` — builds a new array of exactly `rows.length` entries. - `[...rows]` and `rows.slice()` — a full private copy of the data set. - `JSON.stringify(rows)` — reads and parses every element, then holds the whole set again as one string. - `Object.values(rows)` and `Array.from(rows)` — both walk every index and hand back a plain array. - Returning the array from `setup()` — k6 serialises setup data and gives each VU a plain copy, so the default function receives an ordinary array, not a `SharedArray`. | expression in the default function | elements read | what is allocated | |---|---|---| | `rows[i]` | one | one object | | `for (const r of rows)` | all, one at a time | one object at a time, immediately collectable | | `rows.filter(fn)` | all | a retained array of the matches | | `rows.map(fn)` | all | a retained array of `rows.length` entries | | `[...rows]` | all | a full private copy of the data set | | `JSON.stringify(rows)` | all | one string holding the whole set | The difference between `for...of` and `.filter()` is not the number of parses — both do `length` of them — it is **peak memory**: `for...of` lets each copy become garbage immediately, while `.filter()` holds every match at once. ## Where the shaping belongs 1. **Filter and map inside the loader.** k6 calls it at most once per name for the whole process, so the `SharedArray` ends up holding only the rows you want and the per-iteration cost stays at one index read. 2. **If the subset differs per VU, keep the full set shared** and compute an index instead of materialising a subset — a `SharedArray` is cheap to index and expensive to copy. 3. **If you truly need a private array, build it once during init**, not per iteration. Be clear-eyed about it though: init runs in every VU, so that is one copy per VU again, and you have chosen to pay the memory. ## How it shows up in a run The symptom is that memory grows roughly in step with the VU count even though a `SharedArray` is plainly present in the script. That shape is precisely what the type exists to flatten, so when you see it, look at the first few lines of the default function: the culprit is almost always a single `filter`, `map` or spread that never throws and never announces itself. There is a second symptom worth knowing. Iteration duration climbs with the size of the shared set even though the requests themselves are unchanged, because the parsing all happens before any request is issued. A `SharedArray` indexed once per iteration does not behave that way; one that is filtered per iteration does.
- Does for...of over a k6 SharedArray cost the same as calling .filter() on it?It performs the same number of parses, but the copies are produced one at a time and become garbage immediately, so peak memory stays flat. `.filter()` retains every match in a new array for as long as that variable lives, and it is the retention, not the parsing, that scales with VU count.
- Why does JSON.stringify() on a whole k6 SharedArray defeat its purpose?Serialising the array forces k6 to read and parse every element, then build a single string holding the entire data set in that VU's heap. Stringify only the element you actually need, which costs one parse and one small string.
- Is a k6 SharedArray still worth it when each iteration reads only one row?Yes — that is exactly the case it was built for. One index read per iteration is a single `JSON.parse`, while the memory saved is the whole data set multiplied by the VU count.
saying these in an interview costs you the question
- Thinks filter on a SharedArray returns another SharedArray
- Assumes methods that do not throw must therefore be cheap
- Filters the shared rows on every iteration to pick a subset
- Spreads the SharedArray into a plain array for convenience
- Believes only writes, never reads, can undo the sharing