skip to content

A team moved a per-item computation into a dedicated Web Worker, posting each item to the worker and posting each result back, and the page got slower rather than faster. What cost are they paying on every message, and how would you fix it?

level: seniorimportance: should knowfreq 40%

answer

  1. the copy is not free
  2. serialization happens on your thread
  3. two copies per round trip
  4. overhead per message, not per byte only
  5. make the unit of work bigger

basics

~20 s

Each message pays a structured clone: a synchronous serialize on the sender's thread plus a deserialize on the receiver's, twice per round trip, plus a task dispatch on each side. With tiny work per item that overhead dominates. Fix it by batching many items into one message.

solid answer

~50 s

Every `postMessage` copies its payload. The serialization runs **synchronously on the calling thread** — so the main thread is blocked while it happens, even though the "work" is supposedly in the worker — and the deserialization runs on the receiving thread before the handler is invoked. A round trip therefore pays two full copies plus two task dispatches, and the cost scales with the size and node count of the object graph, not just its byte length. If each item takes microseconds to compute, that per-message overhead is the whole runtime, and you have added a thread for nothing. The fixes are all about granularity: send one message carrying a batch instead of one per item, use compact flat payloads such as typed arrays rather than deep object graphs, keep long-lived state resident in the worker instead of shipping context on every call, and measure by wrapping `performance.now()` around the `postMessage` call to see how much main-thread time the copy alone is costing.

code

javascript · 10 lines
javascript
const worker = new Worker('/work.worker.js');
const items = Array.from({ length: 10000 }, (_, i) => i);

// Chatty: 10000 clones and 10000 task dispatches on each side
// for (const item of items) worker.postMessage(item);

// One clone of one compact buffer, one dispatch
const t0 = performance.now();
worker.postMessage({ type: 'batch', items: Float64Array.from(items) });
console.log('main-thread clone cost:', performance.now() - t0);

go deeper

for a junior

Remember that every postMessage copies its payload, so sending data across the boundary is real work rather than a free handoff. Sending many tiny messages costs more than sending one larger one.

for a middle

Explain the cost model: a synchronous serialize on the sender, a deserialize on the receiver, a task dispatch on each side, and scaling with graph node count rather than byte size alone. Then show batching as the fix.

for a senior

Diagnose it concretely — time postMessage with performance.now(), spot long tasks inside message dispatch in a profile — and reshape the protocol: batch, flatten to typed arrays, keep state resident in the worker, and return only what the page needs.

for a principal

Set the boundary policy: define the granularity at which work crosses threads, budget main-thread time spent on serialization, and be willing to conclude that a job does not belong in a worker at all when the payload dominates the computation.

## What a message actually costs It is tempting to picture `postMessage` as handing a pointer across a thread boundary. It is not. The browser runs the structured clone algorithm: it walks the value you passed, and produces an independent reconstruction for the receiver. That has three separately measurable costs. **Serialization, on your thread, synchronously.** `postMessage` does not return until the copy has been made. This is the detail that surprises people: if you post a 40 MB object graph from the main thread, the main thread is blocked for the duration of that copy. The worker has not saved you anything yet — you have added main-thread work in order to move other main-thread work away. **Deserialization, on the receiving thread.** Before the `message` handler runs, the receiver rebuilds the value. That time is charged to the receiver, and if the receiver is the main thread — which it is for every result coming back — it is main-thread jank again. **A task dispatch on each side.** The message is delivered as a task on the receiver's event loop. It waits behind whatever that thread is already doing, and it has fixed per-message overhead of its own. So a request/response round trip costs two clones and two task dispatches. Cost scales with the *shape* of the graph, not merely its size in bytes: a hundred thousand small objects, each with a few properties and a couple of references, is expensive to walk even if the total data is modest. Ten thousand numbers in a `Float64Array` is cheap by comparison, because it is one contiguous buffer rather than a hundred thousand graph nodes. ## Why chattiness is the specific killer Write the arithmetic out. Suppose computing one item takes 20 microseconds, and a round trip's fixed overhead — two clones of a small object plus two dispatches — is on the order of tens of microseconds. Then per item you have replaced 20 µs of main-thread compute with a comparable or larger amount of main-thread copy-and-dispatch time, *plus* latency, *plus* the worker's own overhead. Ten thousand items later, the page is slower and the main thread is no freer. Meanwhile a single message carrying all ten thousand items pays the fixed overhead once and one clone of one compact structure. This is the general rule: **a worker pays off when the unit of work you ship is large relative to the cost of shipping it.** Offloading is a bandwidth-versus-latency trade, and chatty protocols spend the entire budget on the trade itself. ## Diagnosing it You do not have to guess. Two direct measurements settle it: ```js const t0 = performance.now(); worker.postMessage(payload); console.log('clone cost on this thread:', performance.now() - t0); ``` Because serialization is synchronous, that interval *is* the copy cost, attributed to the thread that paid it. Do the same around the point where the reply handler starts versus when the worker posted, to see queueing and deserialization. In a profiler, the same thing shows up as long tasks on the main thread that sit inside the `postMessage` call or inside message dispatch rather than inside your own functions — a strong tell that the payload, not the algorithm, is the problem. ## Fixing it **Batch.** Post one message with an array of items and get one message back with an array of results. This is almost always the largest single win, and it costs a few lines. **Flatten and compact the payload.** Prefer typed arrays and flat records over deep nested graphs with many small objects. If you are moving numeric data, a `Float64Array` or `Int32Array` is dramatically cheaper to clone than the equivalent array of objects. **Keep state in the worker.** If every call ships a large context object so the worker can do its job, invert it: send that context once at startup, let the worker hold it, and send only the small per-call parameters afterwards. Shipping the same structure repeatedly is the clearest sign that ownership sits on the wrong side. **Send less back.** Results are often much smaller than inputs, or can be — if the page only needs an aggregate or an index, compute it in the worker and post that, rather than posting the whole transformed dataset back for the page to reduce. **Move a coarser unit of work.** Rather than "compute this item", make the worker's API "compute this whole page of results" or "parse this file". Coarse APIs amortize the boundary; fine-grained ones are dominated by it. There is also a way to hand certain buffers over without copying them at all, which changes this arithmetic completely for large binary payloads — but that is a distinct mechanism with its own ownership rules, and it does not rescue a design whose real problem is that it sends ten thousand messages instead of one. ## When the honest answer is "no worker" Sometimes the conclusion is that the work does not belong in a worker. If the job is a few milliseconds total, the boundary costs more than the job. If the job is genuinely long, a worker is right — but then make sure the first thing you check is whether the payload crossing the boundary is large, because a slow `postMessage` on the main thread is exactly the jank you were trying to eliminate.

  • Why does posting a very large object from the main thread hurt, even though the worker will do all the work?
    Because `postMessage` serializes synchronously before returning, and that serialization runs on the calling thread. Posting a large graph from the main thread blocks it for the whole copy, producing exactly the long task you moved the computation to avoid. Time it with `performance.now()` around the call — the interval is the copy cost, charged to the main thread.
  • Why is an array of a hundred thousand small objects so much more expensive to post than a typed array of the same numeric data?
    Cloning cost tracks the number of nodes in the object graph, not just the byte count. Each small object means property enumeration and reconstruction on both sides, whereas a typed array is a single object over one contiguous buffer. Flattening numeric payloads into typed arrays is often a larger win than any algorithmic change on either side.
  • How would you decide whether a job belongs in a worker at all?
    Compare the work per message against the cost of the boundary. If the computation is milliseconds and the payload is large, the two clones plus dispatch can exceed the job, and you are better off chunking on the main thread or reducing the work. Workers pay off for long, self-contained jobs over payloads that are small, compact, or resident in the worker already.

saying these in an interview costs you the question

  • Assumes postMessage passes a pointer with no copy
  • Thinks serialization happens on the worker's thread
  • Blames the worker's algorithm without measuring the payload
  • Sends the same large context object on every call
  • Believes more workers fixes per-message overhead

context