skip to content

A per-record body written for a second language runtime dominates a job's time. What does each record pay, and what changes if records cross in batches?

level: seniorimportance: should knowfreq 54%

answer

  1. fixed cost per crossing, not per byte
  2. convert out, cross, cross back, convert in
  3. trivial body, identical crossing
  4. batching divides the fixed cost
  5. batch size trades memory and delay

basics

~20 s

Each record leaves the engine's own runtime and comes back: it is converted into a form both sides read, crosses a process boundary twice, and is converted back, on top of the body's own work. Handing many records over at once pays the crossing cost once for the batch instead of once per record.

solid answer

~50 s

A **cross-runtime hand-off** is records leaving the engine's own runtime so a second language runtime can process them, and coming back. Per record the job pays four things on top of the body's work: converting the value into a form both sides understand, crossing out of the engine's process, crossing back, and converting the result. None of those depends on how much the body does, so a trivial body can be dominated by them entirely. Moving records over in batches amortises the fixed part of the crossing across many records and lets the conversion work on many values at once. The price is that the body must be written to take a batch, both sides hold one in memory at the same time, and in a continuous job a larger batch raises the delay before the first result. What varies: some engines have no second-runtime path at all, and some can execute common expressions natively so the crossing never happens.

go deeper

for a junior

Recall that a body written for another language runtime does not run inside the engine's worker; records must be carried out to it and the results carried back, and that carrying costs something per record.

for a middle

Explain the four per-record costs — convert out, cross, cross back, convert in — and why they are independent of what the body computes, which is what makes batching pay.

for a senior

Separate the three stacked costs before you act: lost rewriting, the per-record crossing, and the body's own work. Say what batch size trades away, and when removing the body beats tuning it.

for a principal

Ask whether the platform should support author-written bodies in a second runtime at all, given the crossing cost, the memory that lives outside the worker's budget, and the operational story when that runtime fails.

## What a hand-off actually is An **opaque step** is a step whose body is ordinary code the engine's plan rewriter cannot look inside, so it can only be called, never reasoned about. A **cross-runtime hand-off** is the harder case of one: records leave the engine's own runtime so a second language runtime can process them, and come back. The engine's worker process cannot simply call the body, because the body is not code the worker can execute. Something has to carry values across. That carrying has a fixed part and a per-byte part, and the fixed part is what hurts. ## The per-record bill For each record handed over, the job pays, in addition to whatever the body itself does: 1. **Conversion out** — the value is turned into a representation both runtimes agree on. 2. **The crossing out** — whatever mechanism connects the two runtimes is exercised: a write, a wake-up, a context switch on the other side. 3. **The crossing back** and **conversion in** — the result makes the same journey in reverse. 4. **A second memory footprint** — the other runtime holds its own copy of whatever is in flight, in memory the engine's worker does not account for. What that does to a worker's memory budget is owned by The Worker's Budget; note only that it is not free. None of those four scales with how clever the body is. A body that adds two numbers pays exactly the same crossing as a body that runs a model. That is why the classic symptom is a job whose profile shows the overwhelming majority of its time in a step that does almost nothing. ## Why batching changes the arithmetic Hand over many records in one crossing and the fixed cost is divided across all of them. Three effects compound: - **The crossing is paid once per batch**, not once per record — usually the largest single win. - **Conversion gets cheaper per value**, because converting many values of the same type together is a tighter loop than converting one at a time, and per-value bookkeeping is shared. - **The second runtime can work over many values at once**, which in most languages is dramatically faster than a loop over single values. It is not free. The body must now be written to accept many records and return many results, which is a different function signature and a different set of bugs. Both sides hold a batch's worth of data at once, so the memory cost rises with batch size. And in a continuous job, a batch that must fill before it crosses adds waiting to the delay between a record arriving and its result appearing — a real trade, not a detail. | | one record per crossing | many records per crossing | |---|---|---| | fixed crossing cost | paid once per record | paid once per batch | | conversion | per value, with per-value bookkeeping | shared across the batch | | body's shape | takes one record, returns one result | takes many, returns many | | memory in flight | one record on each side | one batch on each side | | delay to first result | lowest | rises with batch size | ## What varies between engines This is a place where a claim true of one engine is flatly false of the next. - Some engines in this class offer **no second-runtime path at all**: the body must be written in the engine's own language, and the whole question disappears. - Where a path exists, engines differ in **what carries the records** — a child process alongside each worker, an embedded interpreter, a shared memory region — and therefore in how large the fixed cost is. - Some engines **avoid the crossing entirely for expressions they recognise**, executing a subset of the second language's operations natively. Then the cheapest optimisation is to write the body out of existence. - Where batching exists, whether it is the default, and whether the author must write a different kind of body to get it, both vary. ## The judgment being tested A senior answer does not stop at "it is slow because of the boundary". It separates the three costs that stack in this step: the rewriting the engine gave up because the body is unreadable, the crossing paid per record, and the body's own work. They have different fixes — restructure the graph, batch the hand-off, rewrite the body — and reaching for the third when the first two are dominating is the common and expensive mistake.

  • Why can a body that does almost nothing still dominate the job?
    Because the crossing cost is fixed per record and independent of the body. Conversion out, two process crossings and conversion back are paid whether the body adds two numbers or runs a model, so with a cheap body they are essentially the whole cost of the step.
  • What is the downside of making the batch as large as possible?
    Both runtimes hold a batch at once, so memory in flight grows with it, and in a continuous job a batch that must fill before it crosses adds waiting before the first result appears. Past a point the fixed cost is already amortised and larger batches only buy those two problems.
  • When is the right move to remove the hand-off rather than tune it?
    When the body expresses something the engine already understands — a comparison, an arithmetic expression, a string operation. Written as named operators it stops being unreadable, crosses no boundary and becomes rewritable again. Some engines also recognise a subset of such expressions and execute them natively without the author doing anything.

saying these in an interview costs you the question

  • Blames the second language's speed rather than the crossing
  • Thinks batching removes conversion entirely instead of amortising it
  • Assumes every engine of this class offers a second-runtime path
  • Says a larger batch is always better, ignoring memory and delay
  • Treats the lost rewrites and the crossing cost as one problem
  • Believes the crossing happens once per piece of the input, not per record