skip to content

A job is written entirely as per-record function bodies - which parts would you re-express as named operators, and which would you leave?

level: seniorimportance: should knowfreq 42%

answer

  1. move the seam, do not pick a side
  2. name what already has a name
  3. split the body doing three things
  4. custom logic stays a body, correctly
  5. on some models you buy clarity, not speed

basics

~20 s

Re-express anything with a named equivalent - selecting, discarding, combining, grouping, joining, aggregating, sorting - and keep as a body only work with no name: bespoke parsing, custom scoring, per-key logic that remembers. Split fat bodies so the nameable half becomes visible.

solid answer

~50 s

Go body by body and ask one question: does this work have a name the engine already understands? Selecting fields, discarding rows, computing a column, grouping, joining, aggregating and sorting all do, and moving them onto a **declared-operator surface** - where the author names operations the engine understands rather than handing over a function to call per record - makes them visible to checking and to whatever rewriting the engine does. Leave as a body the work with no named equivalent: parsing an awkward format, a bespoke rule, per-key logic that carries state between records. The highest-value move is usually neither: it is **splitting a fat body** that parses *and* discards *and* aggregates, so the nameable two-thirds become named steps and only the genuinely custom part stays opaque. Do it where it is worth it - on an engine that rewrites little, you are buying legibility, not speed.

go deeper

for a junior

Know the basic sorting rule: if the work has a name the engine already understands, name it; if it is custom logic, a function body is the right home for it.

for a middle

Explain why splitting a body that does three things is usually worth more than converting a body that does one: it turns two invisible operations into visible ones at no cost to the custom part.

for a senior

Demonstrate a procedure rather than a policy, name at least one body you would deliberately leave alone, and say how you would compare results before and after the change.

for a principal

The lead's version is setting the default for a codebase: which surface new work starts on, what evidence is required before anyone hand-writes a body for work that has a name, and how that default changes if the engine underneath is replaced.

## The seam, not the doctrine Every real job of this kind is mixed, and the question is never "which surface is better". It is where the seam falls between the part the engine can reason about and the part it can only call. A job written entirely as per-record function bodies has pushed that seam all the way to the top: the engine holds a chain of **opaque steps** - steps whose bodies are ordinary code the plan rewriter cannot look inside, so they can only be called, never reasoned about - and it will execute them close to literally. Moving the seam downward means converting the work that *has* a name into named operations, so what remains opaque is only what genuinely has no name. ## Going body by body 1. **Does it have a named equivalent?** Keeping or discarding records by a condition, producing a subset of fields, computing a new field from existing ones by arithmetic or string work, grouping by a key, joining two inputs on a key, aggregating, sorting, deduplicating - all of these are in the named vocabulary of a **declared-operator surface** on essentially every engine in this class. If the body is doing one of them, name it. 2. **Is the body doing several things at once?** This is the common and the valuable case. One body that parses a line, throws away malformed records, and accumulates a per-key total is three operations wearing one coat, and the engine can see none of them. Split it: parsing stays a body, discarding and accumulating become named steps. The seam moves without anyone rewriting the hard part. 3. **Does it genuinely have no name?** Bespoke parsing of an awkward text format, a scoring rule specific to the business, a call out to another system or a model, a per-key decision that depends on what that key has seen before. Leave these as bodies. Handing work over is what the function surface is *for*, and contorting it into a chain of named operations that nobody can read is a real cost with no matching gain. ## What you gain and what you pay | | Re-expressed as named operations | Left as a handed-over body | |---|---|---| | Engine's view | a meaning it may act on | a call it must make where you put it | | Checking | references validated against the known record shape before records move | discovered when the body runs | | Expressiveness | limited to the named vocabulary | anything the language can do | | Readability | intent is explicit, if the operation fits the vocabulary | natural for custom logic, opaque for standard logic | | Cost of the change | rewriting and re-testing working logic | none | ## When to leave it alone Three honest reasons not to touch a body: - **The engine would not act on the meaning anyway.** In a **two-phase disk-handoff model** - one grouping step at a time, every intermediate written to storage before the next begins - naming the operations buys checking and clarity, not a different plan. That may still be worth it; just do not sell it as a speed-up. - **The logic belongs on the function surface.** In a **continuous record-at-a-time model**, where one fixed graph stays running and each record passes through as it arrives, per-key logic that remembers things between records is normally written as a body, and the engine offers no named operator that does the same job. Re-expressing it is not an improvement; frequently it is not possible. - **The rewrite is risky and the step is small.** A body that runs after the data has already been reduced to something small is rarely where a job's problem lives, and changing working logic has its own cost in review, testing and defects. ## Sequencing the work When the job is large, do not convert everything at once. Start where the engine's freedom is worth most - near the read, and around the steps that handle the most records - and stop when the remaining bodies are all genuinely custom. Convert one body at a time and compare results between the old and the new job on the same input, because a named operation and a hand-written body are only *approximately* interchangeable: they can differ on missing values, on ordering assumptions the body quietly relied on, and on how each treats records the other would reject. ## What an interviewer is listening for A weak candidate answers with a policy: "rewrite everything onto the named surface". A strong one answers with a procedure - go body by body, name what has a name, split the fat ones, leave the genuinely custom - and then names a case where they would deliberately leave a body alone. Being able to say why re-expression buys less on one engine model than another is what turns this from an opinion into judgment.

  • How would you verify that a converted step still produces the same answer?
    Run both jobs over the same input and compare results, stating which equality you mean: the same rows ignoring order is usually the right bar. Pay attention to missing values and to records one form rejects and the other keeps - that is where a named operation and a hand-written body most often disagree.
  • Which body would you convert first in a large job?
    The one nearest the read that handles the most records, because that is where making the work visible is worth most and where the engine's remaining freedom is greatest. Bodies that run after the data has been reduced to something small are rarely where the job's problem lives.
  • Is there a case where splitting a body makes things worse?
    Yes, when the split is artificial: forcing a single coherent piece of custom logic into two halves that must then pass an awkward intermediate record between them can cost more in clarity than it gains in visibility. Split where the seam is natural - a nameable operation sitting next to a custom one.

saying these in an interview costs you the question

  • Wants every body rewritten onto the named surface as a blanket policy
  • Says per-record bodies should never appear in a production job
  • Promises a speed-up without checking whether the engine rewrites at all
  • Converts working logic without comparing results on the same input
  • Leaves a body doing parsing, discarding and aggregating as one opaque step
  • Tries to re-express per-key logic that remembers, where no named operator exists