skip to content

Macros and Compile-Time Metaprogramming

Rewriting code before it compiles: token pasting versus syntax-tree transformation, the name-capture problem, and expansion order. Interviewers ask about hygiene, the line past text substitution.

on this pageshow

questions

13

When a macro body runs while the program is being compiled, which values can it splice into the code it emits?

level: middleimportance: must knowfreq 58%

answer

  1. two programs in one file
  2. the build's process, then the artifact
  3. arguments arrive as code, not values
  4. a value must render back as code
  5. no live handle crosses the boundary

basics

~20 s

Only values the expanding stage can render back as code: literals, names the emitted program can resolve, or fragments it built by quoting. A live compile-time object cannot be smuggled into run time; it has to be reconstructed.

solid answer

~40 s

A compile-time expansion splits one file into two programs. The **expanding stage** runs inside the build, in the toolchain's process; the **emitted program** is the code it hands back, which the compiler then compiles and which runs much later. The stage is handed its arguments as code fragments, not as values, so it can inspect their shape but cannot know what they will evaluate to. Anything the stage computes lives only in the build process, and reaches run time only by being rendered back into code: a literal, a reference to a name the emitted program can resolve, or a constructor call. Things with no code form -- an open handle, a closure over build-process memory -- cannot cross at all; the stage must emit code that rebuilds them.

code

pseudocode · 10 lines
pseudocode
// stage 1 - runs while the project is being built
macro squares_up_to(n):
    parts = []                     // a list that exists only during the build
    for i from 1 to n:             // this loop runs now, not later
        parts.add(literal(i * i))  // render the computed number as code
    return quote( list_of( splice(parts) ) )

// stage 2 - what the compiler is finally handed at the call site
//   squares_up_to(4)   ->   list_of(1, 4, 9, 16)
// the list named parts is gone; only the rendered literals remain

go deeper

for a junior

Hold on to the shape: one file, two moments. The macro body runs while the project builds; the code it hands back runs later, when the program actually runs.

for a middle

Explain that arguments arrive as code fragments rather than values, and that anything the stage computes reaches run time only by being written back out as code the compiler then compiles.

for a senior

Show where this bites in production: a value folded in at build time that nobody can explain months later, or a stage that tried to hold a live resource and quietly emitted a stale constant instead.

for a principal

Frame it as a cost transfer. Work moved into the build is paid once by every builder and never by a user. Decide which computations deserve that trade, and who owns the build time it consumes.

## Two programs live in one file A source file that uses compile-time expansion really describes two programs that never overlap in time. The first is the **expanding stage**: the macro body itself. It runs *inside the build*, in the toolchain's own process, on whatever machine is compiling. The second is the **emitted program**: the code the stage hands back, which the compiler then compiles into the artifact and which runs later, wherever that artifact is deployed. By the time the artifact exists, the stage has finished and is gone -- there is no process left to call back into. Almost every confusion about macros is really a confusion about which of those two programs a given line belongs to. ## What the stage is actually handed The stage is not called with the values the program will have. It is called with **code fragments**: descriptions of the arguments *as they were written at the call site*. If a caller writes `total(x)`, the stage receives something meaning "a call named `total` applied to a name `x`" -- not a number. So the expanding stage can: - inspect the shape of an argument: is it a literal, a name, a call, a block; - read the value of an argument that genuinely is a literal, or of anything else already known while compiling; - count arguments and read a written-down type description; - decide what code to produce, including producing nothing at all; - place an argument fragment into the emitted code once, several times, or never. And it cannot: - read the value an argument will have when the program runs; - observe anything about the running program, which has not started; - hand the emitted program a reference to an object the stage is holding. ## Crossing the stage boundary The one-way street from build to run time is the whole mechanism. A value computed while expanding reaches the emitted program **only if it can be written back out as code** -- this rendering step is what the stage's output is made of. Most toolchains give the stage a quoting notation for exactly this (a fragment is *quoted* and computed pieces are *spliced* into it, the idea usually called **quasiquotation**), plus a way to turn a plain value into a literal. | Computed while expanding | Reaches the emitted program? | How it crosses | |---|---|---| | A number, string or boolean | Yes | rendered as a literal | | A list of such values | Yes | rendered as a literal sequence | | A name the program can resolve | Yes | emitted as a reference | | A structure with a code form | Yes | emitted as a constructor call | | An open handle or connection | No | emitted code must open its own | | A closure over build-process state | No | must be re-expressed as emitted code | The reverse direction is closed outright: nothing from run time is available while expanding, because run time has not happened. ## What the split buys, and what it costs - **Work moves, it does not vanish.** A computation done by the stage is paid once, during the build, by everyone who builds -- and never by a user. The artifact contains only the result. - **Failures move with it.** A bug in stage code is a broken build, not a production incident. That is usually a good trade, and it is why validating inputs in the stage is worth doing. - **Evidence disappears.** The emitted program shows no sign the stage existed unless the stage emitted something identifying. A constant folded in during expansion looks exactly like a constant somebody typed. - **Only literal-shaped results are cheap.** If the natural result of a stage computation has no code form, the stage ends up emitting the recipe rather than the result, and the work happens at run time after all. - **Toolchains differ in how much the stage may do.** Some restrict the expanding stage to a limited, effect-free subset of the language; others let it run ordinary code with full access to the build machine. Know which kind you are working in before you rely on either. ## Reading it in review When you read an unfamiliar macro, answer three questions in order: 1. **Which lines run during the build?** Everything in the macro body that is not inside a quoted fragment. 2. **What is the emitted text, literally?** Write out the expansion for one real call site by hand. 3. **Which values crossed the boundary, and as what?** Every spliced value should be traceable to a literal, a name, or a constructor call in the emitted text. Anything you cannot account for is a misunderstanding waiting to become a build failure.

  • How does a value known only while expanding become visible to a debugger at run time?
    It does not, unless the stage emitted it. If you want the value inspectable, emit it as a named constant or a field initialiser rather than folding it into the middle of an expression. Then it exists in the artifact like any other declaration, and ordinary tooling can see it.
  • What actually changes when you move a computation from run time into the expanding stage?
    It is paid once, during the build, and disappears from the artifact except for its result. Build time grows, start-up and per-call cost shrink, and a failure in that computation becomes a build failure instead of an incident. The cost is that the source no longer shows the work being done.

A jig in a workshop shapes the part and then stays behind on the bench. Only the part ships, so anything the jig knew has to be stamped into the part itself.

saying these in an interview costs you the question

  • Thinks the expanding stage can read the value of a run-time variable
  • Believes an object built during the build is handed to the running program as-is
  • Assumes any compile-time value can be spliced, whatever its type
  • Says compile-time work is still paid again on every run-time call
  • Treats the expander and the emitted program as one running process
open as a page

Why can a shorthand that substitutes argument text before parsing compute the wrong value inside a larger expression?

level: middleimportance: must knowfreq 62%

basics

~20 s

Substitution happens before the grammar is applied, so the pasted characters are parsed together with whatever surrounds the call. The caller's operators can bind tighter than the body's, regrouping the expression into something the author never wrote.

open as a page

A compile error points at a line nobody typed, inside a macro's output, so what makes that message actionable?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Source positions carried on every emitted node, plus an expansion trace. Text derived from the caller's argument should point back at the call site, text the macro authored at the macro, and the trace should name the chain that produced it.

open as a page

An expansion introduces a temporary binding into the caller's block - how can that silently change the caller's meaning?

level: seniorimportance: must knowfreq 58%

basics

~20 s

The introduced binding shadows a caller name of the same spelling, so the caller's own references - including argument text spliced into the body - resolve to the expansion's temporary instead. This is accidental capture, and it compiles cleanly.

open as a page

When a compile-time expansion substitutes one argument expression at two places in its body, what does the caller lose?

level: middleimportance: should knowfreq 52%

basics

~20 s

Substitution copies the argument's expression, not its value, so the expression is evaluated once per occurrence that control flow reaches. Side effects and cost repeat, and two evaluations can disagree. Bind it once into a temporary instead.

open as a page

Why can a shorthand that expands to two statements leave only its first statement inside a conditional branch?

level: middleimportance: should knowfreq 46%

basics

~20 s

Because substitution inserts statements, not a block. In a grammar where a branch takes one statement unless delimiters group several, only the first expanded statement belongs to the branch and the rest becomes ordinary code after the conditional.

open as a page

When one macro call is the argument of another, does the outer expansion see the inner call already expanded?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It depends on the expander's order. Expanding the enclosing call first hands the outer macro the inner call as written, so it can inspect or discard it. Expanding arguments first hands it finished code instead.

open as a page

What can a transformer that rewrites a parsed tree reject that a text-substituting shorthand cannot?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A tree transformer is handed parsed nodes, so it can refuse a call whose arguments are the wrong number or the wrong kind of construct, and report that at the caller's position. A paste has only characters.

open as a page

Your platform team must decide what a compile-time expansion may read beyond the source it is given, so where do you draw the line?

level: principalimportance: should knowfreq 32%

basics

~20 s

Draw it at declared, version-pinned inputs checked into the repository. Anything else an expanding stage reads is baked into the artifact as a literal, so identical source stops producing identical output and build caches can no longer tell what is stale.

open as a page

Your platform team ships expansions that land inside other teams' blocks on a facility with no automatic renaming - what do you standardise?

level: principalimportance: should knowfreq 26%

basics

~20 s

Standardise what the facility cannot guarantee: a reserved name space for every introduced binding, a bind-once rule for any argument used more than once, no dependence on caller-scope names, and tests that force the collisions deliberately.

open as a page

A macro expands into a call to itself and expansion never stops, so what did the author get wrong?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

The base case was emitted as a run-time test instead of being decided while expanding. The recursive call is present in the emitted text either way, so the next pass rewrites it again and the guard never gets to run.

open as a page

An expansion's body calls a helper by name - in whose scope should that name resolve at expansion?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

In the scope where the body was written, not where it is expanded. Otherwise a caller who happens to define that name silently substitutes their own helper for the author's, changing what the expansion does without touching it.

open as a page

As a lead, how would you decide whether a codebase may use text-level substitution when a tree-rewriting facility exists?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Decide on review cost, not power. Text-level substitution is acceptable only where the body is one fully parenthesised expression; anything contributing statements, or needing to refuse bad input, belongs to a tree rewrite or an ordinary function.

open as a page