skip to content

In DSPy, what does compiling a program with BootstrapFewShot actually produce?

level: seniorimportance: must knowfreq 40%

answer

  1. prompt state, not weights
  2. a new program object comes back
  3. keep traces that pass the metric
  4. harvest each module's own in/out pairs
  5. demonstrations cost tokens forever

basics

~20 s

Compiling returns a new copy of the program whose predictors now carry selected demonstrations (and, with instruction optimizers, rewritten instruction text). No model weights change. That prompt state is the artifact, and it can be saved to JSON and loaded later.

solid answer

~50 s

An optimizer's `compile()` call takes your program, a training set and a metric, and returns a **new program object with optimized prompt state attached to each predictor** — demonstrations, and for instruction optimizers like MIPROv2 also rewritten instructions. Nothing is fine-tuned; the artifact is prompt text and examples, serializable to JSON and loadable into a fresh process. `BootstrapFewShot` builds it by *self-teaching*: it runs the program (by default a copy of itself as teacher) over training inputs, keeps the traces whose final output passes the metric, and harvests each module's own inputs and outputs from those successful traces as that module's demonstrations, capped by `max_bootstrapped_demos` and `max_labeled_demos`. That is how an intermediate module you never labelled gets examples — the label comes from the end-to-end metric, not from an annotator. Practically: from fifty labelled question-answer pairs you can compile a two-module retrieve-then-answer program without ever writing a single retrieval label.

code

python · 8 lines
python
import dspy

def grounded(example, pred, trace=None):
    return example.answer.lower() in pred.answer.lower()

optimizer = dspy.BootstrapFewShot(metric=grounded, max_bootstrapped_demos=4)
compiled = optimizer.compile(ArchiveQA(search), trainset=trainset)
compiled.save("archive_qa.json")

go deeper

for a junior

Know that compiling attaches few-shot examples to the program and returns a new program object, and that no model training or weight update happens.

for a middle

Explain the trace-filtering mechanism: run the program, score the final output, keep passing traces, and take each module's own inputs and outputs from them as demonstrations.

for a senior

Show operational judgment — a weak metric bakes bad exemplars into production, demonstrations cost tokens on every call, and the artifact needs a held-out evaluation before it ships.

for a principal

Own the artifact lifecycle: treat compilation as a build step with pinned inputs, a versioned output, a cost budget, and a rollback story rather than something run ad hoc before a deploy.

## The mental model: compile, not train The word *compile* is chosen deliberately. You hand the optimizer a program, a training set and a metric; you get back a program of the same shape whose prompts have been filled in. The analogy with a compiler is exact in one respect and misleading in another. Exact: source (the declarations) plus a build step produce an artifact you ship. Misleading: the build step is stochastic and involves running the model many times, so it is closer to a profile-guided build than to `javac`. What comes back is **not weights**. Every optimizer in this family leaves the base model untouched and changes only what is put in front of it. That is the first thing to say in an interview, because it is what makes the artifact portable, cheap to store, and inspectable. ## What is attached, concretely Each predictor inside the compiled program ends up carrying two kinds of state: 1. **Demonstrations** — input/output examples rendered into that predictor's prompt as few-shot exemplars. 2. **Instruction text** — the natural-language instruction for that step. `BootstrapFewShot` leaves this as you wrote it; instruction optimizers such as MIPROv2 propose and select rewrites. Because that is all it is, the compiled program serializes cleanly: saving to a JSON file writes out the per-predictor demonstrations and instructions, and loading restores them into a freshly constructed program. You can diff two artifacts and read exactly which examples the optimizer chose — which is far more auditable than most prompt-optimization stories. ## How BootstrapFewShot fills an unlabelled middle The mechanism worth being able to explain is self-teaching by trace filtering. Suppose a two-module program over a museum catalogue: a retrieval step and an answering step, and fifty labelled items of the form *question -> gold answer*. You have labels for the pipeline's ends, and nothing for what should happen in the middle. The optimizer runs the program on the training inputs, using a teacher — by default a copy of the same program, optionally a stronger or differently-configured one. Every run leaves a **trace**: the actual inputs and outputs of each module on that item. The final output is scored with your metric. Traces that pass are kept; traces that fail are thrown away. Then, for each module, the optimizer harvests that module's own inputs and outputs from the kept traces and uses them as candidate demonstrations for that module. The consequence is the point: **intermediate supervision is inferred from end-to-end success**. The answering step gets demonstrations of the form (context, question) -> answer that were produced by the program itself and are known to have led to a correct final answer. Nobody labelled them. This is why fifty end-to-end examples can compile a multi-step program. Two caps bound the result: `max_bootstrapped_demos` limits how many self-generated demonstrations each predictor keeps, and `max_labeled_demos` limits how many raw training examples can be used directly. Both exist because demonstrations cost prompt tokens on every production call, forever. ## The failure modes a senior candidate names - **Metric-shaped demonstrations.** The optimizer keeps whatever passes your metric. A loose metric — substring match, a lenient judge — will happily certify traces with sloppy reasoning, and those become the exemplars baked into production prompts. The demonstrations inherit every weakness of the metric. - **Nothing to bootstrap from.** If the uncompiled program almost never passes the metric, there are few or no successful traces to harvest, and the compile returns something barely better than what you started with. The fix is a stronger teacher or an easier intermediate target, not a bigger training set. - **Overfitting to the training set.** Demonstrations selected to maximize a score on fifty items can encode idiosyncrasies of those items. A held-out evaluation set is not optional. - **Token cost.** Every demonstration is paid for on every request in production. A compile that improves accuracy by two points while tripling the prompt length may be a bad trade, and the comparison should be made explicitly. - **Leakage.** Bootstrapped demonstrations are drawn from your training data. If that data contains anything sensitive, it is now embedded in the shipped prompt. ## What to say when asked "so what did you actually ship?" The correct answer is: a JSON file of per-predictor instructions and demonstrations, produced by a build step, evaluated on a held-out set, and versioned alongside the code and the model identifier it was compiled against. That framing — artifact, build, evaluation, version — is what distinguishes someone who has run this in production from someone who has read the README.

  • Where do demonstrations for a middle module come from if you never labelled that module?
    From the program's own successful runs. The optimizer executes the program on training inputs, scores the final output with the metric, discards failing traces, and takes each module's actual inputs and outputs from the traces that passed. End-to-end success is used as a proxy label for every intermediate step, which is why a small end-to-end training set can supervise a multi-step program.
  • What happens if the uncompiled program almost never passes your metric?
    Bootstrapping has nothing to harvest and the compile barely improves anything, because it can only keep traces that already succeeded. The remedies are to use a stronger teacher configuration to generate the traces, to relax or restage the target so some runs succeed, or to fix the program structure first. Throwing more training examples at it does not help — the bottleneck is the success rate, not the sample count.
  • How does MIPROv2 differ from BootstrapFewShot in what it emits?
    BootstrapFewShot changes only demonstrations. MIPROv2 also proposes instruction rewrites and searches over combinations of instructions and demonstration sets across every predictor jointly, evaluating candidates on batches of the training data. The emitted artifact is the same kind of object — per-predictor instructions and demos — but it costs far more model calls to produce, because it explores a much larger space.
  • Should the compiled artifact live in version control?
    Yes, treat it as a build output that is nonetheless pinned. Commit or store it alongside the model identifier, optimizer configuration, seed and training-set version used to produce it, so a deployment is reproducible and a regression can be bisected. Recompiling on every deploy makes production behaviour non-deterministic in a way you cannot roll back.

saying these in an interview costs you the question

  • Saying compilation fine-tunes or updates the model's weights
  • Thinking every module needs its own hand-labelled training data
  • Assuming demonstrations are free once compiled
  • Shipping without a held-out evaluation set separate from the training set
  • Believing a weak metric still yields good demonstrations

context