skip to content

Compiled Prompt Programs

The DSPy approach: declare what each step takes and returns as a signature, compose modules such as Predict, ChainOfThought and ReAct, then let an optimizer like BootstrapFewShot or MIPRO compile the actual prompt text. Interviewers ask why moving from prompt strings to a compiled program changes how an LLM pipeline is maintained.

on this pageshow

questions

4

In DSPy, what does a signature declare, and why isn't it just a prompt string?

level: middleimportance: must knowfreq 46%

answer

  1. declaration, not wording
  2. what the step takes and returns
  3. the docstring is the instruction
  4. prompt text becomes a build artifact
  5. field names carry semantic weight

basics

~20 s

A DSPy signature declares one step's named inputs and outputs plus a short description of the task — for example context, question -> answer. The wording, formatting rules and examples that make up the actual prompt are generated by the framework, not written by you.

solid answer

~40 s

A signature is a **typed declaration of one LLM call**: the input fields, the output fields, and a natural-language statement of what the step is for. You write it either inline (`"context, question -> answer"`) or as a class with `dspy.InputField()` / `dspy.OutputField()` and a docstring. What you never write is the prompt text — an adapter renders the fields, the instruction and any compiled demonstrations into a real prompt at call time, then parses the reply back into the declared output fields. The point is separation: the *contract* (`answer` is a string, produced from `context` and `question`) is stable source code, while the *phrasing* becomes a build artifact an optimizer can rewrite. Downstream code reads `pred.answer` regardless of whether the instruction was hand-written, rewritten by MIPROv2, or run against a different model.

code

python · 10 lines
python
import dspy

class ArchiveQA(dspy.Signature):
    """Answer questions about museum catalogue records using only the retrieved context."""

    context: str = dspy.InputField(desc="passages from catalogue records")
    question: str = dspy.InputField()
    answer: str = dspy.OutputField(desc="one or two sentences, grounded in the context")

step = dspy.Predict(ArchiveQA)

go deeper

for a junior

Be able to read an inline signature like "context, question -> answer" and say which side is input and which is output, and that you do not write the prompt text yourself.

for a middle

Explain the three carriers of meaning — field names, docstring instruction, field descriptions — and that an adapter renders them into the real prompt and parses the reply back into the declared fields.

for a senior

Show you know the operational consequence: the prompt is a build artifact, debugging means inspecting rendered calls rather than reading source, and a field rename is a prompt change that needs evaluating.

for a principal

Own the judgment call about when the abstraction pays. Argue the indirection cost against the maintenance win, and be clear that a single stable prompt does not justify a compiled program.

## The problem a signature solves In a hand-written LLM pipeline, one string carries at least five jobs at once: the task instruction, the names and meanings of the inputs, the desired output shape, the formatting rules that make the output parseable, and any worked examples. Because those jobs are fused into a single blob of text, changing one of them means hand-editing the others, and nothing in the repository records which part of the string was actually load-bearing. Six months later nobody can tell whether "Answer concisely." is carrying the accuracy or whether it is decoration someone pasted in during a bad afternoon. A DSPy signature separates the **declaration** of a step from its **implementation**. It says what the step consumes, what it produces, and what it is for. It deliberately says nothing about wording. ## Anatomy There are two forms. The inline form is a string: `"context, question -> answer"` names input fields to the left of the arrow and output fields to the right. The class-based form is what real programs use, because it lets you attach types and descriptions: A class-based signature has three parts that carry meaning. **Field names** are semantic, not cosmetic — the framework puts them in the prompt, so a field called `answer` and a field called `out` do not behave identically. The **docstring** is the task instruction, and it is precisely the object an instruction optimizer later rewrites. Per-field **`desc`** text is rendered as guidance next to that field. On top of that, **type annotations** (`str`, `list[str]`, `bool`, a `Literal`, or a Pydantic model) tell the adapter how to ask for the value and how to parse and validate what comes back. ## What is generated rather than written When a module built from the signature runs, an adapter turns the signature into an actual request: a section describing the fields, the instruction, whatever demonstrations the compiled program is carrying, and the formatted inputs. It then parses the reply back into the declared output fields and errors if a required one is missing. At no point did you type "You are a helpful assistant" or "Respond only with JSON." If you want to see what was really sent, you inspect the call history rather than reading a template file. ## A worked example Take a question-answering step over a museum's catalogue archive. Declared as a signature, it says: given `context` (retrieved catalogue passages) and `question`, produce `answer`, and the docstring says answers must stay within the retrieved records. Now consider the changes that pipeline will actually go through. You add a reasoning step. You move from one model to another. An optimizer rewrites the instruction and attaches four demonstrations. A colleague adds a second output field for the record identifier. In every one of those changes, the declaration is what you edit or leave alone deliberately — and the calling code still reads `pred.answer`. ## Why this changes maintenance Three consequences follow, and they are the reason interviewers ask about signatures at all. First, **the prompt becomes an output of the build, not source you maintain**. You maintain the contract and the training data; the text is regenerated. That is the whole premise of compiled prompt programs. Second, **the boundary is checkable**. A missing or malformed output field fails at the step that produced it, with the field name, instead of surfacing three functions later as a `KeyError` on a dict you parsed out of prose. Third, **changes stay local**. Swapping the implementation of one step, or letting an optimizer retune it, does not ripple through the rest of the program, because everything else was written against field names rather than against string positions. ## What it costs The abstraction is not free. You lose direct legibility: you cannot open a file and read the prompt that ships, which is genuinely uncomfortable during an incident, and it makes debugging depend on tracing the actual calls. Some prompt techniques do not fit neatly into a fields-in, fields-out shape, and forcing them through a signature is worse than writing the string. And the semantic weight of field names surprises people — engineers treat them as variable names, rename `answer` to `result` during a refactor, and measurably move their metric. Signatures buy maintainability at the price of indirection, and that is the tradeoff to state out loud in an interview.

  • If the prompt text isn't in the source, how do you debug a step that starts returning garbage?
    You inspect the calls, not the code. DSPy keeps a history of the rendered prompts and raw completions, so you look at what was actually sent — instruction, demonstrations, formatted fields — for the failing input. That usually shows one of three things: a demonstration the optimizer picked that does not generalize, an output the adapter could not parse into the declared field, or context that never made it into the input field at all.
  • Does anything about a signature actually change what the model sees, or is it purely organizational?
    It changes it materially. Field names, the docstring and the per-field descriptions are all rendered into the prompt, and the declared types drive how the request is formatted and how the reply is parsed. Renaming an output field or tightening a description is a real prompt edit, not a refactor — which is why signature changes should be evaluated, not merged on the strength of a code review.
  • When would you not reach for a signature-based declaration at all?
    When there is exactly one prompt, it rarely changes, and you have no labelled data to optimize against. The abstraction earns its keep when there are several steps, a metric, and a reason to retune — a multi-step pipeline you expect to move across models. For a single classification call behind a stable API, a plain string plus a parser is less machinery and easier to read.

saying these in an interview costs you the question

  • Calling a signature just a template with placeholders for variables
  • Assuming field names are arbitrary and can be renamed freely
  • Thinking the docstring is documentation the model never sees
  • Believing signatures make the pipeline model-independent by themselves
  • Claiming you still hand-write the prompt and the signature only validates it

context

open as a page

In DSPy, what does compiling a program with BootstrapFewShot actually produce?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Compiling returns a new copy of the program whose predictors now carry selected demonstrations (and, with instruction optimizers, rewritten instruction text). No model weights change. That prompt state is the artifact, and it can be saved to JSON and loaded later.

open as a page

In DSPy, what changes when you swap Predict for ChainOfThought on a step?

level: middleimportance: should knowfreq 34%

basics

~20 s

Swapping the module keeps the same signature but changes how that step is executed: ChainOfThought extends the declared outputs with a reasoning field the model fills before answering. Calling code, training data and metric are untouched — you change one line and recompile.

open as a page

How do you handle a base-model upgrade for a compiled DSPy program in production?

level: principalimportance: should knowfreq 28%

basics

~20 s

Treat the compiled artifact as a build output pinned to the model it was compiled against. On an upgrade, first evaluate the existing artifact on the new model against a held-out set, then recompile and compare all three options — old artifact, new artifact, and uncompiled — before deciding what to ship.

open as a page