How do you handle a base-model upgrade for a compiled DSPy program in production?
answer
- artifact is tuned to one model
- pin the model id and seed
- evaluate before you recompile
- always include the uncompiled arm
- recompiling is a recurring bill
basics
~20 sTreat the compiled artifact as a build output pinned to the model it was compiled against. On an upgrade, first evaluate the existing artifact on the new model against a held-out set, then recompile and compare all three options — old artifact, new artifact, and uncompiled — before deciding what to ship.
solid answer
~50 sA compiled program is prompt state tuned for one model's behaviour, so a model change silently invalidates the assumption behind it: instructions written to correct an old model's failure mode may be dead weight, and demonstrations that scaffolded a weaker model can constrain a stronger one. The policy I would run is: **pin** the artifact with the model identifier, optimizer configuration, seed and training-set version that produced it; **never recompile implicitly on deploy**, because that makes production behaviour non-reproducible and unrollbackable. On an upgrade, evaluate three candidates on a held-out set — the pinned artifact on the new model, a fresh recompile, and the uncompiled program — and pick on quality *and* cost, since the fresh compile may attach fewer demonstrations and be cheaper per request. Budget the recompile explicitly: an instruction-and-demonstration optimizer like MIPROv2 runs many scored rollouts across every predictor, so the bill scales with training-set size, candidate count and program size.
go deeper
Know that a compiled program is tuned for the model it was compiled against, and that changing models means re-checking it rather than assuming it still works.
Explain what to pin alongside the artifact — model identifier, optimizer configuration, seed, training-set version — and why evaluating the existing artifact comes before recompiling.
Run the three-arm comparison in practice: pinned artifact on the new model, fresh recompile, and uncompiled baseline, judged on metric score, prompt tokens and latency, with a rollback artifact retained.
Own the economics and the cadence: quantify what a recompile costs as a function of program size, training-set size and search breadth, decide how often it is worth paying, and be honest that transfer across model generations is unsettled.
## Why a model upgrade is a compile-invalidating event Compilation optimizes prompt state *against a specific model's behaviour*. Everything the optimizer discovered is a correction for how that model failed on your data: an instruction clause that stops it inventing catalogue identifiers, four demonstrations that teach a format it did not follow naturally, a phrasing that suppresses a particular verbosity. Change the model and those corrections are addressed to someone who is no longer in the room. The outcomes are not symmetrical, which is what makes this a judgment question rather than a procedure: - The artifact may still **help**, because most of it encodes task structure rather than model quirks. - It may become **neutral**, wasting tokens on instructions the new model no longer needs. - It may actively **hurt**. This is the case people miss. Demonstrations that scaffolded a weaker model can anchor a stronger one to a shallower answer style than it would produce unprompted, and defensive instruction clauses can suppress capability. Because all three are plausible, the answer is never "recompile" or "keep it" — it is "measure the three candidates." ## The pin A compiled artifact without provenance is unmanageable. Store, next to the artifact: the base model identifier and version, the optimizer and its configuration, the random seed, the sampling temperature used during compilation, the training-set version or hash, and the held-out score it achieved. That turns "why did quality drop last Tuesday" from an archaeology project into a diff. Note what pinning can and cannot buy you. It gives you a **fixed artifact** — the same prompts ship every time, which is the important half. It does *not* give you a bit-reproducible compile: the compile ran a stochastic model many times, and providers change model behaviour under a stable name. Recompiling with identical inputs yields a similar artifact, not the same one. Claiming full reproducibility here is a red flag; claiming a pinned, versioned output is correct. ## The upgrade procedure 1. **Freeze.** The currently-deployed artifact stays deployed while you evaluate. 2. **Evaluate the pinned artifact on the new model** using the held-out set. This is the cheapest experiment and it often answers the question outright. 3. **Recompile against the new model** with the same optimizer configuration and training set, producing a candidate artifact. 4. **Evaluate the uncompiled program on the new model** as the third arm. On a genuinely stronger model this sometimes wins, and discovering that saves you a recurring token bill. 5. **Compare on the quality-cost frontier.** Include prompt token count per request, latency and metric score. A compiled artifact that adds four demonstrations to every production call is a permanent cost line. 6. **Ship one, keep the loser.** Rollback means redeploying the previous artifact, which only works if you kept it. ## Budgeting the recompile The recompile bill is the part principals are expected to quantify rather than wave at. For a demonstration-only optimizer, the cost is roughly the training-set size times a few rollouts. For a joint instruction-and-demonstration optimizer like MIPROv2, the optimizer proposes instruction candidates for **each** predictor, pairs them with demonstration sets, and scores combinations on batches of training data — so cost scales with the number of predictors in the program, the number of candidates explored, and the evaluation-batch size, multiplied by however many trials the search runs. A two-module program with a fifty-item training set is inexpensive; a six-module program with a thousand-item set and a heavy search setting is a real budget item, and it recurs on every model upgrade. That recurrence is the strategic point. If your provider ships a new model quarterly, you have signed up for a quarterly recompile-and-evaluate cycle. That argues for keeping the training set small and high-quality, keeping the program shallow, and automating the three-arm comparison in CI rather than doing it by hand each time. ## Contested ground, stated honestly There is no settled industry consensus in 2026 on how much prompt optimization survives a model generation. Reported experience varies by task: structured extraction tends to transfer well, while instructions that patched a specific model's reasoning weakness often do not. The defensible position in an interview is not a number — it is the process: pin everything, hold out an evaluation set, always evaluate the uncompiled arm, and treat the artifact as disposable rather than precious.
- Why not simply recompile automatically as part of every deployment?Because it makes production behaviour non-deterministic and unrollbackable. Each compile is a stochastic search, so two deploys of identical code would ship different prompts, and a quality regression could not be attributed to a change you made. It also puts an unbounded, variable cost on the deploy path. Compile deliberately, evaluate the result, commit the artifact, and deploy that.
- Which arm of the comparison do people forget, and why does it matter?The uncompiled program on the new model. Teams assume the compiled artifact is strictly better because it was better once, and never re-test the baseline. On a stronger model the demonstrations sometimes anchor output to a shallower style than the model would produce unprompted, so the plain program wins on both quality and token cost. Skipping that arm can leave a permanent, unnecessary per-request bill in place.
- Can you make a compile bit-for-bit reproducible?Not realistically. You can pin the seed, temperature, optimizer configuration and training-set version, which makes runs similar and makes differences explainable. But the compile executes a hosted model many times, and provider-side behaviour drifts under a stable model name. The achievable guarantee is a pinned, versioned artifact that ships identically every time — not a compile you can replay exactly.
- How would you keep the recompile bill from growing over time?Control the three multipliers: training-set size, program depth, and search breadth. Keep a small, curated, high-signal training set rather than an ever-growing one; keep the number of predictors low so a joint optimizer has fewer sites to tune; and use a lighter search setting for routine upgrades, reserving heavy search for genuine redesigns. Automating the three-arm comparison also removes the human cost, which is often the larger one.
saying these in an interview costs you the question
- Assuming a compiled prompt is model-agnostic and transfers untouched
- Recompiling automatically on every deploy
- Claiming a compile run is exactly reproducible if you set a seed
- Never re-testing the uncompiled program on the new model
- Ignoring the per-request token cost that baked-in demonstrations add