In automated prompt search, how do you choose the proposer model and candidate budget?
answer
- two models, two very different call counts
- N proposals versus N times M scorings
- proposer tokens are the rounding error
- prompts do not transfer between target models
- broad pool early, full evaluation late
basics
~20 sGeneration is usually the cheap half: proposing N prompts costs N calls, while scoring them costs N times the evaluation-set size. That asymmetry argues for a strong proposer and a small, well-chosen pool, with the optimized prompt always validated on the model that will actually serve it.
solid answer
~50 sDo the arithmetic before the aesthetics. Generating one hundred candidates is one hundred model calls; scoring them against a two-hundred-example dev set is twenty thousand. As of mid-2026 that asymmetry means proposer tokens are a rounding error in most optimization runs, so use the strongest available model as proposer even if production runs on a cheap one — better proposals raise the ceiling of the whole search for very little money. The budget lever that matters is the pool size and the dev-set size, because their product sets the cost of every round. Two constraints temper this: prompts optimized against one target model transfer poorly, so the *target* in the scoring loop must be the model you will deploy, not the proposer; and if data cannot leave your environment, the proposer must be a local or self-hosted model, which caps proposal quality and should be planned for rather than discovered late.
go deeper
Know that the model writing candidate prompts and the model those prompts are tested against can be different, and that testing costs far more calls than writing does.
Be able to state the cost equation — candidates times one proposer call, versus candidates times evaluation examples of target calls — and explain why that asymmetry justifies a strong proposer with a modest pool.
Show the operational judgment: score against the model you will actually deploy, re-run the optimization when the target model changes, keep the production prompt in every round as a baseline, and log enough to reproduce the run later.
Own the allocation. Set the split between broad early screening and precise late rounds, define the stopping rule when candidate differences fall below evaluation noise, and decide when the remaining budget is better spent on data, retrieval or evaluation than on more prompt search.
## Separate the two models in your head An optimization run involves two distinct roles, and conflating them causes most of the bad decisions here. - The **proposer** writes candidate prompts. It runs once per candidate. - The **target** executes those prompts during scoring, and is the model that will serve production. It runs once per candidate per evaluation example. They need not be the same model, and usually should not be. ## The cost arithmetic that drives every decision Per round, roughly: - generation cost ≈ *N candidates* × one proposer call - scoring cost ≈ *N candidates* × *M evaluation examples* × one target call With N = 100 and M = 200, that is 100 proposer calls against 20,000 target calls — a 200× ratio, before counting that scoring calls carry the full prompt plus a real input while a proposal call is short. As of mid-2026, under any mainstream pricing, generation is a small single-digit percentage of run cost. Three consequences follow directly: 1. **Spend on proposal quality.** Using a frontier model as proposer while serving a small one in production is a good trade: you pay proposer prices a hundred times, and the resulting prompt is used millions of times. Downgrading the proposer to save money optimizes the wrong term. 2. **The real budget levers are N and M.** If a round is too expensive, cut pool size or evaluation-set size — but know what each costs you. Cutting N narrows what the search can find; cutting M raises evaluation noise, and past a point every candidate looks the same because the measurement cannot resolve them. 3. **Dedup pays twice.** Removing a duplicate candidate saves M target calls, not one. ## Choosing the proposer Beyond raw capability, four criteria: - **Instruction-writing quality.** Some models write crisper, more literal instructions; others hedge and pad. Padding matters because prompt length is a permanent inference cost after the run. - **Diversity of output.** A proposer that returns near-identical phrasings under resampling is a poor generator regardless of its benchmark scores. Two smaller proposers from different families can out-diversify one large one. - **Deployment constraints.** If your labelled data is regulated — patient records, client legal matter — the proposer sees that data whenever you induce instructions or bootstrap exemplars from it. That forces a self-hosted proposer, and the quality cap is a planning input, not a surprise. - **Reproducibility.** Log the proposer identity, version and sampling settings with every candidate. When a prompt regresses six months later, "which generator produced this and under what settings" is the first question, and a provider-side model update can silently change the answer. ## The transfer trap A prompt optimized against one target model frequently loses much of its gain when moved to another — different families respond differently to persona framing, explicit reasoning cues, negative constraints and output-format instructions, and the effect is large enough to erase a hard-won few points. The operational rules that follow: - **Score against the deployment target.** A cheap proxy target makes scoring affordable but optimizes for the wrong model. If you must use a proxy for cost reasons, validate the finalists on the real target before shipping. - **Re-run on model change.** A target-model upgrade invalidates the optimization, not just the benchmark. Budget a re-run as part of any model migration. - **Keep the baseline in the pool.** Always score the current production prompt alongside the candidates, in the same round, on the same data. It is the only honest reference point, and it catches the case where the whole round regressed. ## Allocating a fixed budget across rounds Given a fixed spend, the interesting question is how to split it. Broad-then-narrow generally wins: early rounds use a large, diverse pool scored on a *smaller* evaluation subset to cheaply eliminate the clearly bad, and later rounds use a small pool of survivors scored on the full set for a reliable ranking. This is a screening strategy — it trades early precision for coverage, on the reasoning that early rounds only need to separate bad from plausible, while the final round needs to separate good from best. Finally, decide up front what would make you stop. Automated prompt search has a real floor: once candidates are separated by less than evaluation noise, further rounds buy nothing but cost, and the honest move is to stop and spend the remaining budget on better data, better retrieval, or a better evaluation set instead.
- Would you ever use a weaker model as the proposer on purpose?Occasionally. If the deployed target is small, a frontier proposer may write instructions that assume reasoning the target cannot perform, producing candidates that score poorly for a capability reason rather than a wording one. Matching the proposer closer to the target can help there. It is also the pragmatic answer under data-residency rules, where a self-hosted smaller model is the only one allowed to see the examples.
- How do you budget the evaluation-set size against the candidate pool size?They multiply, so treat the product as the budget. Early screening rounds can use a subsample large enough to reject obviously bad candidates and no larger; final rounds need the full set because the surviving candidates are close together and the ranking must survive noise. Confirm the winner once on a held-out split the search never scored against, which costs one candidate's worth of calls.
- What do you log so an optimized prompt is reproducible a year later?The proposer model and version with its sampling settings, the target model and version used for scoring, the exact data splits, the metric definition, the full candidate pool with each candidate's score, and the baseline score from the same round. Without the baseline in the same round you cannot tell later whether a prompt was genuinely better or the evaluation set simply moved underneath it.
saying these in an interview costs you the question
- Downgrading the proposer to save money when scoring dominates cost
- Assuming an optimized prompt transfers to any target model
- Scoring only against a cheap proxy of the deployment model
- Growing the candidate pool without accounting for the evaluation multiplier
- Not scoring the current production prompt alongside the candidates