You host the tool server your agent under test calls. Between two runs you edit both the tool's description text and its parameter schema, and the attempt starts succeeding. What is wrong with that experiment, and how would you rerun it?
answer
- two edits, one observation = confounded
- n=1 against a sampled model is noise
- factorial cells, one axis each
- hash the served tool definitions
- compare rates, not outcomes
basics
~20 sTwo variables moved, so you cannot say which caused the change — and with a stochastic model, neither may have. Rerun with one edit at a time from a pinned baseline, repeat each configuration enough times to compare success rates, and record which fixture version produced each run.
solid answer
~50 sThe point of hosting the server is that its surface becomes a set of controlled variables; editing two at once throws that away. Worse, agent behaviour is sampled, so a one-run-to-one-run difference can be noise even if you had changed nothing. Rerun it as a small factorial: baseline, description-changed-only, schema-changed-only, both. Each cell gets N attempts with the sampling settings and the rest of the environment held fixed, and you compare success *rates*, not outcomes. Give every fixture state an identifier — a variant id or a hash over the served tool definitions — and stamp it into each run record, so months later a result still names the exact surface it came from. The tradeoff is cost: four cells times N attempts times a metered endpoint. In practice you screen with the cheapest cell first and only expand the grid around a difference that looks real.
go deeper
Spots that two things changed at once, so the cause is unknown, and suggests changing one at a time.
Adds sampling: a single run per configuration proves nothing, so each cell needs repeats and the comparison is between success rates.
Adds bookkeeping and cost — pinning and stamping a fixture identifier, holding the environment fixed, and screening cheaply before spending the full grid.
Sets the standard for what the team is allowed to publish from an experiment like this, and what statistical bar a reported behaviour change must clear.
**Why the experiment is void.** Two independent defects stack here, and each alone would sink it. *Confounding.* The description text and the parameter schema are two separate axes of the served surface. Moving both between runs means the outcome is consistent with either one mattering, with both mattering, or with them cancelling. There is no arithmetic that recovers the attribution afterwards; the information was never collected. *Sampling noise.* Agent behaviour is sampled. At any nonzero temperature the same fixture, the same prompt and the same model can produce a tool call on one run and a refusal on the next, and multi-turn agents amplify this because each turn's sampled output changes the context of the turn after it. A single run per configuration therefore carries almost no information. Run your untouched baseline twice and you will often see it disagree with itself; that is the honest noise floor, and n=1 sits entirely inside it. **How to rerun it.** 1. *Pin and identify the baseline.* Take a content hash over the tool definitions the server serves — names, descriptions, schemas, in order. A hash moves on any change the model can see, including a whitespace edit nobody logged, which a hand-written variant label cannot. 2. *Move one axis per cell.* On this surface the useful axes are the description text, the parameter schema (property names, types, `enum` members, `required`, `additionalProperties`, length caps), the result body returned to the model, and the error text on a rejected call. 3. *Hold everything else fixed.* Sampling settings, model version, the tool list and its order, the seed conversation, and the harness revision. Any of these drifting between cells reintroduces the confound you just removed. 4. *Run N attempts per cell and report attempts and successes*, not outcomes. Two-by-two here is four cells: baseline, description-only, schema-only, both. The "both" cell is not optional — it is what tells you whether the two edits interact. 5. *Stamp the fixture identifier into every run record*, so the comparison is reconstructable months later. ```json { "fixture_id": "toolsrv-2026-08-31-c", "definitions_hash": "sha256:4f0c...", "axes": { "description_variant": "baseline", "schema_variant": "required-relaxed" }, "attempts": 40, "successes": 11, "held_fixed": ["sampling", "model_version", "tool_list_order", "seed_conversation", "harness_rev"] } ``` **What it costs.** The grid multiplies. Each additional axis you want to attribute doubles the number of cells, and every cell costs its own N agent attempts, each of which is a multi-turn loop of five to fifteen model calls. Four cells at forty attempts is 160 runs — roughly one to two million tokens and an hour or more of wall clock on a metered endpoint even with parallelism; three axes at the same N is eight cells and double that. The affordable pattern is to screen first: small N across all cells to find where a difference might be, then spend the real N only around that difference. Exploratory screening against a local open-weights target is cheaper still, provided you confirm the shortlist on the target that actually matters. **Where the number misleads.** A cell-to-cell difference is a comparison of two noisy proportions, and at the sample sizes red teams actually run it is usually not significant. Twelve successes in forty versus eight in forty looks like a 30 percent versus 20 percent effect and is well inside what resampling the same fixture produces; treating it as a finding is the commonest error on this surface. Report attempts and successes rather than a bare percentage, so a reader can see 3/5 was not 60 percent in any useful sense, and compare any candidate effect against your measured baseline-versus-baseline spread before naming it. **What you would check.** Rerun the baseline cell twice, unchanged, and confirm the two numbers agree closely enough that your effect is bigger than the gap. Diff the hashes of every cell to prove each differs from the baseline in exactly the intended place. Confirm the sampling settings and model version recorded in the run metadata are identical across cells. And separate the two questions that this design deliberately keeps apart: "which knob moved the agent" is answered by the grid, while "does that knob exist on the real server" is answered only by diffing the fixture against the real contract. A schema variant the real service would reject can be a perfectly clean experimental result and still be untransferable.
- Why hash the served tool definitions instead of writing the variant name by hand?The hash changes on any edit the model can see, including an accidental one; a hand-written label can silently claim a state the server was not actually in.
- How do you keep the factorial affordable against a metered endpoint?Screen the cheapest cells first with a small N, expand only around an apparent difference, and reuse a local open-weights target for the exploratory grid where the behaviour transfers well enough to shortlist.
Each knob you want to attribute doubles the grid: two knobs is four cells, three is eight, and every cell costs its own N attempts. The experiment gets rigorous linearly and expensive exponentially.
saying these in an interview costs you the question
- Reports the combined edit as the cause without separating the axes.
- Compares single runs and calls the difference a result.
- Cannot say what state the server was in when a past result was produced.
- Changes sampling settings or the tool list between cells without noticing.
- Assumes an effect found on the fixture necessarily exists on the real service.