skip to content

Your stand-in tool server reproduces a hit in the harness every time, but the same attempt against the partner's real tool server produces nothing. Which differences between your stand-in and the real server would you check first?

level: seniorimportance: must knowfreq 48%

answer

  1. stand-in more permissive than real
  2. validation, truncation, escaping
  3. stale definitions vs current contract
  4. reflected error text is a surface
  5. rule out under-instrumented live run

basics

~20 s

Check where your copy is more permissive than the real service: looser schema validation, unsanitised and untruncated result bodies, stale description text, different field names, and errors that echo input back. Then diff your served tool definitions against the partner's current contract and a capture of real responses.

solid answer

~60 s

This is stand-in drift, and it is the standing cost of testing against a server you built. Work the diff in the order that most often explains a fake hit: - **Schema strictness.** Your copy accepts arguments the real service rejects — extra fields, wider enums, an optional parameter that is actually required — so the call that carries the effect never even executes upstream. - **Result body.** The real service escapes, sanitises, truncates or paginates what it returns; yours returns the payload whole. Length limits alone kill many reproductions. - **Definitions.** Your description text and parameter names are a snapshot; the partner shipped a change. - **Error text and metadata.** Yours may reflect input back into the transcript where theirs returns a bare code. - **Surrounding surface.** Tool count, ordering and the presence of neighbouring tools all change what the model does with the same content. The fix is to re-derive the stand-in from the partner's published contract plus a capture of real responses, and re-derive it on a cadence rather than once.

go deeper

for a junior

Notices the two servers are not the same thing and suggests comparing the tool definitions.

for a middle

Lists concrete divergence axes — validation strictness, result truncation and escaping, stale descriptions, error text — and checks them in order.

for a senior

Rules out an under-instrumented live run first, works the diff by likelihood, and decides whether the outcome is a retracted finding or a documented control.

for a principal

Institutionalises it: stand-ins generated from contracts, a periodic re-derivation step, and a rule that no report item ships on fixture-only evidence.

**Frame the situation before diffing anything.** You have a reliable result on your stand-in and nothing on the real server. There are two families of explanation, and they cost very different amounts to chase. Either the stand-in is more permissive than the real service, so the hit was an artefact — or the live attempt did hit and you could not see it. Rule out the second first, because it is an hour of work and the first is a day. On your own server a sensitive handler writes a server-side row the moment it is invoked; against the partner you usually have only the agent's own narration of what it did, which is a weaker detector and can silently miss a real invocation. Confirm the live call actually reached the tool, that the response you assume came back is the one that came back, and that your success criterion was observable at all in that environment. **Then diff the surface, most-likely-first.** 1. *Validation strictness.* Types, `enum` membership, the `required` list, `additionalProperties` handling, string length caps, server-side normalisation. This is the single most common source of a hit that does not transfer: your copy executes a call the real service rejects at the door, so nothing downstream ever runs. 2. *Returned content.* Encoding and escaping, truncation limits, pagination, and whether the field carrying the content is even the same field. A real service that caps a description field at 200 characters kills a large class of reproductions on length alone, with no security control involved. 3. *Definitions.* Diff your served tool definitions against the partner's current contract field by field — names, descriptions, defaults, ordering. Your copy is a snapshot; they shipped a release. 4. *Failure paths.* What the model sees on a rejected call. A stand-in written quickly tends to echo the offending input back in the error string; a production service tends to return a bare code. Reflected error text is its own surface, and it is one you invented. 5. *Environment.* Tool list size and ordering, the auth scope of your token (fields it cannot see at all), rate limiting, latency and retry behaviour. **What it costs.** Chasing this properly is a day or two of engineer time per service, and it is recurring: the diff has to be re-run every time the partner ships. The cheap version is to stop hand-writing the stand-in — generate it from the partner's published contract, keep a stored capture of real responses as a comparison fixture, and run the diff as a pre-engagement step rather than as a post-mortem. The expensive version is what you are doing now, which is discovering the drift after a finding has already been written up. **Where the number misleads.** The fixture produced a clean rate; production produced one non-reproduction. Those two numbers do not belong on the same axis. A single failed live attempt is not a zero percent rate — it is one observation with an unmeasured detector, and quoting it as 0 percent next to a fixture's 85 percent invents a precision neither side has. The inverse error is worse and more common: keeping the fixture rate in the report as the headline and demoting the live contradiction to "inconclusive" in an appendix. It is not inconclusive. It contradicts the claim, and a partner who finds that in your report will discount everything else in it. **Then decide what the finding is.** If the divergence is a real control on the partner's side — output sanitisation, a length cap, strict schema validation — that is a positive finding, not a wasted week. Write it up as a control that holds, name the exact mechanism, and note the agent behaviour underneath it that would resurface if the control were relaxed or bypassed, because that behaviour is genuine and it is the partner's residual risk. If the divergence is an artefact of your copy, retract the item, fix the stand-in, and re-derive it from the contract. What you may never do is keep both stories alive at once. **What you would check to close it out.** Re-run the attempt against the corrected stand-in and confirm it now fails the same way the live one did — a fixture that reproduces the real server's refusal for the real server's reason is the only evidence that you found the actual cause rather than a plausible one. Then add that case to the pre-engagement comparison so the same drift cannot come back unnoticed next quarter.

  • Before blaming drift, what do you rule out about the live attempt?
    That it was observable at all — did the call reach the tool, was the response what you assume, and could your success criterion be detected without the server-side logging your stand-in gave you.
  • The difference turns out to be the partner's output sanitisation. Is that a finding?
    Yes, a positive one: report it as a control that holds, name the mechanism, and note that the agent behaviour behind it is unchanged should the control ever be bypassed or removed.
  • How do you stop the stand-in drifting again next quarter?
    Generate it from the partner's contract rather than hand-writing it, keep a capture of real responses as a comparison fixture, and re-run the diff as a pre-engagement step.

Your stand-in is a key filed to match a photograph of the original. It opens your copy every time, and that tells you nothing about whether it opens the door.

saying these in an interview costs you the question

  • Keeps the finding in the report on fixture evidence and drops the live contradiction.
  • Assumes the live attempt failed without checking it was observable.
  • Cannot name a single concrete way a stand-in is typically more permissive.
  • Treats the partner's sanitisation as a nuisance rather than reporting it as a control.
  • Hand-patches the stand-in until it matches once, with no re-derivation plan.

context