You are asked to red-team a staging copy of a client's LLM application instead of production. What differences between the copy and production would stop your results transferring, and what do you record about the environment you tested?
answer
- config match decides transfer
- system prompt, model id, routing
- corpus is synthetic in staging
- filters disabled to save cost
- signed delta list from client
basics
~20 sStaging results transfer only as far as the configuration matches. Check the system prompt, the model identifier and routing, the retrieval corpus, tool permissions, and whether input and output filters are enabled the same way. Record all of it with the test dates, and say in the report that findings apply to that configuration.
solid answer
~60 sA staging clone is a different agreement and often a different system. The parts that decide whether a finding reproduces in production are: - **The instruction layer** — system prompt and any prompt-template scaffolding. Staging frequently runs an older or stripped-down version. - **The model** — identifier, deployment region, routing between a cheap and an expensive model, sampling settings. - **The retrieval corpus** — staging often carries synthetic documents, so an injection planted in a real production document has no analogue. - **The tool surface** — which tools are registered and what credentials they hold. Staging tools are often stubbed. - **The moderation layer** — input and output classifiers and rail rules are the single most common thing disabled in staging to save cost. The fix is not to guess: capture the configuration as evidence, ideally a hash or a diff against production supplied by the client, and label every finding with the environment and date. Where the moderation layer differed, say so explicitly rather than letting a reader assume the finding survives it.
go deeper
Knows that staging and production can differ and that the report should say which environment was tested.
Enumerates the configuration elements that decide transferability — instruction layer, model identifier and routing, corpus, tool surface, moderation components — and captures them as evidence.
Negotiates a signed delta list against production, tags each finding with the components its reproduction depends on, and arranges a supervised production confirmation for the findings that matter.
Sets the firm's rule for when a staging-only engagement may be sold at all, and what the deliverable is allowed to claim when no production confirmation was permitted.
**Why "clone" is doing too much work.** Staging is a second deployment of the same application, provisioned to resemble production and diverging the moment anyone economises. For an LLM application the layers that decide behaviour are precisely the layers staging tends not to copy faithfully, because they are the expensive ones. A finding transfers to production only as far as each of these matches. | Layer | What it is | Usual staging deviation | What the deviation changes | |---|---|---|---| | Instruction layer | system prompt and any template scaffolding wrapped around user input | older or stripped-down copy | what the model will and will not agree to do | | Model | the identifier, the deployment or region, routing between a cheap and an expensive tier, and sampling settings such as temperature | a cheaper tier, different sampling | the success *rate* outright, not just the outcome | | Retrieval corpus | the index the application searches and where its documents come from | synthetic placeholder documents | removes the untrusted-content channel entirely | | Tool surface | the registered tools and the scopes on their credentials | stubs returning canned data | removes the consequence, and usually the chain that follows it | | Moderation | input classifier, output classifier, rail rules, hosted moderation service | switched off to save per-call cost | whether a successful attempt ever reaches a user | **The error runs in both directions, and one direction is invisible.** If staging has the moderation components disabled, every finding those components would have blocked is reported as live risk the product does not carry. The report *overstates*, the client burns remediation effort on something already mitigated, and when they cannot reproduce it, your credibility goes with it. That failure at least announces itself the first time somebody re-tests. The other direction never announces itself. Staging indexed with synthetic placeholder documents has no untrusted-content channel, so indirect injection through an ingested document cannot be found there. Staging with stubbed tools cannot show you the chain where a tool's output is fed back into the model as fresh instruction. You did not fail to find those things because they are absent; you failed because the environment removed the channel. And yet the deliverable reports a page count, a probe count, maybe a coverage percentage — and every one of those numbers is computed over the surface that *this environment* exposed. The denominator silently shrank. The number that misleads is not a severity; it is coverage. **What it costs.** The saving is real and it is why teams do this: a staging engagement avoids production spend, avoids polluting live safety telemetry, avoids waking an on-call rota, and is far easier to get signed. The offsetting costs are a configuration diff the client has to produce (hours of an engineer's time, and often genuine reluctance), plus a small supervised confirmation window against production for the findings that matter. Budget both at scoping. A staging-only engagement sold as a production assessment is cheaper for exactly the reason it is worth less. **What to capture, concretely.** For every environment you touch: base URL and deployment; model identifier, routing rules and sampling parameters; the system prompt — hashed with the date if the client will not let you store the text; the retrieval index name, its document count and the ingestion source; the registered tool list with the scopes each credential holds; every moderation component and whether it was enabled; and the calendar dates. Then ask the client either to sign that this matches production or to list the deltas. That signed delta list is the most valuable artefact a staging engagement produces, and it is the thing that converts "we tested a copy" into a defensible claim. **What goes in the deliverable.** One environment table, and a per-finding environment tag. For any finding whose reproduction depends on a component that differed, restate the dependency inside the finding itself — a reader who skims will not turn back to the table. For any channel the environment removed, say so in the coverage section, because that is where the coverage number needs its correction. **What you check.** Before testing: request the diff, and specifically ask whether the output classifier and rail rules are enabled here. During: confirm the model identifier the application actually calls, from a response header or the client's logs, rather than from the architecture diagram. After: for each accepted finding, ask which of the five layers it depends on, and whether that layer was verified rather than assumed.
- The client will not let you record the system prompt because it is considered proprietary. What do you do instead?Record a hash of it plus the date, and have the client confirm in writing that the same hash is deployed in production. That gives you a reproducible environment marker without storing the text, and it fails loudly if the prompt changes before re-test.
- You found a prompt-injection path through the retrieval corpus on staging. What is the minimum you need before claiming it in production?Evidence that production ingests documents from the same untrusted source and that the same retrieval and instruction handling applies. Ideally a single supervised confirmation attempt against production with a benign marker payload instead of a harmful one.
saying these in an interview costs you the question
- Assumes staging mirrors production without asking for or checking a configuration diff.
- Reports staging findings as production findings with no environment label.
- Never notices that the moderation layer was switched off in the tested environment.
- Treats stubbed tools as equivalent coverage of the real tool surface.
- Cannot name a single configuration element whose difference would invalidate a result.