What does a payment provider's own sandbox prove that a stub you write yourself cannot?
answer
- somebody else wrote the answers
- agreement by construction versus disagreement
- their validation, not yours
- still a separate deployment
- split the suite, schedule the sandbox part
basics
~20 sA provider's sandbox proves your request is acceptable to code you did not write: their validation, refusal rules and state transitions run for real. A stub only proves your client handles the answers you already imagined. Neither is production.
solid answer
~50 sA stub you author agrees with you by construction — it returns what you told it to return, to the requests you anticipated. That makes a suite repeatable, and it also caps what the suite can discover. A **provider's sandbox** runs the provider's own code: their validation, their required combinations, their signing and authentication, their sequencing rules, their refusal vocabulary. So it can refuse something you never thought to encode, which is the entire value. What it cannot give you is production: it is a separate deployment with its own data, configuration and often a different build, so a green run constrains release day without determining it. It also cannot produce a dropped connection, a truncated reply or a held-open response on demand, and repeating a call until every retry branch has run is exactly what the account's throttling exists to prevent.
go deeper
Be ready to say who wrote the replies in each case: you author a stub's answers, while a provider's sandbox runs the provider's own code against the request you send it.
Explain the evidence each gives you. A stub proves your client handles an answer you already described; the sandbox proves your request survives validation and sequencing you did not write and could not anticipate.
Demonstrate the split in practice — a small scheduled sandbox set answering whether the provider still accepts what you send, and a fast stand-in suite carrying the rest, with sandbox failures triaged by a person.
Own the position that realism is rationed. Argue what the organisation gains from a narrow, well-triaged dependency on a provider's environment, and what it loses by routing every test through it.
## What each arrangement actually is A **stub you write** is a process you start, populate and stop. You author its replies, so it agrees with you by construction: it returns what you told it to return, to the requests you anticipated, in the order you arranged. That is a feature — it is what makes a suite repeatable — and it is also the ceiling on what such a suite can discover. A **provider sandbox** is the third party's own system running in test mode. The provider wrote the code, the validation and the state transitions; you are a guest holding credentials they issued. Nothing in it agrees with you by construction, and that disagreement is the entire product. ## What the sandbox proves - **Your request is acceptable to code you did not write.** Required combinations, mutually exclusive options, encodings, signing and the exact shape of an authenticated call are all enforced by the provider. A stub you authored cannot enforce a rule you never knew existed. - **Their refusals are real refusals.** What comes back is what a real caller gets, phrased in their vocabulary, carrying whatever detail they actually include — often far less than the helpfully complete error body a developer writes into a stub. - **Their sequencing is real sequencing.** Where an operation must follow another, cannot be repeated, or settles asynchronously, the sandbox enforces that ordering. A stub enforces only the ordering you remembered to encode. - **Your credential handling works end to end.** The call is authenticated against a real issuer, so a broken signature, an expired-token path or a misconfigured client surfaces in a test rather than in production. - **Timing is approximately real.** Callbacks arrive when they arrive, over a real network, and the sandbox will not deliver them at your convenience. ## What it cannot prove - **It is not production.** It is a separate deployment with its own data, its own configuration, its own risk and eligibility settings, and often a different build of the provider's service. Passing there constrains release day; it does not determine it. - **A green run says nothing about tomorrow.** Nothing about passing this afternoon restrains what the provider ships tonight, and no commit in your repository will explain the change. - **It cannot reach what they did not expose.** Unusual outcomes are reachable only where the provider publishes a designated trigger for them. A partial outage, a malformed reply, or a connection dropped mid-body is not on the menu; you produce those yourself with a stand-in. - **It cannot exercise your failure handling under repetition.** Hammering an operation until every branch of your retry logic has run is precisely the usage the account's throttling exists to prevent. ## The comparison, laid out | Question | A stub you control | The provider's sandbox | |---|---|---| | Who wrote the answers? | You did | The provider did | | Can it refuse something you never anticipated? | No | Yes | | Is a run repeatable and self-contained? | Yes | No | | Can you reach an arbitrary edge case? | Yes, by authoring it | Only where a trigger is published | | Does a green result mean production is safe? | No | Still no | | Who can change its behaviour without telling you? | Nobody | The provider | ## How teams combine them The mature shape is a split rather than a choice: - The great majority of tests run against a stand-in you control: repeatable, fast, parallel, and free to model faults the provider will never produce for you. - A small, deliberately chosen set runs against the sandbox, usually on a schedule rather than on every commit, and answers a narrow question — does the provider still accept what we send, and do we still understand what they send back? - Failures in that scheduled set are triaged by a person, because a build cannot distinguish your defect from their change or their outage. ## The mistake worth naming The common failure is the team that adopts the sandbox and then routes everything to it, on the reasoning that more realism must always be better. What they get instead is a suite whose runtime is a network round trip per assertion, whose failures are ambiguous, whose parallelism is capped by somebody else's throttling, and which still cannot exercise the failure modes that actually take a service down. Realism at the edge is valuable precisely because it is rationed. Spread thinly across every test, it buys very little and costs the team its fast feedback loop — and, worse, it teaches everyone to distrust a red build, which is the habit that lets a genuine integration break sit unnoticed for a week.
- If the sandbox catches request-shape problems, why keep a stand-in at all?Because the sandbox cannot produce the failures that take services down — a dropped connection, a truncated reply, a response held past your timeout — and it cannot be hammered, because the account's throttling exists to stop exactly that. A stand-in gives you those branches cheaply, repeatably, and in parallel.
- What can you honestly assert about behaviour you did not configure?That the provider accepted or refused a particular request today, and that your client handled the answer it got. You cannot assert why they refused, that they will refuse the same way tomorrow, or that production applies the same rules — the sandbox is a separate deployment with its own configuration.
- How often should the suite that calls the sandbox run?On a schedule rather than on every commit, plus before a release that changes the integration. The question it answers — does the provider still accept what we send, and do we still understand the reply — changes on their timetable, not on yours, so tying it to your commit rate spends the account's allowance for nothing.
A stub is a dress rehearsal in which you also wrote the other actor's lines. A provider's sandbox is a rehearsal in their theatre with their cast, so you finally learn whether your cues land.
saying these in an interview costs you the question
- Says a green sandbox run means production will work
- Believes a stub can catch validation its author never knew
- Thinks the sandbox can produce arbitrary faults on request
- Routes every test to the sandbox for extra realism
- Assumes the sandbox runs the same build as production