Your team's tram fare-inspection sandbox account is shared by parallel CI runs — what breaks?
answer
- isolation you do not own
- the account is the scope
- records persist between runs
- throttling shared across workers
- partition, seed, cap the traffic
basics
~20 sShared sandbox state breaks isolation. Parallel runs see each other's records, share a throttle you cannot raise, and get no teardown at all. The provider owns the data and the outages, so partition per run and stub the rest.
solid answer
~50 sA provider's sandbox is **their** environment, so the isolation a local stub gives you for free is simply absent. Records your run creates persist for the next run and for everyone else on the account, so an assertion such as *"the newest inspection is mine"* becomes order-dependent. Throttling applies to the account rather than the worker, so adding parallelism produces failures that look like product bugs. There is no truncate-and-reseed teardown, because the data lives in a system you cannot administer. The workable shape is to **partition** — a distinct provider-issued identity or scoping value per run, where the provider offers it — **seed** everything your assertions read, never assert on global state, and move high-volume calls onto a stand-in you control, keeping on the sandbox only the interactions that genuinely need the provider's own logic.
go deeper
Be ready to say why a test that reads the newest record in a shared provider account passes on its own and fails in a parallel run. Name the shared, externally owned state as the cause.
Explain that both data and throttling scope to the account rather than to the worker, and that no administrative teardown exists, so isolation has to come from unique per-run references and seed-then-read assertions.
Show how you split the suite: which calls stay on the provider's environment, which move to a stand-in you control, and how a red sandbox stage gets triaged by a person rather than retried.
Own the trade across teams: how many test accounts you request from the provider, who negotiates extra scoping, and whether a dependency on somebody else's environment may ever block a release.
## The arrangement, not the interface A provider-operated sandbox is a third party's own system, running in test mode, on infrastructure they own. The operator of a tram fare-inspection API hosts it, seeds it, patches it and takes it down. Your team is issued credentials and an account inside it, and that account — not your CI worker, not your test class — is the unit that everything is scoped to. Every isolation property a stub you run yourself hands you for free becomes something you must negotiate, work around, or give up: - **Lifecycle**: you cannot start it, stop it, or roll it back to a known state. - **Population**: it holds whatever your team, and often the provider's own demonstration traffic, has left behind. - **Concurrency**: throttling is applied to the account, so your own workers compete with each other. - **Availability**: maintenance is announced to you, never scheduled by you. - **Behaviour**: the answers were written by the provider, so you cannot add a case to reach a branch they chose not to expose. That last point is why the sandbox is worth having, and it is also why it is expensive. You are buying somebody else's judgment about what a fare-inspection request means, at the cost of controlling nothing about the environment that judgment runs in. ## What actually breaks when runs go parallel These are not dramatic failures. They are the slow, flaky-suite failures a team can misdiagnose for weeks. - **Order-dependent assertions.** A check that reads "the most recent inspection record" passes when run alone and fails as soon as another worker writes after it. The assertion was never about your data; it was about global data that happened to be yours at the moment you wrote it. - **Reference collisions.** A sandbox commonly refuses a duplicate external reference within an account. Workers that derive a reference from a committed fixture therefore collide with each other, while each of them, taken in isolation, is perfectly correct. - **Throttle contention.** Adding workers shrinks every worker's share of the account's allowance. The suite begins returning throttling failures that look exactly like a defect in your retry code, and the fastest agent in the fleet becomes the flakiest. - **Teardown that cannot run.** You have no administrative access, so the truncate-and-reseed step a local database or a stub you own would allow simply does not exist. Data accumulates for the life of the account. - **Windows you did not choose.** The environment goes away when the provider decides it should, which reliably includes the evening a release candidate is waiting on a green build. ## Approximating isolation inside somebody else's system You cannot obtain isolation here; you can only approximate it, and the approximations are worth ranking. - **Partition by identity.** Where the provider issues distinct scoping values — separate test accounts, sub-accounts, or per-branch credentials — spend them. This is the only mechanism that genuinely separates runs, and it is worth asking the provider for by name. - **Make every run's data unique.** Derive references from a per-run value rather than from a committed literal. This removes collisions even when no partitioning is available. - **Seed what you read.** A test may assert only on records the same test created, fetched back by its own reference. Never assert on a listing, a total, or a "latest" — those are global, and they belong to everybody on the account. - **Serialise what cannot be partitioned.** A short critical section around a genuinely shared operation costs less wall-clock time than a flaky suite costs in attention. - **Cap the traffic.** Route to the sandbox only the interactions that need the provider's own logic, and stand in for the rest yourself. A suite that sends everything is not testing more; it is spending the account's allowance on repetition. ## Deciding what stays on it The useful question is not "sandbox or stub" but "which calls". Anything whose value is the provider's own validation, refusal or state machine belongs on the sandbox, and should be a small, slow, separately scheduled set. Anything whose value is your client's handling of an answer you already understand belongs on a stand-in you control, where it can be repeatable, fast, and as parallel as your hardware allows. Splitting the suite along that line usually removes most of the flakiness without losing real signal, because the flaky part was rarely the part that needed the provider. ## Operating the split Treat a sandbox failure as ambiguous evidence by default. Your build cannot tell your defect apart from their outage, their throttling, or a change they shipped overnight, so the pipeline needs somewhere to put that ambiguity: a clearly labelled job that may go red without blocking a release, a notification channel somebody actually watches, and a standing expectation that a red sandbox job is investigated rather than retried. The failure to avoid is a team that has learned to press retry, because at that point the stage costs time and attention and proves nothing at all.
- The provider offers no per-run scoping value. What is your fallback?Derive every reference from a per-run seed so runs cannot collide, assert only on records the same test created, serialise the genuinely shared operations, and cut sandbox traffic down to the interactions that need the provider's own logic. Then ask the provider for additional test accounts — that request is often granted, and it is the only real fix.
- How should a pipeline treat a red stage that calls the provider's sandbox?As ambiguous evidence. Put it in a clearly labelled job that may go red without blocking a release, route its notifications somewhere a person watches, and record the expectation that a failure is investigated rather than retried. A team that has learned to press retry has a stage that costs attention and proves nothing.
saying these in an interview costs you the question
- Claims a global teardown can clean the provider's account
- Asserts on the newest record in a shared account
- Blames the retry code when account throttling is the cause
- Thinks more parallel workers always shorten a sandbox suite
- Retries a red sandbox stage instead of triaging it