Most teams refuse to run a full apply-and-destroy integration test on every commit to their infrastructure repository. Why, and how would you decide what runs on a pull request versus nightly versus before a release?
answer
- minutes, money, quota, flakiness
- blocking gates must be deterministic
- cheap and often, expensive and scheduled
- rehearse the upgrade, not just the build
basics
~20 sA real apply costs minutes to hours, real money, and finite account quota, and it fails intermittently for provider reasons. So pull requests get fast, deterministic checks, a scheduled run pays for the expensive real deployment, and pre-release runs the full path once against realistic scale.
solid answer
~60 sFour forces push the expensive level off the commit path. Time: a real environment can take tens of minutes to build and as long to tear down, which no reviewer will wait for. Money: every run bills, and every failed teardown bills indefinitely. Contention and quota: parallel pull requests fight over a single test account's limits and over globally unique names. And flakiness: capacity errors, rate limits and slow readiness make the run intermittently red for reasons unrelated to the change, which destroys trust in the gate. So I place work by feedback value per minute. On a pull request: static checks, plan-level assertions on the intended diff, module tests against faked providers, and contract checks on shared interfaces — all deterministic and credential-light. Nightly: the real apply into a throwaway account, with a janitor sweeping orphans. Before a release: the full path at realistic scale, plus an upgrade rehearsal from the current production version. Everything blocking must be deterministic; anything intermittent gets an owner instead of a gate.
go deeper
Know that deploying real infrastructure takes minutes and costs money, so pull requests run the fast checks and the real deployment happens on a schedule instead.
Be able to place each level on a timeline and justify it: what runs on a pull request, what runs nightly, and why account quota and wall-clock time push the expensive level off the merge path.
Demonstrate that you have run this: triaging flaky provider failures, sweeping orphaned resources, change-scoped triggers, and insisting that anything blocking a merge is deterministic.
Own the strategy and the admission that goes with it — which assurance you are deliberately deferring, what a pool of test accounts costs to build, and when a testing level has stopped earning its keep and should be reduced or retired.
## The four forces Ask why not run everything every time and the answer is not one reason but four, and a strong answer names them separately because each has a different mitigation. **Wall-clock time.** Real infrastructure is slow: managed databases, clusters, certificates and DNS propagation are measured in tens of minutes, and teardown is often as slow as creation. A pull-request gate that takes forty minutes is not a gate; it is a queue, and people learn to merge around it. **Money.** Every run bills for what it creates, and the bill scales with how realistic the environment is. Worse, failed teardown bills *forever* until somebody notices. A dozen commits a day multiplied by a realistic environment is a line item somebody will eventually ask about. **Contention and quota.** A test account has finite limits: addresses, instances, network constructs, API rate. Ten concurrent pull requests each building an environment will exhaust something, and then the failures are about each other rather than about any change. Globally unique names collide. You can buy your way partly out with per-run accounts from a pool, but that is real platform investment, not a pipeline setting. **Non-determinism.** Capacity errors, throttling and eventual consistency make the run fail sometimes for reasons unrelated to the code. This is the force that matters most, because a blocking gate that is wrong one time in five teaches the team to re-run until green — at which point it blocks nothing while still costing everything. ## The placement principle Use one rule: **what blocks a merge must be fast and deterministic; what is slow or intermittent runs on a schedule with a named owner.** Everything else follows. **On the pull request** — target single-digit minutes, no credentials beyond read-only: - static checks: parse, schema, lint, secret and misconfiguration scanning - a generated plan, published where the reviewer can read it, with assertions on the diff: destructive actions, blast-radius size, mandatory tags, forbidden resource classes - module tests against faked providers for anything with real logic - contract checks on shared module interfaces **Nightly or on a schedule** — pays the money once per day: - build and destroy a real environment in a throwaway account - a small behavioural check that the thing actually works, not just that it was created - a janitor sweep for orphaned resources, and a drift check against real environments **Before a release** — the highest-fidelity, lowest-frequency run: - the full path at realistic scale and configuration - an upgrade rehearsal: deploy the version production currently runs, then apply the new one, so you test the *transition* and not just the greenfield build. This is the level most teams skip and the one that finds forced-replacement surprises. - a destroy rehearsal, because the delete path is code too ## Escape hatches worth naming - **Change-scoped triggers.** If only one component changed, deploy only that component's environment. Most commits touch a small area, and this turns an unaffordable run into an affordable one. - **A pool of accounts.** Hand each run its own account or project from a pool so contention and quota stop being shared problems, and teardown failure is bounded by recycling the account. - **Label-driven opt-in.** Let an author request the expensive run on a risky pull request without imposing it on every typo fix. - **Retry budgets by class.** Automatically retry the known-transient provider errors a bounded number of times; never retry a genuine assertion failure. ## Reporting, and the honest part Say out loud what you are accepting. Moving the real apply off the merge path means a change can merge and later fail at the top level. That is a deliberate trade: fast feedback and a trustworthy gate in exchange for a delayed, owned discovery of provider-level problems. Make it survivable — the nightly result alerts a person, the failure is triaged into regression versus provider noise, and a broken nightly blocks the next release rather than being carried indefinitely. ``` PR (minutes) static + plan assertions + module tests blocking nightly (an hour) real apply + behaviour + janitor owned, alerts release (hours) full scale + upgrade + destroy rehearsal blocks the release ``` ## What the interviewer is listening for Not the schedule — the reasoning. That flakiness, not cost, is the strongest argument against gating on the expensive level. That the upgrade path deserves its own rehearsal because greenfield builds hide replacement risk. And that any level you cannot afford to keep green should be reduced in scope rather than left permanently amber.
- What is the single strongest argument against gating merges on a real apply, cost aside?Non-determinism. Provider capacity errors, throttling and eventual consistency make the run red for reasons unrelated to the change, and a gate that is wrong often enough trains everyone to re-run until green. At that point it blocks nothing while still costing full price — worse than not having it, because it manufactures false confidence.
- Why rehearse an upgrade from the currently deployed version rather than only a greenfield build?Because production is never greenfield. Building from empty hides everything about the transition: resources that will be replaced rather than updated, in-place changes that require downtime, migrations of recorded resource identity, and delete-then-create ordering. Applying the new version on top of the old one is the only way to see the change production will actually experience.
- A team wants the full environment test on every pull request and has the budget. Would you agree?Only if it is genuinely deterministic and fast enough that nobody routes around it — which usually means per-run isolated accounts, change-scoped deployment, and a measured false-failure rate near zero. Budget solves cost and contention; it does not solve provider flakiness or forty-minute feedback. Measure the false-failure rate first and decide on that number.
- How do you decide when a whole testing level is not worth keeping?Look at what it has actually caught over a quarter versus what it cost in money and in engineer time triaging it. A level that has produced no unique finding — everything it caught was also caught cheaper below it — is paying rent for nothing. Reduce its scope to the parts that do break, or retire it deliberately and say so, rather than letting it sit permanently amber.
saying these in an interview costs you the question
- Believing more expensive tests everywhere always means safer
- Keeping a flaky gate blocking and re-running until green
- Testing only greenfield builds, never the upgrade path
- Ignoring test-account quota until pipelines start colliding
- Assuming budget alone justifies apply tests on every commit