skip to content

Some infrastructure defects survive every cheap check and only surface when the configuration is genuinely applied into a live account. Which classes of failure are those, and why can no plan or mocked test predict them?

level: seniorimportance: should knowfreq 48%

answer

  1. the plan is a prediction, not a request
  2. the provider has not been asked yet
  3. quota, permissions, capacity, uniqueness
  4. things that fail on the way down too

basics

~20 s

Only a real apply exposes what the provider decides at request time: quota and limit refusals, missing permissions, argument combinations the schema allows but the API rejects, name collisions, region or capacity unavailability, slow eventual consistency, and dependency-ordered failures including destroys that hang or leave orphans.

solid answer

~50 s

A plan is a prediction built from a schema and from what the provider reported a moment ago; a mocked test is a prediction built from answers you invented. Neither has asked the API to do the work. So the failures that escape are precisely the ones the provider owns: quota exhaustion, an identity that lacks a permission, an argument combination that is legal in the schema but rejected by the service, a globally unique name already taken, an instance type unavailable in that zone today, and capacity errors. Then there is time: resources that take many minutes to become ready, health checks that flap, and eventual consistency where a freshly created identity is not yet visible to the next call. Finally there is teardown — deletes that fail because something else attached itself, or resources with protection that refuse to go. You only find these by creating them for real, which is why a throwaway account run exists at all despite its cost.

go deeper

for a junior

Know that a plan predicts changes but does not perform them, so failures like a missing permission or an exhausted quota only appear once the deployment actually runs.

for a middle

Be able to enumerate the categories — authorization, quota, semantic rejection beyond the schema, name collisions, capacity, eventual consistency — and explain why each is invisible to a schema-driven prediction.

for a senior

Show you have operated this: how you triage a flaky nightly, how you keep the test account's permissions honest, and how you handle partial applies and failed teardowns rather than only the happy path.

for a principal

Own the tradeoff between assurance and cost across the estate: which environments earn a real apply test at all, how much intermittent red the organisation will tolerate, and what budget and quota policy the test account needs to stay viable.

## The reason this gap exists at all Every cheap testing level works from a *model* of the provider. Static checks work from the schema. Plan-level assertions work from the schema plus a snapshot of what the provider said existed a few seconds ago. Faked-provider tests work from answers you wrote yourself. None of them has submitted a create request. A cloud API, however, enforces a great deal that lives nowhere in any schema — and that difference is the entire justification for maintaining an expensive integration level. ## The classes of failure **Authorization.** The identity running the deployment may be able to create nine resource types and not the tenth, or may create a resource but not attach the policy it needs. Permission evaluation happens at the API, against a policy that may itself have been changed by someone else this morning. No plan can know this; the plan was produced by a read-only credential in many pipelines, which is doubly blind. **Quota and limits.** Accounts have caps: addresses per region, instances of a family, rules per firewall group, buckets per account, keys per service. A plan happily shows you creating the resource that will be refused. This is the single most common surprise in a first real deployment to a new account, and it is the reason a throwaway account for tests needs its own quota headroom. **Semantic validation beyond the schema.** Schemas are coarse. A service may accept a field syntactically and reject the combination — this storage class with that replication mode, this database engine version with that instance family, this feature in this region. The schema says string; the API says no. **Uniqueness and collisions.** Some names are globally unique or unique per account. A plan sees no conflict because the conflicting object belongs to someone else, or was created after the last refresh. **Capacity and availability.** The requested machine type may simply not be available in that zone right now. This is not a configuration error at all; it is a runtime condition, and it makes integration tests intermittently red in a way no amount of code review prevents. **Time and eventual consistency.** Real resources take real time — clusters, databases and certificates can take tens of minutes. Worse, a resource can be reported as created while it is not yet visible to a subsequent call, so the next step fails with "not found" for something that certainly exists. Retry and readiness behaviour is only exercised by an actual apply. **Ordering and partial failure.** When step forty of sixty fails, you are left half-built. The system's ability to resume, and your ability to reason about a partially applied change, is only tested for real here. **Destroy-side failures.** Teardown is a first-class source of trouble and is frequently forgotten in interviews. Deletes fail because another resource attached itself outside your configuration, because a protection flag was set, because a dependent object was created by a controller rather than by you, or because deletion is asynchronous and times out. A test that creates fine and fails to destroy leaves paid resources behind. ## What this implies for how you run it Because these failures are real, the integration level has to be real: a separate account, project or subscription with its own quota, its own credentials, and a naming scheme that makes orphans identifiable. Because it is expensive and intermittently flaky through no fault of your code, it does not belong on every commit — it belongs on a schedule and before a release, with an owner who triages failures into "our bug", "provider capacity", and "teardown leak". And because teardown fails, you need a janitor: a scheduled sweep that finds resources tagged with a test run older than N hours and deletes them, plus a cost alarm on the test account. Teams that skip this discover the cost later, as a bill. ``` plan says: create 12, change 2, destroy 0 apply says: LimitExceeded on resource 9 of 12 left behind: 8 created, 1 failed, 3 never attempted ``` ## The interview signal Anyone can say "integration tests catch real problems". The senior answer names the categories — authorization, quota, semantic rejection, uniqueness, capacity, timing, partial failure, teardown — and then draws the operational conclusion: this level is valuable *and* intermittently red for reasons unrelated to the change, so it must be scheduled, owned and swept, not wired into every pull request.

  • Your nightly ephemeral-environment test is red one morning in three. How do you keep it useful?
    Triage by cause before touching the code. Split provider capacity and rate-limit errors from real regressions, retry the known-transient classes automatically with a bounded budget, and alert differently for the two. Track the ratio over time: if genuine regressions are rare and noise dominates, shrink the scope to the resources that actually break rather than letting the whole suite go amber permanently.
  • How do you stop failed teardowns from quietly costing money?
    Tag every resource a test creates with the run identifier and a creation timestamp, then run a scheduled janitor that deletes anything older than a few hours in the test account. Add a budget alarm on that account so a leak surfaces in hours, not on the monthly invoice, and treat a teardown failure as a build failure rather than a warning.
  • Does running the deployment identity with reduced permissions in a test account help or hurt?
    It helps if the test account's permissions mirror production's. The point of the exercise is to discover missing permissions before production does, so an over-privileged test identity destroys the signal — everything passes there and fails in the environment that matters. Mirror the production role, and let the test account be the place that surface the gap.

saying these in an interview costs you the question

  • Assuming a clean plan means the apply will succeed
  • Ignoring teardown failures because the test already passed
  • Running integration tests with admin credentials in the test account
  • Blaming code for every red run without triaging provider errors
  • Believing mocked provider tests cover permissions

context