You own hundreds of Terraform root modules, and refresh on every plan is now saturating provider API rate limits. How do you keep refresh and drift detection affordable?
answer
- one read per resource instance, per plan
- state size is the first lever, not the flag
- tier drift cadence by risk of the module
- stagger the cron; do not self-DoS
- never gate an apply on an unrefreshed plan
basics
~20 sTreat refresh cost as proportional to resources per state and runs per hour, then attack both: keep root modules small, tier how often each is fully refreshed, stagger schedules so runs do not collide, and reserve unrefreshed plans for cheap feedback rather than for gated applies.
solid answer
~40 sRefresh issues roughly one provider read per resource instance in state, so total API load is resources-per-state times runs-per-hour across the estate — and the first lever is state size, not flags. I split oversized root modules so each plan touches fewer objects, which also shrinks blast radius and lock contention. Then I tier the cadence: high-risk, frequently-clicked estate gets a nightly refreshed drift plan, stable low-churn modules get a weekly one, and everything is staggered rather than firing on the same cron minute. `-refresh=false` earns a place on cheap draft-PR feedback, but never on the plan that gates an apply, because that plan must be computed against reality. I also make throttling visible — retries and plan duration are metrics — so we tune before someone else's deploy pipeline starts failing.
go deeper
Know that a plan re-reads every resource Terraform manages in that state, so a very large root module means a slow plan and a lot of API calls.
Explain the arithmetic — reads scale with resource instances per state and with how often you plan — and why splitting a root module shortens refresh, locks and blast radius together.
Show you would tier and stagger drift runs by risk, keep unrefreshed plans off the gated pre-apply path, and watch plan duration and retry counts as the leading indicator of throttling.
Own the estate-wide policy: which tiers get monitored how often, how state sizing and account boundaries partition both blast radius and API quota, and the standing rule that any human-approved plan is a refreshed one.
## Model the cost before choosing a knob Refresh cost is not mysterious. Each plan reads every managed resource instance in that root module's state, roughly one provider read per instance, sometimes more for resources the provider assembles from several calls. So the estate's steady-state API load is approximately: resources_per_state x plans_per_state_per_hour x number_of_states Three multiplicands, three levers. People reach for the flag (`-refresh=false`) because it is the visible one, and it is the weakest of the three. ## Lever one: fewer resources per state A 1,500-resource root module pays 1,500 reads for a one-line change, every time anyone plans. Splitting it into focused states — per service, per environment, per lifecycle rate-of-change — cuts refresh time roughly proportionally, and pays other dividends: smaller blast radius on a bad apply, shorter lock windows, plans a reviewer can actually read. Split state is the structural fix, and it is the one worth doing even if the rate limiter never complained. The cost is coupling: split states must read each other's outputs, which is a separate design conversation. But refresh economics is one of the strongest arguments in favour of doing it. ## Lever two: fewer runs, better spread Not every module deserves the same attention. Tier the estate: - **High-risk, high-touch** — production networking, IAM, anything with console access and an on-call habit of clicking. Nightly refreshed drift plan. - **Stable, low-churn** — a data pipeline nobody touches. Weekly is plenty. - **Ephemeral** — preview environments that live for hours. Do not monitor drift at all; they are recreated, not repaired. Then stagger. Hundreds of jobs on the same cron minute is a self-inflicted denial of service against your own account. Spreading the same total work across the hour changes peak load without changing coverage, and it is usually the single cheapest fix available. ## Lever three: where an unrefreshed plan is honest `-refresh=false` is legitimate where the plan's job is fast feedback, not a safety gate: - Draft-PR plans that a human will re-run properly before merge. - A re-plan moments after a refreshed plan on the same module. - An emergency change where refresh itself is timing out under throttling. It is not legitimate on the plan that gates an apply. That plan's whole purpose is to say what will happen to the real world, and a diff computed against stale state cannot say that. If your gated plans are so slow that people want to skip refresh, that is lever one telling you something. ## Provider-side and process-side relief - **Credentials and quota partitioning.** Where a provider's limits are per-account or per-role, spreading modules across accounts spreads the limiter too — often the same boundary you already want for blast radius. - **Make throttling observable.** Plan wall-clock time and provider retry counts belong on a dashboard. Rate-limit pain shows up first as slow plans and only later as failures, so measure the leading indicator. - **Cap concurrency at the pipeline level.** Two dozen pipelines planning at once will find the limiter no matter how well each one behaves individually. - **Fewer resources, not just fewer reads.** Estates accumulate; modules that manage a thousand near-identical objects sometimes want a different abstraction rather than a bigger quota. ## The judgment an interviewer is listening for They want to hear that you do not treat drift detection as free and do not treat it as optional either. Blanket hourly refresh across everything is expensive and creates alert fatigue; turning refresh off to make the pain stop means you no longer know what production looks like. The principal-level answer is a *tiered* policy tied to risk, sized by state, spread over time, and measured — plus a clear rule that the plan a human approves before an apply is always a refreshed one. ## What to say out loud Cost is resources-per-state times runs-per-hour. Shrink states first because it fixes blast radius and lock contention too, tier and stagger the drift schedule by risk, allow unrefreshed plans only for throwaway feedback, partition across accounts where the limiter is per-account, and watch plan duration and retry counts so you tune before it becomes an incident.
- Why is splitting state a better answer to slow refresh than defaulting to -refresh=false?Because it removes the work rather than hiding it, and it pays elsewhere: smaller blast radius, shorter lock windows, and plans a reviewer can read end to end. `-refresh=false` leaves the same thousand resources in one state and simply stops looking at them, so the next full plan is just as slow and now potentially surprising.
- Which modules would you exclude from scheduled drift detection entirely?Short-lived environments that are recreated rather than repaired — preview and test stacks whose remedy for any problem is destroy and re-apply. Monitoring them generates alerts nobody acts on. Spend the API budget and the human attention on long-lived production estate where a console change can persist unnoticed for months.
- How do you tell whether provider rate limiting is actually hurting you yet?Watch plan wall-clock time and provider retry counts per module over time, not just failures. Throttling degrades gradually — retries lengthen plans long before anything errors — so a rising duration trend is the leading indicator. Failures are the point at which someone else's deploy pipeline has already started breaking.
- Is there a case for making the gated pre-apply plan skip refresh?No. That plan exists to state what will happen to the real world, and a diff against stale state cannot make that statement — a resource deleted or modified out of band simply will not appear. If the gated plan is too slow to tolerate, the answer is a smaller root module, not a blinder plan.
saying these in an interview costs you the question
- Turns off refresh globally to make the pipeline faster
- Runs hourly refreshed drift checks on every module regardless of risk
- Fires every scheduled plan on the same cron minute
- Treats refresh cost as fixed rather than proportional to state size
- Requests a quota increase without shrinking oversized root modules