skip to content

A large `terraform apply` keeps failing partway through with rate-limit errors from the cloud provider's API. What does Terraform's `-parallelism` option control, and how would you use it here?

level: seniorimportance: nice to knowfreq 35%

answer

  1. how many things at once
  2. the default is ten
  3. one run's graph walk, not the whole account
  4. lower it when the API pushes back
  5. raising it seldom makes anything faster

basics

~20 s

Terraform's -parallelism sets how many resource operations one run performs concurrently while walking the dependency graph, defaulting to 10. Lowering it thins the burst of API calls the run makes, which is the fastest way to get a throttled apply through.

solid answer

~40 s

`-parallelism=n` limits how many nodes Terraform works on at once as it walks the dependency graph; the default is 10. It is a per-run concurrency limit, not a per-provider or per-account one, and it changes nothing about the plan — only how fast the calls go out. When a provider API starts returning throttling errors, dropping to something like `-parallelism=3` for that run usually gets it to complete, at the cost of wall-clock time. I would pair that with the provider's own retry setting, such as the AWS provider's `max_retries` argument, so individual throttled calls back off instead of failing the run. And I would treat both as tactical: a root module big enough to saturate an API is usually a signal to split it into smaller states.

code

bash · 5 lines
bash
terraform apply -parallelism=3

# or, for every apply in a pipeline:
export TF_CLI_ARGS_apply="-parallelism=3"
terraform apply

go deeper

for a junior

Know that Terraform applies several resources at the same time rather than one by one, and that a command-line option caps how many, with a default of 10.

for a middle

Explain that the cap applies to one run's walk of the dependency graph, that it never changes ordering or the resulting infrastructure, and why lowering it reduces the burst of API calls.

for a senior

Diagnose the throttling: identify it in the error output, lower concurrency to get the run through, raise the provider's retry setting so individual calls back off, and say what the wall-clock cost will be.

for a principal

Frame it as a shared-quota problem. Concurrency limits are per run, so several pipelines still sum against one account; the durable answer is smaller root modules and staggered schedules, not a number everyone tunes locally.

## What the graph walk is doing After Terraform builds the dependency graph, apply is a concurrent walk over it: any node whose dependencies are already satisfied is eligible to run. Without a limit, a wide configuration — two hundred DNS records, fifty IAM policies, a big fan-out of subnets — would fire an enormous number of provider API calls at once. `-parallelism=n` caps how many of those operations are in flight simultaneously. The default is 10. ```bash terraform apply -parallelism=3 ``` ## What it does and does not affect It affects **concurrency only**. The graph is unchanged, the ordering constraints are unchanged, and the resulting infrastructure is identical: a run at `-parallelism=1` produces the same outcome as one at 10, just serialized and slower. It is also a *per-run* setting, not a global budget — it says nothing about other pipelines, other engineers, or other tools hitting the same API. Two runs at `-parallelism=5` are ten concurrent operations from the provider's point of view. It is likewise unrelated to state locking. Terraform prevents two applies from racing on the same state by holding a lock, which is a separate mechanism entirely; parallelism governs what happens *inside* one run. ## Why lowering it helps a throttled apply Cloud APIs enforce request rates, often per account, per region and per service. A wide apply is exactly the shape that trips those limits: many identical calls arriving in the same instant. Halving or quartering the concurrency spreads the same work over more seconds and typically drops the run under the threshold. It is the quickest lever because it needs no code change, no state surgery and no coordination — just re-running with a smaller number. The cost is time. If the graph is genuinely wide and the provider is fast, wall-clock time rises roughly in proportion. On a deep, chained graph the cost is far smaller, because dependencies were serializing most of the work anyway. ## Why raising it rarely helps The common misuse is bumping parallelism to speed up a slow apply. It usually does nothing, for two reasons. First, most slow applies are slow because of *dependency depth* or because individual resources take minutes to become ready — a database, a certificate, a load balancer — and more concurrency cannot shorten a chain or make a provisioning wait finish sooner. Second, pushing more calls at an API that is already near its limit converts a slow apply into a failing one. ## The complementary knobs - **Provider retries.** Providers implement their own backoff. The AWS provider, for instance, takes a `max_retries` argument in its provider block; raising it makes individual throttled calls retry rather than abort the whole run. Retries and parallelism solve different halves of the same problem — one absorbs the errors, the other stops causing them. - **`TF_CLI_ARGS_apply`.** If a repository must always run reduced, `TF_CLI_ARGS_apply="-parallelism=3"` in the pipeline environment applies the flag without editing every command line. - **Splitting the configuration.** The structural fix is fewer resources per root module. A state that manages so many objects that it saturates a provider API is also a state with a slow plan, a large blast radius and a long lock hold. ## What to say in an interview Name the default (10), say plainly that it is per-run concurrency over the graph walk and not a correctness knob, reach for lowering it as the incident-time lever, mention provider retries as the complement, and then say the structural thing: a configuration that reliably rate-limits the provider has outgrown its root module. The senior signal is not knowing the flag — it is knowing that the flag buys time and does not fix the shape of the estate.

  • Does lowering parallelism change the result of the apply in any way?
    No. It only limits how many operations run concurrently during the graph walk. Ordering constraints come from the dependency graph and are unaffected, so the resulting infrastructure and the recorded state are identical — the run simply takes longer. That is what makes it a safe lever to pull mid-incident.
  • Two pipelines run against different states in the same account, each at the default parallelism. Does the flag protect the shared API?
    No. Parallelism is scoped to a single run, so the provider sees the sum of both. State locking keeps runs off the same state but does nothing about a shared account quota. Protecting a shared API means lowering it in both pipelines, staggering their schedules, or reducing how much each root module manages.
  • Would raising parallelism speed up an apply that takes forty minutes?
    Usually not. Long applies are normally dominated by dependency depth and by resources that simply take minutes to become ready, neither of which extra concurrency shortens. It also risks turning a slow run into a throttled, failing one. Look at what the run is actually waiting on before touching the number.

saying these in an interview costs you the question

  • Thinks -parallelism controls how many Terraform runs execute at once
  • Raises parallelism as the standard fix for a slow apply
  • Believes lowering it can change apply order or the final result
  • Assumes the default is unlimited concurrency
  • Treats provider retry settings and parallelism as the same knob

context