skip to content

The AWS CDK writes a file called `cdk.context.json` in your project. What goes into it, and should it be committed to version control?

level: middleimportance: should knowfreq 45%

answer

  1. answers from AWS, frozen in a file
  2. so synth needs no credentials
  3. reproducible builds, reviewable changes
  4. the price is going stale
  5. reset one key, then re-synth

basics

~20 s

It caches the results of synthesis-time lookups against a real AWS account — VPC ids, availability zones, hosted zones, AMI ids, SSM values. Commit it, so synthesis is reproducible without AWS access; the cost is staleness until you reset an entry.

solid answer

~40 s

Some CDK calls query AWS while the template is being built — `ec2.Vpc.fromLookup`, `route53.HostedZone.fromLookup`, `ec2.MachineImage.lookup`, `ssm.StringParameter.valueFromLookup`, and the availability-zone list. The CLI makes those calls with your credentials on the first synth and writes the answers into `cdk.context.json`, keyed by provider plus account and region. Later synths read the cache instead of calling AWS, which is the point: a pull-request build can synth and diff deterministically without credentials for the target account, and two engineers get the same template. So yes, commit it. The tradeoff is staleness — if the looked-up VPC or AMI changes, you keep getting the cached value until you run `cdk context --reset <key>` (or `cdk context --clear`) and re-synth with credentials. Also remember it puts account ids and resource ids into the repository.

code

bash · 9 lines
bash
# Inspect the cached lookup results
cdk context

# Drop one stale entry by its listed number, then re-synth with credentials
cdk context --reset 3
cdk synth

# Nuclear option: re-query every provider on the next synth
cdk context --clear

go deeper

for a junior

Know that a few CDK calls named fromLookup or lookup query AWS while synthesizing, and that their answers are cached in cdk.context.json, which belongs in version control.

for a middle

Explain the caching contract: first synth queries with your credentials, later synths read the file, so builds are reproducible and credential-free — and describe how to reset a stale entry.

for a senior

Weigh synthesis-time against deploy-time resolution, and explain why a value frozen in a reviewed commit is often preferable to one that can change production with no code change.

for a principal

Decide the estate rule for which inputs are pinned in the repository versus resolved dynamically, and who is allowed to refresh context, so that infrastructure inputs cannot change outside review.

## Which calls are lookups Most of the CDK never talks to AWS. A handful of calls do, during synthesis, and they are all named in a way that hints at it: - `ec2.Vpc.fromLookup(...)` — finds an existing VPC and its subnets and AZs; - `route53.HostedZone.fromLookup(...)` — finds a hosted zone id by domain name; - `ec2.MachineImage.lookup(...)` — resolves an AMI id by filter; - `ssm.StringParameter.valueFromLookup(...)` — reads a parameter's value at synthesis; - the stack's availability-zone list. Each is a **context provider**. The CLI runs it, gets an answer, and the answer is baked into the template as a literal — a real `vpc-…` id, a real AMI id, a real string. ## Why the results are cached A CloudFormation template has to contain literal values, and the CDK's contract is that synthesis is a pure function from your code to a template. If lookups ran on every synth, that contract would break in two ways: you would need credentials for the target account every time you built (including in a PR job that has no business holding them), and the template could change underneath you because someone else created a subnet. So the first synth performs the missing lookups and records them in `cdk.context.json`, keyed by provider, account, region and the query arguments. Every later synth uses the cache and makes no AWS call. ```json { "vpc-provider:account=111122223333:filter.vpc-id=vpc-0abc123:region=eu-west-1": { "vpcId": "vpc-0abc123", "availabilityZones": ["eu-west-1a", "eu-west-1b"] } } ``` ## Commit it The file is meant to be in version control. The reasons are the reasons the cache exists: - **Reproducibility.** A given commit synthesizes to a given template on any machine. - **Credential-free builds.** CI can synth, run `cdk diff` against a template, and run policy checks without access to the target account. - **Reviewability.** When a looked-up value changes, it changes in a commit that a human approves, rather than silently between two builds. That last point is the important one and it is the answer to "but it goes stale". Staleness is not an accident of the design; it is the design. You are choosing to change infrastructure inputs deliberately. ## Refreshing an entry When the underlying resource really has changed — the VPC was replaced, a new AMI was published — the cache must be invalidated explicitly: - `cdk context` lists the cached entries with numbers; - `cdk context --reset <key-or-number>` drops one; - `cdk context --clear` drops all of them. Then re-synth with credentials for the right account, review the resulting diff in `cdk.context.json` and in the template, and commit both. The reason to prefer `--reset` over `--clear` is that clearing everything re-queries every provider and produces a large, hard-to-review diff. A common confusion: `cdk.context.json` is the *cache*, while context **values** you set yourself (`cdk.json`'s `context` block, `--context key=value`, `node.tryGetContext`) are inputs. Both live in the same context mechanism, which is why the file sometimes contains feature flags too — those you edit deliberately rather than resetting. ## Synthesis-time versus deploy-time values The alternative to a lookup is to defer the value to deploy time. SSM makes the contrast concrete: `ssm.StringParameter.valueFromLookup` reads the parameter during synthesis and inlines the string, whereas `ssm.StringParameter.valueForStringParameter` emits a CloudFormation dynamic reference that CloudFormation resolves during the deployment. The tradeoff is visibility versus freshness. A synthesis-time value is visible in the template and in `cdk diff`, so a review shows exactly what will be deployed; it goes stale. A deploy-time value is always current; it is opaque to `cdk diff`, and a change to the parameter can alter your infrastructure with no code change at all — which is either the flexibility you wanted or an unreviewed change to production, depending on the parameter. ## Gotchas worth naming Lookups need a stack with a concrete account and region, so an environment-agnostic stack cannot use them at all. If the CLI cannot perform a lookup — no credentials, wrong account — the synth produces obviously fake placeholder values so it can complete, and those templates must never be deployed. And the committed file contains account ids, VPC ids, AZ names and AMI ids, which is usually fine for a private repository but worth a moment's thought before open-sourcing a CDK app.

  • A colleague replaced the VPC and your CDK app still synthesizes the old vpc id. What do you do?
    The old id is cached. Run `cdk context` to find the entry, `cdk context --reset <key>` to drop it, then synth with credentials for that account so the provider re-queries. Review the diff in both `cdk.context.json` and the template — a VPC id change usually means resource replacement — and commit the updated cache with the change.
  • Why does the CDK resolve these values at synthesis rather than at deploy time?
    Because a CloudFormation template must contain literals, and because a value fixed at synthesis is visible in `cdk diff` and reviewable. Deploy-time resolution exists where it matters — `ssm.StringParameter.valueForStringParameter` emits a dynamic reference CloudFormation resolves — but then the value is invisible to diff and can change your infrastructure with no code change.
  • What does a synth produce when a lookup cannot be performed and nothing is cached?
    The CLI records the missing context and returns obviously fake placeholder values so synthesis completes rather than crashing. The resulting template contains dummy ids and must not be deployed; the fix is to run synth with credentials for the target account so the provider actually answers and the result is cached.
  • Is it ever right to leave `cdk.context.json` out of version control?
    Rarely, and it costs you: every build then needs credentials for the target account and can synthesize a different template than the last one. Some teams do it for a throwaway sandbox app. For anything promoted between environments, committing it is what makes a diff mean something.

saying these in an interview costs you the question

  • Calls it a local cache file that should be gitignored
  • Thinks lookups run at deploy time
  • Says stale context is a bug rather than the deliberate tradeoff
  • Suggests hand-editing the file instead of resetting the entry
  • Ignores that account and resource ids end up in the repository

context