skip to content

An AWS Lambda handler calls Secrets Manager on every invocation to fetch a database password. What problems does that create, and how would you fix it without hard-coding the secret?

level: seniorimportance: nice to knowfreq 40%

answer

  1. one call per request, at every request
  2. cost, latency, and a rate quota
  3. move it out of the handler
  4. reuse across warm invocations
  5. cache with a bound, not forever

basics

~20 s

Every invocation pays an extra API round trip, an API charge, and a share of the account's Secrets Manager rate quota, so a high-concurrency function throttles itself. Fetch the secret once per execution environment and cache it with a time-to-live.

solid answer

~50 s

Three costs, all of which scale with traffic: added latency on every request, a per-call charge, and consumption of the account's `GetSecretValue` rate quota — under heavy concurrency the function starts receiving `ThrottlingException` from its own credential lookup, which looks like a database outage. The fix is to move the fetch out of the handler into module scope so it runs once when the execution environment initialises, then reuse the cached value across the many invocations that environment serves. Cache with a TTL rather than forever, because a warm environment can live for hours and a rotated password would otherwise fail until it is recycled; a belt-and-braces version also refetches once on an authentication error before giving up. If you would rather not write that yourself, the AWS Parameters and Secrets Lambda Extension does it as a layer, serving cached values over `http://localhost:2773` with TTLs set through environment variables. Either way the execution role still needs `secretsmanager:GetSecretValue`, plus `kms:Decrypt` when the secret uses a customer managed key.

code

python · 18 lines
python
import os, time, json, boto3

_sm = boto3.client("secretsmanager")
_cache = {}
_TTL_SECONDS = 300

def _get_secret(secret_id):
    entry = _cache.get(secret_id)
    now = time.monotonic()
    if entry and now - entry[1] < _TTL_SECONDS:
        return entry[0]
    value = json.loads(_sm.get_secret_value(SecretId=secret_id)["SecretString"])
    _cache[secret_id] = (value, now)
    return value

def handler(event, context):
    creds = _get_secret(os.environ["DB_SECRET_ID"])
    return {"user": creds["username"]}

go deeper

for a junior

Know that fetching a secret on every invocation is wasteful and that the value should be fetched once and reused. Say that the secret still belongs in a secrets store, not in the code.

for a middle

Explain the mechanism — code outside the handler runs once per execution environment and its variables persist across invocations — and name the three costs: latency, per-call charges, and the API rate quota.

for a senior

Show the production judgment: throttling that scales with concurrency masquerading as a database outage, a bounded TTL plus invalidate-on-auth-failure for rotation, and the IAM plus network path the lookup depends on.

for a principal

Own the estate-wide answer — whether caching is a shared layer or per-team code, how rotation windows are coordinated so caches can turn over safely, and what the blast radius is when a single secret backs many functions.

## Why the naive version is a problem Fetching inside the handler means one extra network call, plus TLS and signing work, on the critical path of every single request. At low traffic it is a barely noticeable tens of milliseconds. It stops being harmless as volume rises, in three separate ways. **Latency.** The added round trip is pure overhead on a request that the caller is waiting on. For a function fronting an API, it is a fixed tax on every p50. **Cost.** Secrets Manager bills per API call as well as per stored secret. A function doing a million invocations a day is doing a million credential lookups a day for a value that changed last quarter. **Throttling — the one that actually pages you.** `GetSecretValue` is subject to an account-wide request rate quota. Concurrency multiplies the call rate directly: a thousand concurrent executions each fetching on entry produce a burst of a thousand near-simultaneous calls. When you cross the quota, Secrets Manager returns `ThrottlingException`, the handler cannot obtain a password, and the failure surfaces as "the database is down" even though the database never noticed. The tell is that the failure rate rises with traffic and disappears when traffic falls — a self-inflicted, load-correlated outage. ## The fix: fetch once per execution environment A Lambda execution environment serves many invocations over its lifetime. Code at module scope — outside the handler function — runs once, when the environment initialises, and anything it puts in a module-level variable survives into every subsequent invocation that environment handles. That is the natural place for a credential fetch. ```python import os, time, json, boto3 _sm = boto3.client("secretsmanager") _cache = {} _TTL_SECONDS = 300 def _get_secret(secret_id): entry = _cache.get(secret_id) now = time.monotonic() if entry and now - entry[1] < _TTL_SECONDS: return entry[0] value = json.loads(_sm.get_secret_value(SecretId=secret_id)["SecretString"]) _cache[secret_id] = (value, now) return value def handler(event, context): creds = _get_secret(os.environ["DB_SECRET_ID"]) ... ``` The call rate collapses from once per invocation to roughly once per environment per TTL, which is orders of magnitude smaller and, crucially, no longer proportional to traffic. ## Why the cache needs a TTL The obvious version — fetch once, keep forever — trades one problem for another. Warm environments can survive for a long time, so a function that cached at start-up may still be holding a password that was rotated an hour ago. Every invocation on that environment then fails to authenticate while a freshly created environment works fine, producing intermittent, environment-dependent failures that are miserable to reproduce. Two defences, ideally both: - **A TTL** measured in minutes. It bounds staleness to something you can reason about and still removes essentially all of the call volume. - **Invalidate on failure.** Catch the database's authentication error, clear the cache entry, refetch once, and retry. This handles rotation the moment it bites rather than waiting out the TTL. Guard it so a genuinely wrong credential cannot become a retry loop. If the secret is rotated on a schedule you control, you can also stage the change: Secrets Manager's rotation keeps the previous version usable for a window, so both old and new credentials authenticate while caches turn over. ## The managed alternative AWS publishes the **Parameters and Secrets Lambda Extension** as a layer. It runs alongside your function and exposes a small local HTTP cache, so your code makes a localhost call instead of an AWS API call: ```bash curl -s -H "X-Aws-Parameters-Secrets-Token: ${AWS_SESSION_TOKEN}" \ "http://localhost:2773/secretsmanager/get?secretId=prod/db/password" ``` It listens on port 2773 by default (`PARAMETERS_SECRETS_EXTENSION_HTTP_PORT`), and TTLs are configured with `SECRETS_MANAGER_TTL` and `SSM_PARAMETER_STORE_TTL`. The token header is required — the extension refuses unauthenticated local requests, which stops other code in the environment from trivially reading cached secrets. The trade is a layer and a little memory in exchange for not maintaining cache code in every runtime you use. Powertools for AWS Lambda offers an equivalent in-process utility for teams that prefer a library. ## Permissions, which people forget Whichever route you take, the execution role needs `secretsmanager:GetSecretValue` on the specific secret ARN — and `kms:Decrypt` on the key if the secret is encrypted with a customer managed key. The same applies to Parameter Store: `ssm:GetParameter` or `ssm:GetParameters`, plus `kms:Decrypt` for a `SecureString`. A missing `kms:Decrypt` is the classic confusing failure here, because the Secrets Manager permission looks correct and the error still says access denied. Also remember that if the function is attached to a VPC, reaching either service needs an egress path or a VPC endpoint. ## What a strong answer sounds like Name all three costs, with throttling as the one that turns into an incident; explain module scope versus handler scope as the mechanism; insist on a TTL and give the rotation reason for it; mention the extension as the managed option; and close on the two IAM actions plus the network path.

  • Why is caching the secret forever a bad idea?
    Because execution environments can live for hours. If the password rotates, an environment that cached at start-up keeps presenting the old one and fails to authenticate, while newly created environments succeed — an intermittent failure that tracks environment age rather than anything you can reproduce. A TTL of a few minutes bounds the staleness, and invalidating the cache on an authentication error handles rotation the moment it bites.
  • The role has secretsmanager:GetSecretValue but the call is still denied. What is missing?
    Almost always `kms:Decrypt`. A secret encrypted with a customer managed key requires both the Secrets Manager action and decrypt permission on that key, and the key policy has to allow the principal too. The same pairing applies to a Parameter Store `SecureString`. The error message points at the retrieval call, which sends people re-reading the wrong policy for a while.
  • When would you use the Parameters and Secrets Lambda Extension instead of caching in code?
    When you want one consistent caching behaviour across several runtimes and teams without maintaining that code in each. It is a layer that serves cached values over localhost with TTLs set by environment variables, so the caching policy becomes configuration rather than code. The costs are an extra layer, a little memory, and a component to keep up to date.

saying these in an interview costs you the question

  • Assumes Secrets Manager calls are free and instant
  • Ignores that concurrency multiplies the lookup rate into throttling
  • Caches the value forever so rotation breaks the function
  • Thinks module-level code runs on every invocation
  • Forgets kms:Decrypt when the secret uses a customer managed key

context