skip to content

A Lambda function calls GetSecretValue against AWS Secrets Manager at the top of every invocation. What goes wrong as traffic grows, and how do you fix it without pinning a stale credential forever?

level: seniorimportance: should knowfreq 40%

answer

  1. per-invocation network call and decrypt
  2. horizontal scale multiplies the call rate
  3. module scope survives warm invocations
  4. a layer with a localhost cache
  5. the clock is not the only freshness signal

basics

~20 s

Every invocation pays a network round trip and a KMS decrypt, the API calls are billed per request, and at high concurrency the account starts getting throttled. Cache the value per execution environment with a bounded lifetime, and refetch on an authentication failure so rotation still lands.

solid answer

~50 s

Fetching inside the handler means one Secrets Manager call plus a KMS decrypt on every single invocation. That adds latency to every request, it is billed per API call, and at concurrency in the thousands it will hit the account's request rate and start throttling — which surfaces as intermittent failures under load, not a clean error. The fix has two halves. **Cache**: move the fetch outside the handler so it runs once per execution environment and is reused across warm invocations, or use the **AWS Parameters and Secrets Lambda Extension**, a layer that runs a local HTTP cache on the sandbox and serves both Secrets Manager and Parameter Store reads with a configurable TTL. **Stay fresh**: bound the TTL, and — the part people skip — catch authentication failures, invalidate the cached value, refetch once and retry. Without that retry, a rotation turns every cached environment into an error source until its TTL expires.

code

python · 20 lines
python
import os, time, boto3

_sm = boto3.client("secretsmanager")
_TTL = 300
_cache = {"value": None, "at": 0.0}


def _secret(force=False):
    now = time.monotonic()
    if force or _cache["value"] is None or now - _cache["at"] > _TTL:
        _cache["value"] = _sm.get_secret_value(SecretId=os.environ["SECRET_ID"])["SecretString"]
        _cache["at"] = now
    return _cache["value"]


def handler(event, context):
    try:
        return query(_secret())
    except AuthenticationError:
        return query(_secret(force=True))

go deeper

for a junior

Know that a secret should be fetched once and reused rather than on every request, and that code outside the Lambda handler runs once per cold start while the handler body runs every time.

for a middle

Explain the concrete costs — added latency, per-call billing, account request limits — and describe module-scope caching with a TTL, plus what the AWS Parameters and Secrets Lambda Extension does instead.

for a senior

Demonstrate the freshness reasoning under rotation: bounded TTL combined with invalidate-and-retry on authentication failure, single-flight refetch to avoid a stampede, and why intermittent load-only failures point at a rate limit.

for a principal

Own the pattern across services: a shared retrieval library or the extension as the standard, a default TTL policy, and how rotation is exercised in staging so the refetch path is proven before a production rotation depends on it.

## What the naive version costs ```python def handler(event, context): secret = boto3.client("secretsmanager").get_secret_value(SecretId="prod/db") ... ``` This is wrong in four separate ways, and they compound as traffic grows. **Latency.** A `GetSecretValue` is a network round trip inside the Region plus a KMS decrypt. Tens of milliseconds added to *every* request, on the critical path, for a value that changes once a month. **Cost.** Secrets Manager bills per API call in addition to the monthly per-secret charge. At a few thousand requests per second the API charges quietly exceed the storage charge by orders of magnitude. The same is true for Parameter Store above the free standard throughput. **Throttling.** Both stores have account-level request rates. Lambda's whole model is to scale out horizontally, so a traffic spike multiplies the call rate by the concurrency — precisely the shape that trips a rate limit. The symptom is nasty: intermittent rate-exceeded errors under load only, which look like a downstream problem rather than a configuration one. **Client construction.** Building the SDK client inside the handler adds credential-chain resolution and connection setup on top. Clients belong at module scope regardless of caching. ## Fix one: cache outside the handler A Lambda execution environment is reused across invocations. Anything at module scope runs once per cold start and survives for the life of that environment, which may be many minutes and many thousands of invocations. ```python import boto3, time _sm = boto3.client("secretsmanager") _cache = {"value": None, "fetched_at": 0} TTL = 300 def _get_secret(): now = time.time() if _cache["value"] is None or now - _cache["fetched_at"] > TTL: _cache["value"] = _sm.get_secret_value(SecretId="prod/db")["SecretString"] _cache["fetched_at"] = now return _cache["value"] ``` That drops the call rate from once per invocation to roughly once per environment per TTL. The cost is staleness, bounded by the TTL. ## Fix two: the AWS Parameters and Secrets Lambda Extension AWS publishes a Lambda **extension** — attached as a layer — that does this for you and does it once per sandbox rather than once per language runtime process. It runs a small HTTP server on localhost (port 2773 by default, overridable with `PARAMETERS_SECRETS_EXTENSION_HTTP_PORT`) and serves both Secrets Manager and Parameter Store lookups from an in-memory cache. Your code makes a local HTTP call instead of an SDK call, passing the function's `AWS_SESSION_TOKEN` in the `X-Aws-Parameters-Secrets-Token` header so the extension knows the request came from inside the sandbox. TTLs are configured with environment variables — `SECRETS_MANAGER_TTL` and `SSM_PARAMETER_STORE_TTL`, both defaulting to 300 seconds — along with `PARAMETERS_SECRETS_EXTENSION_CACHE_SIZE` and `PARAMETERS_SECRETS_EXTENSION_CACHE_ENABLED`. The function's execution role still needs `secretsmanager:GetSecretValue` (or `ssm:GetParameter`) and `kms:Decrypt`; the extension uses the function's own credentials and grants nothing extra. What you get over hand-rolled caching: no per-language implementation, a cache shared by every process in the sandbox, and a TTL you can change with an environment variable instead of a deploy. ## The half everyone forgets: refetch on failure Caching and rotation are in direct tension. A TTL of 300 seconds means that after a rotation, up to five minutes of invocations in each warm environment present a credential the target no longer accepts. Shortening the TTL narrows the window but never closes it, and it gives back the savings you cached for. The correct pattern is TTL **plus** invalidate-on-auth-failure: 1. Use the cached credential. 2. If the operation fails with an authentication or authorization error from the target, drop the cached value. 3. Refetch once and retry the operation. 4. If it fails again, fail for real. That converts rotation from an outage into a single retried request, and it lets you run a generous TTL — because freshness is now driven by the failure signal, not only by the clock. Guard the refetch so a stampede of concurrent invocations does not all refetch at once, and cap it at one retry so a genuinely revoked credential does not turn into a retry loop against a throttled API. ## Beyond Lambda The same reasoning applies to containers and EC2, with one variation: ECS can inject a secret as an environment variable at task start, which is effectively a cache with a TTL of "until the task is replaced". That is fine for values rotated rarely and disastrous for values rotated hourly, and it is why long-lived services usually read through an in-process cache anyway.

  • How does caching interact with a secret that rotates every hour?
    A clock-based TTL alone forces a choice between staleness and call volume, and neither is good at hourly rotation. Keep a moderate TTL for the common case but make authentication failure the real invalidation trigger: drop the cached value, refetch once, retry. With alternating-users rotation the old credential also stays valid through the changeover, so most invocations never see a failure at all.
  • Why is the extension's cache better than caching in a module-level variable?
    It is per sandbox rather than per runtime process, so every process in the execution environment shares one cached copy and one upstream fetch. It is also language-independent — the same layer and the same environment variables work for every runtime — and the TTL is a configuration change rather than a code change.
  • What is the risk of caching a secret for the whole lifetime of the execution environment with no TTL at all?
    An environment can live for a long time under steady traffic, so a credential rotated an hour ago may still be in use. Without a TTL your only recovery path is the auth-failure retry, and if that is missing the function fails until Lambda happens to recycle the sandbox — which is not an event you control or can predict.

saying these in an interview costs you the question

  • Fetching the secret inside the handler because "it's only one call"
  • Caching forever and treating rotation as someone else's problem
  • Assuming Lambda automatically caches secret lookups
  • Solving throttling by adding retries with no cache
  • Shortening the TTL to seconds and calling that freshness

context