skip to content

Statelessness & State Management

Function instances are ephemeral and unshared, so anything in memory or /tmp can vanish between requests. You will learn to externalise state into a database, cache, object store or workflow orchestrator, which is the design constraint serverless imposes.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a serverless function-as-a-service platform like AWS Lambda, why is it unsafe to cache a logged-in user's session data in a plain in-memory global variable inside the function's code?

level: juniorimportance: must knowfreq 78%

answer

  1. no sticky sessions
  2. in-memory = one instance only
  3. global var lost on cold start/new instance
  4. warm reuse is opportunistic caching only
  5. externalize to shared store

basics

~10 s

Because each request might land on a different, freshly-started copy of your function that never saw that variable, so the data would just be missing sometimes.

solid answer

~40 s

FaaS platforms run many independent instances of your function concurrently and can freeze, reuse, or kill any instance at any time. A global variable only lives inside one instance's memory, so a session stored there is invisible to every other instance, and even the instance that set it can be recycled between requests. Under load, the platform spins up dozens of instances, and each user request is routed to whichever one is free, not the one that 'remembers' them. This causes inconsistent behavior that appears to work in local testing (single instance) but breaks under real traffic. Session data has to live somewhere all instances can reach: a shared store like Redis or DynamoDB.

go deeper

for a junior

Should recognize that each function call might run somewhere completely different and give a basic reason (no shared memory) without needing to name specific platform internals.

for a middle

Should explain warm vs cold starts, that reuse is not guaranteed or controllable, and name at least one correct externalization option (e.g., a database or cache).

for a senior

Should distinguish safe-to-lose caching from correctness-critical state, explain why the platform intentionally avoids session affinity for scalability, and describe the production fix pattern (stateless tokens + shared store) with its own trade-offs.

for a principal

Should connect this to broader system design: how statelessness underpins the platform's ability to autoscale and multi-tenant efficiently, and how to design APIs/data models from the start so no component silently depends on process-lifetime memory, including cross-team guidance to prevent regressions.

## How the platform actually runs your code Serverless FaaS platforms like AWS Lambda, Azure Functions, and Google Cloud Functions do not run your code as a single long-lived process the way a traditional application server does. Instead, the platform provisions one or more **independent execution environments** — for Lambda these are lightweight micro-VMs built on Firecracker — each with its own isolated memory space, filesystem, and network namespace. - When your function is invoked, the platform routes that single request to one of these environments. - If ten requests arrive within the same second, the platform may spin up ten separate environments to handle them in parallel, because there is no guarantee any one environment can serve more than one request at a time. - A **global variable** declared at module scope lives entirely inside the memory of whichever single environment it was set in. It is invisible to every other environment, full stop — there is no shared address space, no network link, and no mechanism for one environment to read another's memory. ## Why the platform is built this way This model exists because the whole value proposition of serverless is elastic, near-instant horizontal scaling without the operator managing servers. The platform achieves that elasticity precisely by treating each execution environment as **disposable** and by refusing to guarantee **routing affinity** between a given client and a given environment. If the platform guaranteed 'requests from user X always go to instance Y,' it would need sticky-session infrastructure, reintroducing the scaling bottlenecks that serverless is designed to eliminate. So the platform load-balances every invocation independently, and it aggressively recycles environments: - after a period of inactivity an idle environment is frozen and eventually torn down; - under bursty load, the platform may also create and destroy environments rapidly to match demand. None of this is visible or controllable from your code. ## The trade-off The trade-off is that this statelessness is what makes serverless cheap and scalable — you pay only for actual compute time, and the platform can pack thousands of tenants' environments across its fleet — but it means your code must not depend on process-lifetime memory for anything that must be correct. There is a legitimate, narrower use of the same in-memory space: caching things that are safe to lose and merely expensive to recompute, such as a warmed-up database connection pool or a downloaded reference dataset. | What lives in that space | Why the distinction holds | |---|---| | A connection pool or a reference dataset | This is often called 'container reuse' or 'warm start' optimization, and it's fine specifically because losing it only costs a little latency, never correctness | | Session data, cart contents, or auth state | These are different: losing them produces a wrong answer, not just a slow one | ## The failure mode The concrete failure mode is intermittent and load-dependent, which makes it a nasty bug to catch. 1. In local development or low-traffic staging, the platform frequently reuses the same warm environment for consecutive requests, so a global-variable session appears to work perfectly. 2. The moment real traffic causes concurrent invocations, some fraction of requests land on a fresh, cold environment that never saw the session data, and the user gets logged out or their cart empties, apparently at random. 3. Worse, if an engineer tries to 'fix' this by caching per-user data more aggressively to survive warm reuse, they can accidentally leak one user's data to a different user who happens to hit the same warm environment afterward. ## The standard production fix The standard production fix is to externalize anything that must be durable and shared: - encode session state in a signed, stateless token such as a **JWT** that the client carries and the server merely verifies (common with API Gateway plus Lambda authorizers); - or write it to a shared store every instance can reach over the network, such as a `DynamoDB` table with a **TTL** attribute for expiry, or an ElastiCache Redis cluster for lower-latency reads. AWS's own serverless guidance is explicit that functions should treat local memory as a cache, never as a system of record.

  • If a warm Lambda instance can persist memory between invocations, why can't you just accept that some requests will 'miss' the cache and treat it as an acceptable trade-off?
    Because you cannot control or predict which requests will hit a warm instance versus a cold one, and the platform gives no guarantee of routing affinity, so the miss rate is not a fixed, tunable percentage — it depends on concurrency, traffic shape, and platform-internal scaling decisions. For anything where a miss means visibly wrong behavior rather than just slower, that unpredictability is unacceptable; it's fine only for pure performance caches where a miss just costs one recomputation.
  • Does using a stateless JWT for session data eliminate the need for a shared store entirely?
    It removes the need to look up session existence for many auth checks, since the token itself carries verified claims, but it doesn't eliminate shared state everywhere — you typically still need a shared revocation list or refresh-token store, and mutable data like cart contents still has to live in a database, so JWTs solve the identity/auth slice, not general application state.
  • How does container/instance reuse interact with a memory leak in a Lambda function?
    A leak that only grows within a single invocation is harmless because the environment usually only handles one invocation at a time and memory is capped, but a leak in module-level state that grows across invocations on the same warm instance can eventually hit the memory limit and cause repeated OOM crashes on that instance specifically, which looks like intermittent, hard-to-reproduce failures because only long-lived warm instances are affected.

Like assuming your order will always be handled by the same barista at a busy coffee shop with a dozen baristas working in parallel — the moment your name is called, any free barista serves you, and none of them share a memory of what you asked for last time unless it's written on a shared ticket.

saying these in an interview costs you the question

  • Claims a global/module-level variable is a valid session store because 'it worked in my testing'
  • Assumes Lambda instances are shared across all requests like a single running server
  • Doesn't distinguish between safe opportunistic caching (connection pools) and correctness-critical state (sessions)
  • Proposes sticky routing/session affinity as the fix instead of externalizing state
  • Unaware that warm instances can be reused across different users' requests

context

open as a page

AWS Lambda gives each function instance a writable /tmp directory (up to a configurable size, e.g. up to 10 GB). Why is writing files to this directory not a substitute for durable storage, even though multiple invocations sometimes reuse the same warm instance and can still see files written by an earlier invocation?

level: middleimportance: must knowfreq 70%

basics

~10 s

/tmp is just a scratch disk attached to one function instance; it disappears when that instance is recycled, and other instances never see it, so it can't reliably hold data you need to keep.

open as a page

A team implements a multi-step order-fulfillment process (validate payment, reserve inventory, ship, notify) as a single long-running Lambda function that keeps intermediate results in local variables across the steps, retrying the whole function from the top on any failure. What's wrong with this design, and how would a workflow orchestrator like AWS Step Functions change it?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Keeping progress in a running function's local variables means a crash or timeout loses everything and forces starting over, possibly repeating things like a payment charge; a workflow orchestrator instead saves each step's result externally so the process can resume exactly where it stopped.

open as a page

When a stateless function needs to remember something between separate invocations — e.g., a shopping cart or a rate-limit counter — what are the trade-offs between externalizing that state to a key-value store like DynamoDB or Redis versus an object store like S3?

level: middleimportance: should knowfreq 60%

basics

~10 s

DynamoDB/Redis are built for fast lookups of small structured records, while S3 is built for storing large files cheaply; picking the wrong one costs you either speed, correctness, or money.

open as a page

Two concurrent invocations of the same stateless Lambda function both read a counter value of 5 from a DynamoDB item, increment it locally to 6, and write 6 back. What went wrong, and what two DynamoDB mechanisms would prevent it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Both invocations read the same starting number before either wrote back, so one increment silently overwrites the other and the counter ends up one short. DynamoDB fixes this with atomic updates or conditional writes that check the value hasn't changed.

open as a page

A stateless serverless API scales out to 3,000 concurrent function instances during a traffic spike, and each instance opens its own connection to a relational database used as the externalized state store. What failure results, and what architectural patterns address it?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

3,000 function copies each opening their own database connection can overwhelm the database's max-connection limit, causing connection errors; the fix is a shared connection pooler between the functions and the database, or a database designed for many simultaneous connections.

open as a page