A service keeps a large warm in-memory cache and holds long-lived connections to a database and a message broker. The team wants to move it to AWS Lambda. What architectural objections do you raise?
answer
- one cache becomes N cold caches
- isolated environments, no shared memory
- connections follow concurrency, not instances
- frozen after the handler returns
- this is a redesign, not a migration
basics
~20 sLambda gives each concurrent invocation its own isolated environment and freezes it once the handler returns, so a shared warm cache cannot exist, connection count scales with concurrency instead of instance count, and background work after the response stops running.
solid answer
~50 sThree objections, in order of severity. First, **the cache does not survive the move**: concurrent invocations run in separate execution environments with no shared memory, so a cache that was one warm copy becomes N cold copies, and each new environment starts empty. Hit rate collapses and the load you were shielding the database from lands on it. Second, **connections invert**: a fleet of long-lived processes holds a bounded pool, whereas function concurrency can fan out far wider, so connection count now tracks traffic and can exhaust the database — which is why a connection proxy exists as a mitigation. Third, **work after the response stops happening**: the execution environment is frozen when the handler returns, so background threads, timers and in-process schedulers do not run reliably. If those behaviours are load-bearing, the honest answer is that this is a long-lived process and belongs on containers or instances — or it needs redesigning around an external cache and event-driven triggers, which is a rewrite, not a migration.
go deeper
Know that each function invocation runs in its own isolated environment, so anything the process was holding in memory between requests does not carry across.
Explain why hit rate collapses across N environments and why connection count starts tracking concurrency rather than a pool size you chose.
Show you can classify the state as essential or incidental, name the mitigations and their costs, and recommend keeping a long-lived process when the state is genuinely load-bearing.
Own the framing that this is a redesign with its own budget and risk, and decide whether the organisation is buying a platform benefit worth a rewrite or moving for its own sake.
## Why "can it run" is the wrong question The code will almost certainly run. The objection is not feasibility, it is that three properties the service depends on are properties of a *long-lived process*, and the target model does not provide them. Recognising that this is a redesign rather than a lift-and-shift is the whole point of the question. ## Objection 1: an in-process cache stops being one cache A long-lived server holds one cache in one address space, shared by every request that instance serves. Under a per-invocation model, each concurrent execution environment is isolated with its own memory. Ten concurrent executions mean up to ten separate copies of that cache, each populated independently, and a newly created environment starts with nothing. Two consequences follow. The hit rate falls, sometimes dramatically, because each copy sees only the subset of traffic routed to it. And the traffic your cache was absorbing reappears at whatever is behind it — usually the database — precisely when concurrency is highest, which is the worst possible moment. The fix is to externalise the cache into a shared store. That is a real and common design, but notice what has happened: an in-memory lookup measured in nanoseconds became a network round trip measured in a fraction of a millisecond, plus a new dependency to operate and pay for. Whether that is acceptable is a per-service question. It is not free. ## Objection 2: connections scale with concurrency, not with instances A container fleet's connection count is bounded by design: instances × pool size, both numbers you chose. Under an elastic per-request model, concurrency is decided by traffic, and each environment that needs a connection opens one. Connection count now follows load. This is one of the classic production failures on this platform: a traffic spike produces a concurrency spike, which produces a connection spike, which exhausts the database's connection limit — and now every client of that database fails, not just the service that scaled. Connection *churn* also costs, because establishing a connection (handshake, authentication, TLS) is expensive relative to a short handler. Mitigations exist. A managed connection proxy in front of the database pools and multiplexes connections so the fan-out is absorbed outside your code. Reserving concurrency caps the fan-out at the source. Both are extra moving parts you did not previously need. A persistent broker subscription is worse than the database case, because it is not a mitigation problem — a function has no place to hold a subscription open. Consuming messages under this model means the platform polls the source and invokes you with a batch, which is a different consumption model with different ordering, retry and failure semantics. That is a rewrite of the consumer, not a configuration change. ## Objection 3: nothing runs after the handler returns A long-lived process routinely does work outside the request path: flushing metrics on a timer, refreshing a token in the background, running a periodic reconciliation, finishing an async task after the response is written. Under a per-invocation model the execution environment is frozen once the handler returns and may be resumed much later, or never. Background threads therefore run at unpredictable times or not at all — and this failure is *silent*, which makes it the nastiest of the three. Metrics simply go missing; a refresh simply does not happen. Everything of that kind has to become explicit: scheduled invocations for periodic work, a queue for deferred work, the platform's own mechanisms for flushing telemetry. ## What to recommend Be concrete about the fork. If the state is genuinely load-bearing — a large working set, tight latency budget, persistent connections, real background work — then this is a long-lived process, and containers give you the whole ladder's operational benefits without demanding the redesign. If the state is convenience rather than necessity — a modest cache, a pool that exists because the framework created one, a timer flushing metrics — then externalise it and the move is reasonable. And say what you would measure before deciding: current cache hit rate and what the database sees without it, peak concurrency implied by traffic and latency, and an inventory of every background task in the process. Those three numbers turn an argument into a decision.
- The team says environments get reused, so the cache will stay warm. What is wrong with relying on that?Reuse is an optimisation, not a guarantee. You cannot control how many environments exist, how long each lives, or which one serves a given request, so warmth is a probability rather than a property. Correctness and capacity planning must hold when every invocation starts cold.
- How would you keep the database from being overwhelmed if you did make the move?Put a managed connection proxy between the functions and the database so pooling happens outside your code, and cap fan-out by reserving concurrency for the function. Together they bound connection count and protect other clients of the same database from a traffic spike in this one service.
- Which of the three objections would you raise first with a sceptical team?The background-work one, because it fails silently. A missing cache shows up as latency and a connection spike shows up as errors, but a timer that stops firing produces no signal at all until someone notices metrics or refreshes went missing days later.
saying these in an interview costs you the question
- Assumes concurrent invocations share the same in-memory cache
- Thinks environment reuse guarantees a warm cache across requests
- Ignores that connection count now scales with traffic
- Expects background threads and timers to keep running after the response
- Calls it a lift-and-shift when the state model changes completely