skip to content

Different FaaS platforms manage instance concurrency differently - AWS Lambda by default routes exactly one invocation to an execution environment at a time, while Google Cloud Functions gen2 (built on Cloud Run) and Azure Functions can be configured to handle multiple concurrent invocations per instance. What are the practical implications of this difference for how you write handler code and size downstream connection pools?

level: principalimportance: should knowfreq 55%

answer

  1. Lambda default: 1 invocation per environment, isolated processes
  2. Cloud Run/Cloud Functions gen2 & Azure: concurrency setting, multiple requests per instance
  3. shared-instance model = thread-safety burden on your code
  4. isolated model = more DB connections needed at scale (RDS Proxy exists for this)
  5. trade-off: isolation/safety vs connection/resource efficiency

basics

~20 s

Some FaaS platforms give each running function instance exactly one request at a time, while others let a single instance handle several requests at once, similar to a small web server. That difference changes whether your code needs to worry about two requests running at the same time inside the same process, and how many database connections your system ends up opening.

solid answer

~50 s

AWS Lambda's default concurrency model is strictly one-invocation-per-environment: if 50 requests arrive at once, Lambda spins up (up to) 50 separate environments rather than routing multiple requests into one running instance, which means handler code never has to worry about thread-safety between concurrent invocations, but it also means connection pools are effectively per-environment (so N concurrent invocations can mean N separate database connections). Google Cloud Functions gen2/Cloud Run and Azure Functions (depending on hosting plan) allow a single instance to process multiple concurrent invocations - a setting like Cloud Run's concurrency setting controls how many requests one instance's process handles simultaneously - which means the platform needs fewer total instances (and can reuse a single connection pool across concurrent requests, which is more resource-efficient), but it puts the burden on the developer to write thread-safe/concurrency-safe handler code, since two invocations can now genuinely run at the same time in the same process.

go deeper

for a junior

Not expected to know this distinction in depth; a basic awareness that 'how many requests one running function instance can handle at once' varies by platform is enough.

for a middle

Should know that Lambda handles one request per environment by default and that this affects how many database connections a burst of traffic can generate.

for a senior

Should be able to explain both sides of the trade-off - isolation/safety versus resource efficiency - and know at least one concrete mitigation for each failure mode (e.g., RDS Proxy for Lambda connection fan-out, thread-safety review for shared-instance platforms).

for a principal

Should treat this as a first-class factor in choosing between FaaS providers or configuring concurrency settings for a given workload, reasoning about blast radius, downstream resource sizing at scale, and the operational risk of migrating code between platforms with different concurrency semantics without re-auditing for thread-safety.

## The design decision behind the difference FaaS platforms differ in a design decision that has outsized practical consequences: how many concurrent invocations a single running execution environment (instance) is allowed to handle at once. AWS Lambda's default and most common model is **single-concurrency-per-environment**: each environment processes exactly one invocation at a time, and when concurrent requests exceed the number of already-warm environments, Lambda's scheduler provisions additional environments (up to the account/function's concurrency limit) rather than routing a second concurrent request into an environment that's already busy. So if 100 requests hit a Lambda function simultaneously with only 5 warm environments available, Lambda cold-starts up to 95 more environments to cover the burst (subject to burst/scaling limits), and each of those 100 requests runs in complete isolation from every other one, in its own process, at the same moment. ## The concurrency-per-instance model Google Cloud Functions (2nd generation) is built on top of Cloud Run under the hood, and Cloud Run's model is fundamentally different: each instance is a container that can be configured, via a **concurrency setting**, to accept multiple simultaneous requests - up to 1000 per instance on Cloud Run, with a smaller default depending on the product. Azure Functions similarly allows concurrent execution within a single instance depending on the trigger type and hosting plan (the Consumption plan's HTTP trigger and various non-HTTP triggers process multiple invocations per instance by default, governed by settings like `maxConcurrentRequests` in `host.json`), and its Premium/Dedicated plans behave even more like traditional application servers. ## Thread-safety and shared mutable state This difference cascades into two very concrete practical implications for how you write code and size infrastructure. The first is thread-safety and shared mutable state. - **Under Lambda's one-at-a-time model**, your handler code never has to reason about two invocations executing concurrently inside the same process - even though module-scope state persists across invocations, it's only ever touched by one invocation's logic at any given instant, so a naive, non-thread-safe in-memory counter or cache is safe from race conditions (though still subject to the cross-invocation staleness issues that come with warm-state reuse). - **Under a concurrency-per-instance model** like Cloud Run-backed Cloud Functions or a multi-worker Azure Functions plan, two or more invocations can genuinely be executing your handler's code at the same moment inside the same process, meaning any shared mutable state - a global cache, a counter, a client object with internal mutable state that isn't documented as thread-safe - is a live race-condition risk that simply doesn't exist under Lambda's model, and you must actively write (or verify) thread-safe code, use language-level concurrency primitives (locks, concurrent collections), or explicitly avoid shared mutable state. ## Downstream resource sizing The second implication is downstream resource sizing, particularly database connection pools, and it points in the opposite direction. - **Under Lambda's model**, if you have N environments running concurrently, you effectively have N independent processes, each with (if you built one) its own connection pool - so a burst to 1,000 concurrent Lambda invocations, each holding even just 1-2 database connections, can slam a database with 1,000-2,000 simultaneous connections, frequently exceeding what a traditional relational database can handle (this is exactly why AWS built RDS Proxy - a connection-pooling proxy layer specifically to absorb this fan-out pattern from Lambda before it reaches the database). - **Under a concurrency-per-instance model**, far fewer total instances are needed to serve the same request volume (since each instance handles many requests concurrently), so a single shared connection pool per instance serves many concurrent requests, dramatically reducing the total connection count against the database for the same throughput - a genuine efficiency advantage of the shared-instance model, provided your pool is sized and your driver used correctly for concurrent access. ## The trade-off The trade-off, then, isn't that one model is strictly better: - **Lambda's isolation model** trades connection/resource efficiency for maximum safety and blast-radius containment (one invocation's bug, memory leak, or crash can never affect a concurrently-running sibling invocation, since they're in different processes). - **The shared-concurrency model** trades that isolation for resource efficiency and fewer cold starts under high fan-out (since fewer total instances are needed). ## A concrete real-world scenario A concrete real-world scenario: a team migrates an API from AWS Lambda to Google Cloud Functions gen2 for cost reasons and, without changing their handler code, starts seeing intermittent data corruption under load - because their handler had a module-scope, non-thread-safe cache object (safe under Lambda's one-at-a-time model) that was now being mutated by multiple concurrent invocations inside the same Cloud Run instance. The fix required either making that cache thread-safe (e.g., using a concurrent data structure with proper locking) or explicitly configuring the Cloud Run service's concurrency to 1 to replicate Lambda's isolation semantics, trading away the connection-pooling efficiency gain to restore correctness quickly while a proper thread-safe rewrite was scheduled.

  • Why did AWS build RDS Proxy specifically for Lambda-to-database workloads?
    Because Lambda's one-invocation-per-environment model means a burst of concurrent invocations can each open their own database connection, quickly exceeding a relational database's max-connections limit under high fan-out. RDS Proxy sits between Lambda and the database, pooling and multiplexing a large number of Lambda-side connection attempts down to a much smaller, stable pool of actual database connections.
  • If a team sets Cloud Run's concurrency setting to 1 for a function that was previously handling many concurrent requests per instance, what are they trading away and what do they gain?
    They give up the resource efficiency of sharing a connection pool and process across many concurrent requests per instance, meaning they'll need more total instances (and likely more cold starts and cost) to handle the same load. In exchange, they gain the same strong isolation guarantees Lambda provides by default - no risk of shared mutable state races between concurrent requests in the same process.

Like the difference between giving each customer their own private cashier who serves only them start to finish (Lambda's isolation - safe but you need a lot of cashiers and a lot of till drawers open at once) versus one cashier juggling several customers' orders concurrently at a shared register (concurrency-per-instance) - more efficient use of the register, but the cashier now has to be careful not to mix up orders.

saying these in an interview costs you the question

  • Assumes all FaaS platforms have identical concurrency-per-instance behavior
  • Writes handler code with shared mutable state without checking the platform's concurrency model first
  • Doesn't connect Lambda's isolation model to the database connection fan-out problem
  • Thinks a 'concurrency setting' just controls total scaling limits rather than per-instance request handling
  • No awareness that migrating between providers can silently introduce or remove thread-safety requirements

context