Relational database engines serve concurrent client sessions with a process per connection, a thread per connection, or a pool of worker threads. Compare these architectures and their practical consequences.
answer
- Process = isolation, ~MBs each, hundreds of connections
- Thread = cheaper, shared address space, one crash kills all
- Worker pool = connections decoupled from concurrency
- Long transactions starve pooled workers
- Connections are not concurrency; cores bound throughput
basics
~20 sProcess per connection gives strong isolation but the highest per-session cost, so connection counts must stay low. Thread per connection is cheaper but shares one address space, so a crash or leak is riskier. Worker pools multiplex many sessions onto few threads, scaling to many connections at the cost of complexity and restrictions on session-bound work.
solid answer
~60 s**Process per connection** (classic PostgreSQL): each session is a separate OS process. Isolation is excellent - one backend crashing does not corrupt others' memory - and it simplifies the engine's memory model, but each process costs megabytes of private memory plus fork and context-switch overhead. A few hundred connections is a realistic ceiling, so external pooling is mandatory. **Thread per connection** (classic MySQL): each session is a thread in one process. Creation is cheaper and memory per session is smaller, so thousands of connections are viable, but everything shares one address space - a memory error is fatal process-wide - and shared structures need careful locking. **Worker/thread pool** (thread pool plugins, SQL Server's scheduler model): a bounded set of workers executes tasks from many sessions. Connection count decouples from concurrency, which prevents thrashing under overload, but a session that blocks a worker (long transaction, session-scoped temp state) can starve others, so the engine needs careful task scheduling. Practically, the model determines how expensive a connection is - and therefore how aggressively you must pool.
go deeper
Know that the three models exist and that a connection costs the server memory and an execution unit, so connection counts are limited.
Compare cost and isolation concretely and explain why the model drives the practical connection ceiling of the engine you use.
Reason about behaviour past the knee: thrashing versus queueing, per-session buffer multiplication, and what pins a worker in a multiplexed engine.
Treat it as a capacity and reliability trade-off - isolation and blast radius versus session density - and design the connection topology of the fleet around whichever model the engine uses.
## The problem being solved A database server must serve many concurrent sessions, each with its own authenticated identity, transaction, and workspace memory. How the engine maps a session onto an operating system execution unit is one of the oldest architectural choices in the field, and it leaks directly into how applications must be built. ## Process per connection The server's listener accepts a connection and forks a dedicated OS process to serve it. Each backend has its own address space and private memory for parsing, plan caching, sorting, and hashing; shared data such as the buffer pool lives in shared memory that all backends map. *Advantages*: hard isolation. A backend that crashes or corrupts its own heap is killed without taking the whole server down (the engine typically still recycles other sessions defensively after a crash, to be safe about shared memory). It also makes the engine's own code simpler - much of it can assume single-threaded execution. *Costs*: process creation is relatively expensive, and each session carries several megabytes of private memory that is never shared. Kernel scheduling and TLB pressure grow with process count. This is why such engines list connection limits in the hundreds and treat external pooling as a deployment requirement rather than an optimisation. ## Thread per connection One server process, one thread per session. Thread creation is cheaper than fork, and per-session memory is smaller because code, plan structures, and caches can be shared inside one address space. Engines usually add a thread cache so disconnect/reconnect cycles reuse threads instead of destroying them. *Advantages*: higher connection counts (thousands) with less memory per session; shared caches are trivially accessible. *Costs*: no isolation - a bug that corrupts memory in one session can take down the whole server. All shared state needs locks or atomics, so contention becomes the scaling limit. Per-session buffers (sort, join, temp) are still allocated per thread, so a high connection count with heavy queries can still exhaust memory. ## Worker / thread pool multiplexing Here the number of connections is deliberately decoupled from the number of executing threads. Sessions submit work items; a bounded set of workers picks them up. Some engines run their own user-mode scheduler that treats a query's blocking points as yield points. *Advantages*: overload becomes queueing rather than thrashing. Ten thousand mostly idle connections do not translate into ten thousand runnable threads competing for CPU. Throughput stays flat instead of collapsing past the knee of the curve. *Costs*: real complexity. Session-scoped state (open transactions, temp objects, session variables) must be tracked and rebound to whichever worker resumes the session, and any operation that blocks a worker for a long time - a long-running transaction, a lock wait, an external call - can starve queued sessions. Some engines therefore refuse to multiplex sessions with open transactions. ## Why this matters to application developers The model sets the price of a connection, which determines your architecture. On a process-per-connection engine, an application that opens hundreds of connections per instance is a production incident waiting to happen; pooling is not optional. On a thread-per-connection engine you have more headroom but the same qualitative limit, because per-session sort and join buffers still multiply. Under a worker pool the per-connection cost is lowest, but you inherit rules about what a session may do while multiplexed. A useful mental model: **connections are not concurrency.** Useful concurrency is bounded by CPU cores and storage parallelism regardless of the model; the architecture only decides how gracefully the server degrades when connections exceed that bound. ## Direction of travel Engines historically built on process per connection have been moving toward built-in multiplexing precisely because the operational burden of external pooling is high, while thread-per-connection engines have added optional thread pools for the same overload reasons. The trade-off never disappears - it is isolation and simplicity against density.
- Why does a process-per-connection engine effectively require an external connection pooler?Because each session costs a full OS process plus megabytes of private memory, so its practical connection ceiling is in the hundreds while a fleet of application instances can easily want thousands. A pooler multiplexes many client sessions onto a small set of server sessions, keeping the server inside its safe range. Without one, connection growth turns into memory pressure and context-switch thrashing rather than throughput.
- Under a worker-pool architecture, what kind of workload behaves badly?Anything that occupies a worker without making progress: long transactions, sessions blocked on locks, or applications that hold a session open across user think time. These pin workers and force other sessions to queue, so latency rises for everyone. Engines mitigate this by refusing to multiplex sessions with open transactions and by capping how long a task may run before yielding.
Process per connection is giving every guest their own private kitchen; thread per connection is one kitchen with many cooks; a worker pool is a few cooks working a ticket queue for a full dining room.
saying these in an interview costs you the question
- Claiming more connections always means more concurrency or more throughput
- Assuming thread-per-connection has no per-session memory cost
- Thinking a thread pool removes the need to bound connections at all
- Saying process-per-connection is simply obsolete, ignoring its isolation benefit