Your team runs Apache httpd with mpm_event in front of an application, and someone proposes replacing it with an event-driven server such as nginx 'for performance'. How would you decide, and where does a process-and-thread server genuinely lose to an event loop?
answer
- what is actually slow, and where
- connections versus active requests
- a stack per request, or a struct per connection
- the event MPM already closed most of the gap
- one in front of the other, not instead of
basics
~20 sDecide from the traffic's shape, not the reputation. A thread-per-request server loses when concurrent connections vastly outnumber active requests — each needs a thread and its stack, versus a few KB of state in an event loop. If mpm_event already keeps up, the bottleneck is usually the application.
solid answer
~50 sFirst establish what is actually slow, because 'Apache is slow' is usually an application or dependency finding wearing a web-server costume. Then compare the concurrency profiles: a thread-per-request server needs one thread, with its stack, for every request in flight, while an event-driven server keeps a few kilobytes of state per connection and multiplexes them over a handful of workers. The genuine losses are at very high connection counts of mostly-idle or slow clients, where thread stacks and scheduler pressure dominate, and in raw static-file and TLS-terminating throughput. But `mpm_event` already removed the historical gap — idle keep-alive connections no longer hold workers — so on a normal request mix the difference is often not what decides your latency. Against that, weigh what you would give up: per-directory override files, an enormous module ecosystem, and your team's operational familiarity. A common outcome is putting the event-driven server in front for connection absorption and TLS rather than replacing what works.
go deeper
Understand the core contrast: a thread or process per request in flight versus a small pool of workers multiplexing many connections in an event loop.
Be able to explain why per-connection memory differs by orders of magnitude, and that the event MPM removed the idle keep-alive part of Apache's historical disadvantage.
Show that you measure connections versus active requests and where latency accumulates before proposing a swap, and that slow clients and TLS volume are the concrete indicators.
Own the decision as a risk and cost trade — migration risk in rewrite and module configuration, team familiarity, and the layered option of fronting rather than replacing — with a measurement that justifies whichever you choose.
## Start by refusing the premise "Replace it for performance" is a conclusion, not a measurement. The first question is which number is unacceptable and where it is produced. In most stacks that reach this conversation, time-to-first-byte is dominated by the application and its dependencies, and the web server contributes a fraction of a millisecond. Swapping the web server in that situation changes the logo on the config file and nothing a user perceives. So: what is the concurrency, what is the request mix, where does the latency accumulate, and is any resource on the web tier actually near a limit? ## The architectural difference, stated honestly A thread-per-request server allocates a unit of execution per request in flight. Under `mpm_event` that is a thread with a stack — virtual address space measured in megabytes, resident memory typically far smaller but far from free — plus its share of scheduler attention. The kernel context-switches between them, and as the number of runnable threads grows, so does the switching and cache pressure. An event-driven server inverts this. A small number of worker processes each run a poll loop over many connections, and per-connection state is a data structure of a few kilobytes rather than a stack. Ten thousand mostly-idle connections is a memory question of tens of megabytes rather than ten thousand threads. This is the C10K argument, and it is real. The important qualifier: that argument is about *connections*, not *requests*. If your traffic is a modest number of connections each doing real work, the event loop's advantage largely disappears, because the work still has to happen somewhere and both servers hand it to the same application. ## Where the process/thread model genuinely loses - **Very high connection counts with low per-connection activity.** Long-lived connections, streaming, mobile clients with idle periods, or anything approaching tens of thousands of sockets. Thread stacks and scheduler pressure become the constraint before CPU does. - **Slow clients.** A worker occupied writing a response over a poor link is a worker not serving anyone. An event loop absorbs slow readers as cheaply as fast ones. This is the strongest single argument for putting an event-driven server at the edge. - **Static file and TLS throughput at high volume.** Event-driven servers were designed for exactly this path and generally do more of it per core. - **Sudden spikes.** Adding capacity means forking children and creating threads, bounded by the spare-thread window; an event loop simply admits more connections into its existing loop until it hits its own limits. ## Where the loss is smaller than the reputation suggests `mpm_event` closed the historical gap that made this comparison one-sided. Under prefork, an idle keep-alive connection held a whole process, so the memory-per-connection comparison was catastrophic. With event's listener thread absorbing idle connections, Apache pays for requests being *processed*, which is a much narrower and much more defensible cost. Benchmarks that circulate for this comparison are frequently run against prefork, or against a static-file workload that resembles nobody's application. And both servers are the same beyond the front door: if each request spends 80 ms in the application, no front end makes it 40. ## What you would be giving up - **Per-directory override files.** Apache lets configuration be dropped into a directory, which many shared-hosting and legacy applications assume. Event-driven servers deliberately have no equivalent; migrations often discover this late. - **The module ecosystem, and in-process modules generally.** Long-standing integrations, authentication modules and rewrite rule sets are Apache-specific assets. A rewrite corpus in particular does not port mechanically, and a subtly mistranslated rule is a security incident, not a performance regression. - **Operational familiarity.** Your on-call engineers know how to read Apache's logs and status page at three in the morning. That is worth real latency. ## How to decide Measure first: peak concurrent connections versus concurrent active requests, the ratio of static to dynamic, TLS handshakes per second, memory per worker times the worker ceiling against installed RAM, and whether workers wait on the app or the client. If connections vastly exceed active requests, or slow clients are eating workers, the event-driven model addresses your problem directly. If they are similar and the application dominates latency, you would be paying migration risk for a number nobody measures. Then consider the option that is usually best and rarely proposed: put the event-driven server in front, terminating TLS and absorbing slow and idle connections, and leave the existing Apache configuration serving the application behind it. You get the connection-handling win where it matters, keep the rewrite rules and modules that already work, and can retire the back tier later if it turns out to be worth it. Whole-fleet replacement is a large, error-prone project justified by a measurement — so take the measurement before committing to it.
- What single measurement most cleanly indicates that the event-driven model would help you?The ratio of concurrent open connections to concurrently active requests. When connections vastly exceed in-flight requests — many idle, slow or long-lived clients — you are paying thread-per-request costs for sockets doing nothing, and an event loop addresses that directly. When the two numbers are close, the work dominates and the concurrency model barely matters.
- Why is running the event-driven server in front of Apache often better than replacing it?It puts each server where its model wins. The front tier absorbs slow clients, idle connections and TLS at low cost per connection; the existing Apache keeps serving the application with its rewrite rules, modules and per-directory overrides intact. You capture the real benefit immediately and defer, or avoid, a risky configuration migration.
- Which part of an Apache configuration is riskiest to migrate, and why?The rewrite and access-control rules. They encode years of accumulated behaviour, they do not translate mechanically to another server's matching model, and a subtly wrong translation fails silently by exposing something or breaking a URL nobody tests. Unlike a performance regression, that class of mistake does not show up on a dashboard.
saying these in an interview costs you the question
- Asserts one server is simply faster without naming a workload
- Cites benchmarks run against mpm_prefork as evidence about mpm_event
- Ignores that the application usually dominates latency
- Treats rewrite and access rules as a mechanical port
- Frames it as replace-or-not, missing the front-tier option