For a high-traffic API gateway on Laravel Octane, how do you choose between FrankenPHP, Swoole and RoadRunner, and how many workers do you start?
answer
- Swoole-only APIs decide first
- FrankenPHP: Caddy, HTTPS, admin port
- RoadRunner: rr binary plus RPC
- auto means different counts
- workers times memory under RAM
basics
~20 sSwoole or Open Swoole if you need Octane's tasks, ticks, cache or tables; otherwise FrankenPHP (one Go binary with built-in HTTPS) or RoadRunner (a Go process manager). Size --workers from measured memory and I/O wait, not cores alone.
solid answer
~40 sFirst ask whether the gateway needs Octane's Swoole-only APIs (concurrent tasks, ticks, the `octane` cache store, tables); if so the answer is Swoole or Open Swoole, a PHP extension you must build into the image. Otherwise compare operations: FrankenPHP is one Go binary on Caddy that can terminate TLS itself (`--https`, HTTP/2 and HTTP/3) and is reloaded through an admin port; RoadRunner is the `rr` binary plus two Composer packages, configured with `.rr.yaml` and reloaded over an RPC port. For workers, each worker handles one request at a time, so concurrency equals worker count. `--workers=auto` means CPU count (cgroup-aware) on Swoole, RoadRunner's logical-CPU default, and FrankenPHP's own default of twice the cores. An I/O-bound gateway usually wants more workers than cores, capped by memory: per-worker memory times workers must fit, measured under load.
code
bash · 9 lines# FrankenPHP facing the internet, terminating TLS itself
php artisan octane:start --server=frankenphp \
--host=0.0.0.0 --port=443 --https --http-redirect --workers=16
# Swoole behind a proxy, sized above the core count after load tests
php artisan octane:start --server=swoole --workers=16
# RoadRunner with its own config file and RPC port
php artisan octane:start --server=roadrunner --rr-config=.rr.yaml --rpc-port=6001go deeper
Recall the three drivers and that Swoole alone unlocks Octane's concurrent tasks, ticks, cache and tables.
Explain the operational differences: binary vs extension, TLS on FrankenPHP, the reload channels, and what --workers=auto resolves to on each server.
Show that you size workers from load tests and per-worker memory for an I/O-bound gateway and treat Swoole-only APIs as a lock-in decision.
Frame the choice as a platform decision: image build, TLS ownership, team familiarity and portability of code, not a single benchmark number.
## The scenario An **API gateway** built on Laravel mostly authenticates a request, calls one or more upstream services, reshapes the result and returns JSON. Most of its time is spent **waiting on I/O**, and its load is high and bursty. Running it on **Laravel Octane** removes the per-request framework boot, but two decisions remain: which server driver, and how many workers. ## Step 1: do you need a Swoole-only feature? Octane's docs mark four features as requiring Swoole (Open Swoole offers the same set): **concurrent tasks** (`Octane::concurrently`), **ticks and intervals**, the **`octane` cache store**, and **Swoole tables**. If the gateway's design depends on any of them - for example a shared in-memory table of rate counters across workers - the choice is made: Swoole or Open Swoole. That choice also ties the code to Swoole; switching servers later means rewriting those parts. (How those APIs work is a separate topic.) ## Step 2: compare the operational shapes | | FrankenPHP | Swoole / Open Swoole | RoadRunner | |---|---|---|---| | What you install | a Go binary (Caddy-based), downloaded by `octane:install` | a PHP extension you build or install yourself | the `rr` Go binary plus `spiral/roadrunner-http` and `spiral/roadrunner-cli` | | TLS | can terminate TLS itself: `--https` enables HTTPS, HTTP/2 and HTTP/3 with automatic certificates | usually behind a proxy | usually behind a proxy | | Server config | generated Caddyfile, or your own via `--caddyfile` | options merged from `octane.swoole.options` | `.rr.yaml` plus `-o` overrides Octane passes | | Reload channel | Caddy admin API (`--admin-port`; 2019 when serving on port 8000) | `SIGUSR1` to the master process | `rr reset` over the RPC port | | Swoole-only APIs | no | yes | no | Practical questions an interviewer expects: - **Extensions**: a static FrankenPHP build ships a fixed set of PHP extensions; Octane's docs point to the official Docker images when you need more. Swoole needs its extension compiled for your PHP version. - **Edge**: FrankenPHP can face the internet directly; the others are normally behind Nginx or a load balancer, which also serves static files. - **Operations**: each driver has its own reload and stop path, which your deploy tooling must match. ## Step 3: size the workers An Octane worker serves **one request at a time** (Octane even starts Swoole with `enable_coroutine` off). So the number of requests in flight is capped by the worker count; the rest wait in the server's queue. What `--workers=auto` means differs: 1. **Swoole**: Octane sets the count to the CPU count, reading the container's **cgroup CPU quota** first, so a container limited to 2 CPUs gets 2 workers. 2. **RoadRunner**: Octane passes `num_workers=0`, which RoadRunner treats as the number of logical CPUs. 3. **FrankenPHP**: Octane passes no count, and FrankenPHP applies its own default, which its docs give as twice the number of CPU cores. Octane's docs summarise `auto` as a worker per CPU core; on FrankenPHP the source leaves the choice to FrankenPHP. For a gateway that waits on upstreams, one worker per core leaves the CPU idle while every worker blocks on I/O. The usual approach: - start above the core count and **load-test**, watching latency, CPU and queueing; - bound the count by **memory**: each worker holds a booted application, so measured per-worker memory times the worker count, plus headroom, must fit the machine or container; - keep `max_execution_time` sensible so one slow upstream cannot pin workers indefinitely; - keep worker recycling (`--max-requests`, default 500) as the safety net for gradual memory growth. ## A sizing walk-through Illustrative numbers, not a recommendation: a container with 4 CPUs and 4 GB of memory, where a load test shows each worker settles near 120 MB after warm-up. Leaving about 1 GB for the operating system, the server process and spikes, memory allows roughly 25 workers. If latency and CPU stay healthy at 16 workers and queueing starts at 24, a production value around 16-20 leaves headroom. The same arithmetic, redone after each major dependency upgrade, keeps the count honest; per-worker memory tends to grow as the application does. ## What a strong answer avoids - Claiming one server is always the fastest; the right choice depends on the features, image and operations you need, and benchmarks should be your own. - Treating `auto` as the same number on every server. - Sizing only from CPU cores for an I/O-bound workload, or only from traffic without checking memory.
- Why can more workers than CPU cores make sense for a gateway?Each Octane worker handles one request at a time. A gateway spends most of each request waiting on upstream calls, so with one worker per core the CPU idles while all workers wait. Extra workers keep the CPU busy; the ceiling is memory, since each worker holds a booted app, and upstream capacity.
- What does Octane's Swoole auto worker count do inside a CPU-limited container?Octane reads the cgroup CPU quota (`cpu.max` on cgroup v2, the CFS quota files on v1) and rounds up, so a 2-CPU limit gives 2 workers even on a 16-core host. Only without a quota does it fall back to Swoole's CPU count.
- When would you put Nginx in front of FrankenPHP anyway?When the rest of the fleet already terminates TLS at a proxy or load balancer, or when you need proxy features your team runs in Nginx. FrankenPHP's built-in HTTPS is optional; Octane then binds it to a local port like the other drivers.
saying these in an interview costs you the question
- --https on octane:start terminates TLS on every Octane server
- --workers=auto starts exactly one worker per core on every server
- Octane refuses to start more workers than CPU cores
- Each Octane worker runs many requests concurrently as coroutines
- The worker count only needs CPU numbers, not memory measurements