skip to content

How do PHP's curl_multi_* functions run several HTTP requests concurrently, and how do you read each transfer's result and error?

level: seniorimportance: should knowfreq 32%

answer

  1. one thread, many transfers
  2. add easy handles to a multi handle
  3. curl_multi_exec with curl_multi_select loop
  4. curl_multi_info_read gives msg, result, handle
  5. curl_multi_getcontent needs RETURNTRANSFER

basics

~10 s

Add configured CurlHandles to a curl_multi_init() handle, then loop curl_multi_exec() and curl_multi_select() until nothing is running. curl_multi_info_read() reports each finished handle with its CURLE_* result; curl_multi_getcontent() returns its body when RETURNTRANSFER was set.

solid answer

~40 s

`curl_multi_*` drives several easy handles over non-blocking sockets in the same PHP thread. I configure each `CurlHandle` as usual, including `CURLOPT_RETURNTRANSFER` and timeouts, and `curl_multi_add_handle()` it to a `CurlMultiHandle`. Then I loop: `curl_multi_exec($mh, $running)` does whatever work is ready without blocking, and `curl_multi_select($mh)` sleeps until a socket is ready (its timeout defaults to 1.0 second). The loop ends when `$running` is 0 or the exec status is no longer `CURLM_OK`. `curl_multi_exec()`'s return value is the multi handle's own status, not a transfer result. Per-transfer outcomes come from `curl_multi_info_read()`, which returns `['msg' => CURLMSG_DONE, 'result' => CURLE_*, 'handle' => $ch]`. I read the body with `curl_multi_getcontent($ch)`, the status with `curl_getinfo()`, then `curl_multi_remove_handle()`. Total time is roughly the slowest call, not the sum.

code

php · 26 lines
php
<?php
declare(strict_types=1);

$mh = curl_multi_init();
$handles = [];
foreach (['carrier-a', 'carrier-b', 'carrier-c'] as $carrier) {
    $ch = curl_init("https://$carrier.example.test/rates?zip=10115");
    curl_setopt_array($ch, [CURLOPT_RETURNTRANSFER => true, CURLOPT_CONNECTTIMEOUT => 2, CURLOPT_TIMEOUT => 4]);
    curl_multi_add_handle($mh, $ch);
    $handles[spl_object_id($ch)] = $carrier;
}

do {
    $status = curl_multi_exec($mh, $running);
    if ($running) {
        curl_multi_select($mh);
    }
} while ($running && $status === CURLM_OK);

$quotes = [];
while (($info = curl_multi_info_read($mh)) !== false) {
    $ch = $info['handle'];
    $ok = $info['result'] === CURLE_OK && curl_getinfo($ch, CURLINFO_RESPONSE_CODE) === 200;
    $quotes[$handles[spl_object_id($ch)]] = $ok ? curl_multi_getcontent($ch) : null;
    curl_multi_remove_handle($mh, $ch);
}

go deeper

for a junior

Know that curl_multi lets several requests wait on the network at the same time, and that each request is still an ordinary CurlHandle.

for a middle

Walk through add_handle, the exec and select loop, info_read and getcontent, and say why each easy handle needs RETURNTRANSFER and its own timeouts.

for a senior

Separate the multi status from per-transfer results, cap connections with CURLMOPT_* options, avoid busy loops, and design partial-failure behaviour for the page.

for a principal

Decide when hand-rolled curl_multi is worth owning versus a promise-based client, weighing fan-out limits, observability and how failures of one upstream affect the whole response.

## Why a multi handle A single `curl_exec()` blocks until its transfer ends. Calling three quote APIs (three carriers for shipping rates, say) one after another costs the **sum** of their latencies. PHP's **multi interface** lets libcurl drive many transfers at once using non-blocking sockets in the same thread, so the wall time approaches the **slowest** call instead. It is concurrency, not parallel execution: PHP code does not run in several threads, libcurl just interleaves the network waits. ## The moving parts | Function | Role | |---|---| | `curl_multi_init(): CurlMultiHandle` | creates the multi handle (an object since PHP 8.0) | | `curl_multi_add_handle($mh, $ch): int` | attaches a configured easy handle; returns a `CURLM_*` code | | `curl_multi_exec($mh, &$still_running): int` | performs whatever work is ready, sets the count of running transfers | | `curl_multi_select($mh, float $timeout = 1.0): int` | waits until a socket is ready or the timeout passes; `-1` on failure | | `curl_multi_info_read($mh, &$queued_messages = null): array\|false` | returns the next completion message, or `false` when none remain | | `curl_multi_getcontent($ch): ?string` | the body of a handle that used `CURLOPT_RETURNTRANSFER`, else `null` | | `curl_multi_remove_handle($mh, $ch): int` | detaches a finished handle | PHP 8.5 also adds `curl_multi_get_handles($mh)`, which returns every `CurlHandle` currently attached. ## The driving loop 1. Create each `CurlHandle` with `CURLOPT_RETURNTRANSFER => true` and its **own** `CURLOPT_CONNECTTIMEOUT` and `CURLOPT_TIMEOUT`. One slow upstream must not hold the whole batch. 2. `curl_multi_add_handle()` each one, keeping a map from handle to meaning (which carrier it was). 3. Loop: call `curl_multi_exec($mh, $running)`. If transfers are still running, call `curl_multi_select($mh)` so the process sleeps until there is work instead of spinning the CPU. 4. Stop when `$running` reaches 0, or when the status returned by `curl_multi_exec()` is no longer `CURLM_OK`, which means the multi handle itself failed. 5. Drain completions with `curl_multi_info_read()` in a `while` loop. Each message has `msg` (`CURLMSG_DONE`), `result` (a `CURLE_*` code, `CURLE_OK` on success) and `handle` (the `CurlHandle`). 6. For each handle: check `result`, read `curl_multi_getcontent()` and `curl_getinfo($ch, CURLINFO_RESPONSE_CODE)`, then `curl_multi_remove_handle()`. ## Where the errors live This is the part people get wrong: - `curl_multi_exec()` returns a `CURLM_*` code about the **multi handle**. A `CURLM_OK` there says nothing about whether any transfer succeeded. - Each transfer's outcome is the `result` key of its `curl_multi_info_read()` message. In PHP's implementation, reading that message is also what records the error on the easy handle. After it, `curl_errno($ch)` and `curl_error($ch)` describe that transfer. Code that skips `curl_multi_info_read()` and calls `curl_errno()` directly can see `0` for a transfer that failed. - HTTP status is separate, exactly as with `curl_exec()`: a 500 is `CURLE_OK` at the transfer level. - `curl_multi_getcontent()` returns `null` if the handle did not use `CURLOPT_RETURNTRANSFER`, and `""` if it returned an empty body. ## Limits and production concerns - **Fan-out control.** Adding hundreds of handles opens hundreds of connections. `curl_multi_setopt()` with `CURLMOPT_MAX_TOTAL_CONNECTIONS` or `CURLMOPT_MAX_HOST_CONNECTIONS` caps them, or you can feed handles in batches. - **Busy loops.** Skipping `curl_multi_select()` makes the loop spin at 100 % CPU while waiting on the network. - **Partial results.** Decide up front what the page shows if one carrier fails: the other quotes, or nothing. - **When to reach for a library.** A client library wraps this loop in promises and pools; the mechanics underneath are the same multi interface. ## Measuring the gain The benefit is easy to check. Run the three carrier calls sequentially with `curl_exec()` and record `CURLINFO_TOTAL_TIME` for each; the page pays their sum. Run them through the multi handle and time the loop with `hrtime(true)`; the loop takes roughly as long as the slowest transfer, plus a little overhead. If one carrier regularly takes far longer than the others, the batch is only as fast as that carrier, which is why each easy handle needs its own `CURLOPT_TIMEOUT`. A carrier that times out then reports `CURLE_OPERATION_TIMEDOUT` in its `result`, and the page shows the quotes that did arrive. The gain also has a limit. The transfers share the same process, network interface and, if the calls go to one host, possibly the same upstream rate limit. Concurrency shortens waiting; it does not raise the upstream's capacity, and a burst of simultaneous calls can trigger throttling that sequential calls never hit.

  • Why can curl_errno($ch) return 0 for a transfer that failed inside a multi handle?
    In PHP's implementation, the easy handle's error is recorded when `curl_multi_info_read()` returns that handle's completion message. If the code never drains the messages, the handle's saved error stays unset, so `curl_errno()` reports 0. Read `result` from the message, or call `curl_errno()` only after `curl_multi_info_read()` has returned that handle.
  • What goes wrong if the loop calls curl_multi_exec() repeatedly without curl_multi_select()?
    `curl_multi_exec()` returns immediately when no socket is ready, so the loop spins, burning a full CPU core while it waits on the network. `curl_multi_select()` blocks until activity or its timeout (1.0 second by default), so the process sleeps instead of polling.
  • Does the multi interface make PHP code run in parallel threads?
    No. libcurl interleaves the transfers' network I/O over non-blocking sockets in the same thread; your PHP code still runs sequentially. You gain because the waits overlap: total time approaches the slowest request, not the sum. CPU-heavy work in PHP is not sped up at all.

A waiter who takes three tables' orders to the kitchen at once and serves each dish as it comes out, instead of standing at the pass until the first dish is ready before taking the second order. The kitchen cooks in parallel; the waiter just stops waiting in line, so the meal takes as long as the slowest dish.

saying these in an interview costs you the question

  • curl_multi_exec() returning CURLM_OK means every request succeeded.
  • curl_multi runs each request in its own PHP thread.
  • curl_multi_getcontent() works without CURLOPT_RETURNTRANSFER.
  • A loop without curl_multi_select() is fine because curl_multi_exec() blocks.
  • One shared timeout on the multi handle bounds each transfer.