In Laravel, how do a queue connection's retry_after and queue:work --timeout interact, and how can a misconfiguration make one transcode job run twice?
answer
- retry_after: 90 seconds per connection
- --timeout: 60 seconds per worker
- timeout must be several seconds shorter
- pcntl alarm kills the worker
- SQS uses its visibility timeout
basics
~20 sretry_after is how long a reserved job may run before the backend re-offers it; --timeout is how long the worker allows before killing itself. If a job outlives retry_after and --timeout is longer, a second worker runs it too.
solid answer
~50 s`retry_after` (90 seconds on the skeleton's database, beanstalkd and redis connections) is enforced by the **backend**: a job reserved longer than that is treated as available again, whatever the first worker is doing. `--timeout` (default 60) is enforced by the **worker**: with the `pcntl` extension it sets an alarm, and when a job exceeds it the worker marks the attempt, fires `JobTimedOut` and exits so Supervisor restarts it. The docs require `--timeout` to be several seconds shorter than `retry_after`, so a frozen worker dies before its job is handed out again. A 20-minute transcode on a connection with `retry_after` 90 and `--timeout=1800` gets picked up by a second worker after 90 seconds and encoded twice. The fix is a dedicated long-running connection — for example `retry_after` 1800 and `--timeout=1700` — plus a job-level `#[Timeout]` below `retry_after`.
code
php · 11 lines<?php
// config/queue.php (connections)
'redis-long' => [
'driver' => 'redis',
'connection' => env('REDIS_QUEUE_CONNECTION', 'default'),
'queue' => 'transcode',
'retry_after' => 1800,
'block_for' => null,
'after_commit' => false,
],go deeper
Recall that retry_after lives on the connection, --timeout on the worker, and that the timeout must be the shorter one.
Explain how each backend reserves a job and re-offers it after retry_after, and what the worker does when its alarm fires.
Diagnose duplicate runs of long jobs, set up a long-running connection and pool, and keep the job safe to repeat anyway.
Set queue timing conventions across teams so long-running work gets its own connections, pools and idempotency expectations.
## Two clocks, owned by different parties A queued job has two independent time limits, and they are enforced in different places. | | `retry_after` | `--timeout` | |---|---|---| | Where it is set | per connection in `config/queue.php` | per worker on `queue:work`, or per job | | Who enforces it | the queue backend | the worker process | | Default | 90 seconds (database, beanstalkd, redis) | 60 seconds | | What happens | the job becomes available to another worker | the worker kills itself | | Needs | nothing extra | the `pcntl` PHP extension | ## retry_after: the backend's lease When a worker pops a job, the backend records a reservation: - the **database** driver sets `reserved_at`, and its pop query also selects rows whose `reserved_at` is older than `retry_after` seconds; - the **redis** driver moves the job to a `:reserved` sorted set scored with now plus `retry_after`, and migrates expired entries back to the main list; - **SQS** has no `retry_after`; the message becomes visible again when its visibility timeout, configured in AWS, runs out. The backend cannot tell a slow worker from a dead one. When the lease expires, the job is simply handed to the next worker that asks. ## --timeout: the worker's alarm With `pcntl` loaded, the worker calls `pcntl_alarm()` before each job with the job's timeout — a `#[Timeout(seconds)]` attribute or `$timeout` property on the job wins over the `--timeout` flag. If the alarm fires: 1. the job is told about the signal and the timeout is recorded against it (and it is failed if it has no attempts left); 2. a `JobTimedOut` event is dispatched; 3. the worker process exits with an error status, and the process manager restarts it. Laravel 13.33 added `Worker::$killOnTimeout`: setting it to `false` makes the worker throw the timeout exception instead of exiting. Without `pcntl`, no alarm is set and `--timeout` has no effect; with `--once`, the docs note, it has no effect either. ## How a transcode runs twice The transcoding app puts `TranscodeVideo` on the `redis` connection with the skeleton's `retry_after` of 90, and runs workers with `--timeout=1800` because encodes are long: 1. worker A pops the job at t=0 and starts encoding; 2. at t=90 the reservation expires and the job returns to the queue; 3. worker B pops it and starts a second encode of the same video; 4. both write the same rendition, and the second run may count as a new attempt against the job's tries. With the ordering reversed — `--timeout` shorter than `retry_after` — a stuck encode is killed at the timeout, its worker restarts, and the job is retried only after `retry_after`, never concurrently. ## Configuring long-running jobs - Put long jobs on their own connection, such as a `redis-long` copy of `redis` with `retry_after` set above the longest expected run. - Run a dedicated worker pool for that queue with `--timeout` a margin below it, for example 1700 against 1800. - Put `#[Timeout(1700)]` on the job so the limit travels with the class, still below `retry_after`. - Keep Supervisor's `stopwaitsecs` above the longest job, or stopping workers kills encodes midway. - Give outbound calls their own timeouts; the docs warn that blocking I/O such as sockets may not honour the alarm. ## Choosing the numbers Start from the longest legitimate run of the job, measured rather than guessed — say encodes take up to 25 minutes (1500 seconds) at peak. - `#[Timeout(1700)]` on `TranscodeVideo` and `--timeout=1700` on its workers: a margin above the longest real run, so healthy encodes are never killed. - `retry_after` of 1800 on `redis-long`: a further margin above the timeout, so a frozen worker is killed and restarted before the job is re-offered. - Supervisor's `stopwaitsecs` of 1800 or more, so a stop waits for an encode to finish. - Keep the mail and default workers on the ordinary `redis` connection with `retry_after` 90 and `--timeout=60`, so a stuck mail job is retried within two minutes instead of half an hour. The cost of a long `retry_after` is slow recovery: if a transcode worker dies hard, its job waits up to 30 minutes before another worker picks it up. That is why long jobs get their own connection instead of stretching the defaults for every job. ## When a double run still happens Even with correct settings, a worker that dies hard — out of memory, a host reboot — leaves its job reserved until `retry_after` expires, and then another worker runs it again. That is at-least-once delivery. The transcode job should therefore be safe to repeat: write renditions to deterministic paths, and check whether the output already exists before encoding. The `WithoutOverlapping` job middleware and a job's failure policy are separate tools for the same risk.
- What happens to --timeout if the pcntl extension is missing?Nothing enforces it. The worker only registers its alarm handler when `pcntl` is loaded, so a frozen job keeps the worker busy indefinitely, and the backend hands the job to another worker once `retry_after` expires. The docs list PCNTL as a requirement for job timeouts.
- Why does an SQS connection in Laravel have no retry_after key?SQS enforces the lease itself: a received message stays invisible for the queue's visibility timeout, configured in AWS, and then reappears. Laravel's `--timeout` must therefore be shorter than that visibility timeout rather than a `retry_after` value.
saying these in an interview costs you the question
- --timeout should be longer than retry_after so slow jobs are never killed
- retry_after is how long the worker waits before retrying a failed job
- The worker enforces retry_after by killing jobs that exceed it
- --timeout works the same with or without the pcntl extension
- A job-level timeout is ignored when the worker passes --timeout