skip to content

One way to handle a full work queue is to execute the rejected task on the thread that submitted it instead of queuing it. Explain why that creates backpressure, and what can go wrong with it in a real system.

level: seniorimportance: should knowfreq 38%

answer

  1. submitter busy running ⇒ submitter not submitting
  2. intake rate converges to drain rate, no tuning knob
  3. stalls the accept loop / event loop it belongs to
  4. self-deadlock if the inline task needs the same pool
  5. rejections stay 0 — count inline runs instead

basics

~20 s

While the submitter runs the task itself it cannot submit anything else, so the intake rate is throttled to the pool's drain rate automatically. The risks: whatever else that thread was responsible for stalls — accepting connections, an event loop, ordering guarantees — and if the task needs the same pool, it can deadlock.

solid answer

~60 s

**Why it throttles**: submission and execution become the same thread's work. A producer that spends T seconds executing a task is a producer that issues no new submissions for T seconds, so the offered rate falls to roughly the pool's completion rate. It is a feedback loop with no configuration, no error path, and no loss — degradation is gradual rather than a cliff. **What goes wrong**: - The submitter usually has *another* job. If it is an accept loop or event loop, it stops accepting connections or servicing other callbacks — you have converted a bounded pool problem into a whole-node stall, and upstream queues (kernel accept backlog, load balancer) fill instead. - Latency for the submitter's own path spikes to include a full task execution. - If the task blocks on results from that same pool, the submitter can deadlock itself. - Execution now happens on a thread with different context — different thread-locals, different priority, possibly different security context — and ordering assumptions break. Bounded blocking with a timeout is often the safer shape.

code

text · 10 lines
text
# throttling
producer loop:
  if !queue.offer(task): task.run()   # 50ms spent here = 50ms not submitting
=> offered rate falls toward pool drain rate automatically

# self-deadlock
queue full -> submitter runs taskA inline
taskA: submit(taskB); await(taskB.result)
taskB queued behind a full queue, or run inline by ... the blocked submitter
=> submitter waits for taskB forever

go deeper

for a junior

Explain the core idea: while the submitting thread runs the task itself, it cannot submit more work, so the input rate slows down to match what the pool can finish.

for a middle

Add the mechanism (intake rate converges to drain rate) and the main hazard: the submitting thread usually has another job, such as reading requests, which stops while it executes.

for a senior

Discuss the operational picture — backlog migrating to kernel/broker queues, tail latency, self-deadlock with dependent tasks, context and ordering changes, and the fact that rejection metrics stay at zero so you must count inline runs.

for a principal

Frame it as choosing where backpressure surfaces: inline execution pushes the queue upstream into infrastructure you observe less. Compare against fail-fast with client backoff and bounded blocking, and decide per thread role and per criticality class.

## The mechanism Run-on-submitter (often called a caller-runs policy) says: when the queue is full, the submitting thread executes the task inline rather than enqueueing it. Nothing is lost and nothing errors, yet the system slows down. Why? Because the producer's capacity to *produce* and its capacity to *consume* are now the same resource. Suppose tasks take 50 ms and a single producer thread submits in a loop. Normally it can offer thousands per second. Once the queue is full, each submission costs it 50 ms of execution, so it can offer at most 20 per second — and during each of those 50 ms windows the workers are draining the queue. The result is a self-regulating loop: intake rate converges toward drain rate with no tuning parameter at all. This is genuine backpressure, and it has three properties people like: - **No loss.** Every submitted task runs. - **No error path.** Callers do not need retry logic. - **Gradual degradation.** Throughput plateaus and latency rises smoothly instead of a sudden wall of rejections. ## Why it is still dangerous The policy quietly assumes the submitting thread has nothing better to do. In real systems it usually does. **The submitter is often the intake.** In a server, the thread submitting work is frequently the one accepting connections, reading from a socket, or polling a message broker. Occupy it for 50 ms and, for those 50 ms, no connections are accepted and no messages are polled. The backlog does not disappear — it moves to a place you control less: the kernel accept queue, the broker, the load balancer, or the client's connect timeout. You have converted an observable, bounded, in-process queue into an unobservable one somewhere upstream. **It is catastrophic on an event loop.** If the submitter is a single-threaded event loop, running a task inline blocks *every* connection that loop serves, not just this one. A policy meant to slow one producer stalls thousands of unrelated operations. **Latency for the submitter's own work explodes.** The submitting request now pays the full service time of someone else's task on top of its own. Tail latency is where this shows up: p50 looks fine, p99 doubles. **Deadlock.** If the inline task waits on a result produced by the same pool — a sub-task, a barrier, a dependent stage — the submitter is now holding itself hostage: it cannot finish until the pool progresses, and the pool cannot get the item it needs because the submitter is not submitting. Any dependency between pool tasks makes this reachable. **Context and ordering change.** The task runs on a thread with different ambient context: different thread-local state, tracing or request context, priority, possibly a different security or transaction context. Tasks that were guaranteed to execute on pool threads now sometimes do not — a class of "works 99% of the time" bugs. Ordering assumptions break too: an inline task may complete *before* items queued ahead of it, so a queue that appeared FIFO is not. **It can mask a capacity shortfall.** Because nothing errors, the rejection metric stays at zero. Unless you count inline executions separately, the system looks healthy while it is running in a degraded mode. ## Where it is a good fit - A **dedicated feeder thread** whose only responsibility is producing for this pool — a batch loader, an ETL reader, a file scanner. Slowing it is exactly the goal, and it has no other duties. - Work that is **independent** — no task in the pool waits on another — so self-deadlock is impossible. - Pipelines where **no loss is acceptable** and there is no useful place to send a rejection. ## Where to prefer something else - Request-serving threads, accept loops, event loops, or any thread with a latency SLA: prefer fail-fast plus client backoff, or blocking *with a timeout* so the stall is bounded and reportable. - Pools with internal task dependencies: never run inline; the deadlock is not hypothetical. ## Making it safe if you do use it 1. Count inline executions as a first-class metric — it is your real saturation signal, since rejections stay at zero. 2. Assert (in tests, or in code) that the submitter is never a pool worker of the same pool. 3. Bound the pain: prefer offer-with-timeout, then run inline only if the timeout expires, so short bursts still queue normally. 4. Verify context propagation: if the system relies on request or tracing context, make sure inline execution carries the same context that a pool thread would.

  • Your service uses run-on-submitter and, under load, connection timeouts spike while the pool's rejection count stays at zero. What is happening?
    The submitting thread is the one accepting or reading connections, and it is spending its time executing tasks inline. During those windows nothing is accepted, so the backlog accumulates in the kernel accept queue or at the load balancer and clients time out before being served. Rejections are zero by construction, which is why the pool looks healthy — the fix is to count inline executions and to move the policy off the intake thread.
  • How is run-on-submitter different from simply blocking the submitter until queue space appears?
    Both stall the producer and both propagate backpressure, but running inline does useful work during the stall and completes that task without ever queueing it, so throughput is slightly higher and nothing waits. Blocking, by contrast, can be given a timeout, which bounds the stall and lets you report or degrade — a property inline execution lacks, since a long task cannot be interrupted midway. On an event loop both are unsafe.
  • When is run-on-submitter clearly the right choice?
    When the submitter is a dedicated feeder thread with no other responsibility — a batch reader, ETL loader, or file scanner — and the pool's tasks are independent of one another. Slowing that thread is precisely the desired effect, there is no accept loop or SLA to protect, and with no inter-task dependencies self-deadlock is impossible.

A restaurant where, when the kitchen's ticket rail is full, the host has to go cook the next order. Orders do slow down — but while the host is at the stove nobody is seating guests, and the queue simply moves onto the sidewalk where you cannot see it.

saying these in an interview costs you the question

  • Calling it "free backpressure" without asking what else the submitting thread is responsible for.
  • Using it on an event loop or accept loop, stalling unrelated connections.
  • Assuming zero rejections means the system is healthy — inline executions are invisible unless counted.
  • Ignoring self-deadlock when pool tasks submit to and wait on the same pool.
  • Believing it preserves FIFO ordering — an inline task can finish before items queued ahead of it.

context