skip to content

Socket activation is often described as giving restarts that never refuse a connection. On a systemd host, what exactly keeps clients from seeing "connection refused" while the service restarts, and where does that property break down?

level: seniorimportance: should knowfreq 45%

answer

  1. PID 1 never restarts with your service
  2. refused becomes queued, not instant
  3. the pending queue has a floor and a ceiling
  4. stopping the service does not free the port
  5. crash loops can trip a rate limit

basics

~20 s

The listening socket is held by systemd, not by the service process, so restarting the service never closes it. Arriving connections sit queued on that socket and are served once the new process inherits the same descriptor and starts accepting.

solid answer

~50 s

Nothing about the listener changes during the restart, because the listener was never the daemon's. systemd opened it, and `systemctl restart myapp.service` only stops and starts the process behind it. Clients that connect in the gap complete their handshake against the still-open socket and wait in the kernel's pending-connection queue; the new process inherits the very same descriptor and drains them. Three things break the property. Anything already accepted by the old process is lost — there is no state transfer, so in-flight requests still need a graceful shutdown path in the application. The queue is finite, so a restart that takes long enough under enough load starts dropping arrivals. And `systemctl restart myapp.socket` or `stop myapp.socket` genuinely closes the listener, which is the command you use when you actually want the port to go away.

code

bash · 9 lines
bash
# the process cycles; the listening socket in PID 1 never closes
systemctl restart myapp.service

# what actually removes the endpoint
systemctl stop myapp.socket
systemctl stop myapp.service

# confirm which socket unit triggers the service
systemctl status myapp.service

go deeper

for a junior

Know the one-line reason: the socket belongs to systemd, so restarting the service does not close the port. That is enough at this level; leave the limits to the follow-up.

for a middle

Explain the mechanics — the descriptor stays open in PID 1, arriving connections wait in the kernel's pending queue, and the new process inherits the same descriptor and accepts them. Note that already-accepted connections are not covered.

for a senior

Show where it breaks in production: finite queue depth, client timeouts outlasting a slow start, the need for a SIGTERM drain, and the fact that stopping the service leaves the port live while stopping the socket is what takes it down. Recognise the activation rate limit as a cause of a vanished port.

for a principal

Position it correctly in a deployment strategy — it removes the refused-connection window on a single host and costs nothing, but it does not do version skew, health gating or cross-host traffic shifting. Be explicit about which half of the availability problem you are still buying elsewhere.

## Why the connection is not refused "Connection refused" on TCP means the host answered a SYN with a RST, which is what the kernel does when nothing is listening on that port. Under socket activation there is always something listening: systemd created the socket, and systemd is PID 1, which does not restart when your service does. `systemctl restart myapp.service` tears down and re-spawns the process, and the descriptor stays open in PID 1's file table the whole time. So a client connecting during the gap completes its handshake normally and its connection sits in the kernel's queue of established-but-not-yet-accepted connections. When the new process starts, it is handed the same descriptor as fd 3, calls `accept()`, and picks those connections straight up. From the client's point of view the request was slow, not failed — and "slow" is a far better failure mode than "refused", because clients retry storms and circuit breakers usually trigger on refusal. ## What this does not give you It is worth being precise, because the phrase "zero-downtime" oversells it. **Accepted connections die.** Anything the old process had already accepted belongs to that process. When it exits, those sockets close. Long-lived connections, in-flight HTTP requests and open WebSockets are all dropped unless the application handles `SIGTERM` by finishing what it has in hand before exiting. Socket activation protects the arrival path; the application still owns the drain path. **The queue is finite.** The pending queue on the listening socket has a limited depth, which `Backlog=` in the socket unit sets. A restart that takes two seconds under a trickle of traffic is invisible; the same restart under heavy arrival rate overflows the queue, and once it is full new arrivals are dropped rather than queued. Socket activation converts a certain refusal into a bounded wait, not into an infinite buffer. **Nothing is buffered above the socket.** systemd does not proxy, read or replay any bytes. It only holds the descriptor. If your client sends a request and gives up after 500 ms, the fact that the connection was accepted does not help you. ## The commands that behave differently than people expect This is the operational half of the topic, and it is where interviews go: - `systemctl restart myapp.service` — the socket stays open, connections queue. This is the case the feature exists for. - `systemctl stop myapp.service` — stops the process, but the socket is still listening, so the very next connection starts the service again. Operators who meant "take this out of service" are surprised to find it back a second later. - `systemctl stop myapp.socket` — this is what actually closes the listener. To take the endpoint down properly you stop the socket and then the service. - `systemctl restart myapp.socket` — closes and reopens the listening socket, which drops the queued connections. Not what you want during a rolling restart. ```bash systemctl restart myapp.service # process cycles, listener persists systemctl stop myapp.socket # the endpoint really goes away systemctl status myapp.service # look for the TriggeredBy: line ``` ## The crash-loop guard There is a second failure mode worth knowing. If the activated service keeps dying immediately, every queued connection triggers another start attempt, and you get a very fast activation loop. Socket units have a rate limit for exactly this: `TriggerLimitIntervalSec=` and `TriggerLimitBurst=` (defaulting to a couple of hundred activations in a two-second window). Exceed it and the socket unit itself enters the failed state and stops listening — which looks, from outside, exactly like the port disappearing. The fix is to look at why the service crashes, then `systemctl reset-failed myapp.socket` and start the socket again. An engineer who has been paged for this recognises "the port vanished after the service started crash-looping" instantly. ## How it compares with the alternatives Socket activation gives you restart continuity on one host with no coordination and no extra moving parts, and it needs no second copy of the process. What it cannot do is shift traffic between hosts, hold two versions of the service up at once, or gate traffic on a health check — for that you need something in front. The honest framing in an interview is that socket activation removes the connection-refused window during a local restart, and that everything else about safe rollout lives elsewhere.

  • A colleague ran systemctl stop on the service to take it out of rotation, and it came back within seconds. What happened?
    The socket unit was still listening. Stopping the service only ends the process; the next connection that arrives on the socket triggers systemd to start it again, which is the whole point of on-demand activation. To take the endpoint down you stop the socket unit as well — stop the socket first, then the service, or the service can be re-triggered in between.
  • Does socket activation mean the application no longer needs a graceful shutdown path?
    No — it covers the opposite half of the problem. Socket activation keeps new arrivals from being refused, but connections the old process already accepted are closed when it exits. Requests in flight are still cut off unless the process handles SIGTERM by finishing current work and then exiting. You need both: the socket for the arrival path, a drain on SIGTERM for the accepted path.
  • Under what conditions will clients still see failures during a socket-activated restart?
    When the restart outlasts the client's own timeout, so a queued connection times out waiting to be accepted; and when the arrival rate multiplied by the restart duration exceeds the pending queue depth set by `Backlog=`, after which new connections are dropped instead of queued. Both are about how long the process is down, which is why a slow-starting service erodes the benefit.

saying these in an interview costs you the question

  • Claims in-flight requests survive the restart too
  • Thinks systemd buffers the request bytes and replays them
  • Believes the pending connection queue is unbounded
  • Says systemctl stop on the service frees the port
  • Treats it as a substitute for rolling deploys across hosts

context