skip to content

Kamal

Deploying containers to plain servers over SSH with zero-downtime rolling releases behind a Traefik proxy - a deliberately small alternative to a cluster. Comes up when an interviewer wants to see whether you can size the platform to the problem instead of reaching for Kubernetes by reflex.

on this pageshow

explore

questions

4

In a Kamal deploy, how does the new container take over traffic from the one already running without dropping requests?

level: middleimportance: must knowfreq 48%

answer

  1. two versions coexist for a moment
  2. boot first, route later
  3. the healthcheck is the gate
  4. drain before stopping the old one
  5. boot: limit and wait pace the fleet

basics

~20 s

Kamal boots the new container alongside the old one, waits for its healthcheck endpoint to pass, and only then tells its proxy to switch traffic over, draining in-flight requests from the old container before stopping it. A container that never reports healthy never receives traffic.

solid answer

~50 s

Every release is tagged with the git revision, so the new container is a distinct version, not a restart of the old one. On each host Kamal pulls the image, boots the new container, and leaves it off the proxy while a healthcheck polls its endpoint — `/up` by default. Once it passes, the proxy atomically switches new requests to the new container, lets in-flight requests on the old one finish, and then stops the old container. If the healthcheck never passes, the deploy aborts on that host and the old container keeps serving. Across several hosts, `boot: limit:` and `boot: wait:` turn this into a rolling release rather than a simultaneous swap. As of Kamal 2 the proxy is `kamal-proxy`, which performs the health polling and the switch; Kamal 1.x used Traefik plus a temporary healthcheck container.

code

yaml · 12 lines
yaml
proxy:
  ssl: true
  host: app.example.com
  app_port: 3000
  healthcheck:
    path: /up
    interval: 3
    timeout: 3

boot:
  limit: 25%
  wait: 2

go deeper

for a junior

Learn the order and say it in order: pull, boot the new container beside the old, wait for the healthcheck, switch the proxy, then stop the old one.

for a middle

Explain why the healthcheck gate is what makes it zero-downtime, what boot: limit: and wait: change across a fleet, and what happens to the deploy when a container never reports healthy.

for a senior

Show judgment about the gate itself: what the healthcheck should and should not assert, how a partially failed rollout leaves mixed versions, and how migrations must tolerate both versions running at once.

for a principal

Own the release policy above the mechanism. Decide how much of the fleet may move at once, what evidence beyond a 200 response should be required before continuing, and where an external load balancer or a real canary belongs instead.

## Why this is a cutover, not a restart A naive deploy stops the old container and starts a new one, which leaves a gap where nothing is listening. Kamal avoids the gap by making the two versions coexist for a few seconds. Each release is tagged with the git revision, so `myapp:9f3c1a2` and `myapp:4b70de1` are different containers that can run side by side on the same host; the switch is a decision the proxy makes, not a restart of a process. ## The sequence on one host 1. **Pull.** The host pulls the image tag for the new version from the registry. 2. **Boot, but do not route.** Kamal starts the new container with the app's environment. At this moment it is running but the proxy is still sending everything to the old one. 3. **Health gate.** The new container is polled at its healthcheck path — `/up` by default, the Rails convention, configurable — until it responds successfully or the attempts run out. This is the load-bearing step: a container that boots but cannot reach its database will fail here and the deploy stops. 4. **Switch.** The proxy is told to route to the new container. New requests go to the new version from that instant. 5. **Drain and stop.** Requests already in flight on the old container are allowed to finish before it is stopped, so no in-progress response is cut off. The stopped container is retained for a while so a rollback can reuse it. If step 3 never succeeds, steps 4 and 5 never happen. The old version keeps serving and you get a failed deploy rather than an outage — which is the behaviour you want, and worth saying explicitly in an interview. ## Across many hosts Run the same sequence on twenty servers at once and you briefly double your memory footprint everywhere and expose every user to the new version simultaneously. The `boot:` settings pace it: ```yaml boot: limit: 25% wait: 2 ``` `limit` caps how many hosts are being deployed at any moment (a count or a percentage) and `wait` inserts a pause between groups. That is what makes a Kamal release a genuine rolling release across the fleet rather than a simultaneous swap — each group is fully healthy before the next begins. ## The proxy, and the version caveat As of Kamal 2 (released late 2024) the proxy is **kamal-proxy**, a small purpose-built reverse proxy that runs as its own container on the host, owns ports 80 and 443, terminates TLS with automatic Let's Encrypt certificates when you set `ssl: true` with a `host`, performs the healthcheck polling itself, and does the atomic switch. Kamal 1.x instead ran **Traefik** with container labels, and booted a separate temporary healthcheck container to verify the new version before wiring it up. The user-visible behaviour is the same shape; the configuration keys are not, so say which version you are describing. ```yaml proxy: ssl: true host: app.example.com app_port: 3000 healthcheck: path: /up interval: 3 timeout: 3 ``` ## What this does and does not buy you It buys you zero-downtime releases of the application container, gated on the application's own opinion of its health. It does not make the release *safe* in any deeper sense. The healthcheck is one endpoint on one container — if `/up` returns 200 whenever the process is alive, without touching the database or the queue, the gate is decorative and you will happily cut over to a broken version. A useful healthcheck asserts the dependencies the app genuinely cannot serve without, and nothing more (checking a third-party API in it turns their outage into your failed deploy). It also does not coordinate anything outside the app container. Database migrations run separately and are not part of the cutover, so during the switch both versions are briefly live against the *same* schema — which is the real constraint on how you write migrations under any rolling release, Kamal included: each migration must be compatible with the version on either side of the switch.

  • What happens if the new container's healthcheck never passes on one host out of ten?
    That host's deploy fails and its old container keeps serving — the proxy is never switched there. Because `boot: limit:` deploys in groups, the run stops rather than continuing through the remaining hosts, so you are left with a partially updated fleet: some hosts on the new version, the rest on the old. You fix forward or roll back the ones that moved.
  • Both versions are briefly live at once. What does that imply for database migrations?
    Every migration must be compatible with the version on both sides of the switch. Add columns before code reads them, keep old columns until no running version writes them, and split renames into add–backfill–switch–drop across two releases. Kamal does not coordinate migrations with the cutover, so the discipline is entirely yours.
  • What makes a healthcheck endpoint useful rather than decorative here?
    It should fail when the app genuinely cannot serve requests — typically a cheap check that the database connection and any critical local dependency are usable — and it should not reach out to third-party services, or their outage becomes your failed deploy. An endpoint that returns 200 whenever the process is running gates nothing and will happily cut traffic over to a broken release.

saying these in an interview costs you the question

  • Says Kamal stops the old container first, then starts the new one
  • Thinks the cutover happens as soon as the container starts
  • Believes zero-downtime deploys make migrations automatically safe
  • Assumes traffic is switched simultaneously on every host by default
  • Treats an always-200 healthcheck endpoint as sufficient

context

open as a page

A Kamal deploy has shipped a bad release. What does `kamal rollback` actually do, and in which situations will it not save you?

level: seniorimportance: should knowfreq 33%

basics

~20 s

kamal rollback takes a previous version tag, brings that container back up on each host, and switches the proxy to it — no rebuild, no registry round trip, so it is fast. It restores only the app container: not the database, not accessories, and not host state.

open as a page

Your team runs a handful of services on a few VMs and is debating whether to adopt Kubernetes. Make the case for deploying with Kamal instead, and be explicit about what you give up.

level: principalimportance: should knowfreq 40%

basics

~20 s

Kamal fits when your capacity is known, your service count is small, and nobody is paid to run a control plane: you get zero-downtime container deploys over SSH for one config file. You give up scheduling, autoscaling, rescheduling on host failure, and continuous reconciliation of drift.

open as a page

What does Kamal need on a target server before it can deploy your app there, and how does the container image actually get onto that host?

level: juniorimportance: nice to knowfreq 35%

basics

~20 s

Kamal needs only an SSH login that can run Docker on the host — no agent and no control plane. It builds your image, pushes it to a container registry, then drives docker commands over SSH so each server pulls and runs it.

open as a page