skip to content

Your team runs a handful of services on a few VMs and is debating whether to adopt Kubernetes. Make the case for deploying with Kamal instead, and be explicit about what you give up.

level: principalimportance: should knowfreq 40%

answer

  1. do you have a scheduling problem?
  2. cost of owning, not installing
  3. dead container restarts, dead host does not
  4. capacity planning becomes human again
  5. the choice stays reversible

basics

~20 s

Kamal fits when your capacity is known, your service count is small, and nobody is paid to run a control plane: you get zero-downtime container deploys over SSH for one config file. You give up scheduling, autoscaling, rescheduling on host failure, and continuous reconciliation of drift.

solid answer

~60 s

The honest framing is that Kubernetes is a scheduler and a control plane, and you should adopt it when you actually need scheduling. If you run a handful of long-lived services on servers you have already sized, no scheduler is doing useful work for you — but you are still paying to operate, upgrade and secure the control plane, and to maintain a YAML estate that no small team maintains well. Kamal gives you the part you wanted: health-gated zero-downtime releases, a proxy with automatic TLS, rolling deploys across hosts, and a fast rollback, described in one `config/deploy.yml`. What you give up is real and you should name it unprompted: no bin-packing or autoscaling, no rescheduling when a host dies — a crashed container restarts on the same machine, but a dead machine's capacity is simply gone until you replace it — no continuous reconciliation of drift, and no built-in RBAC, secret store or multi-tenancy. Capacity planning becomes a human activity again, and whoever deploys holds SSH access to production.

go deeper

for a junior

Be able to say the tradeoff in one line: Kamal deploys containers to servers you manage yourself, which is simpler to learn but does nothing to place or move workloads for you.

for a middle

Explain concretely what a cluster provides that Kamal does not — scheduling, rescheduling after node failure, autoscaling — and which of those a small, predictable workload genuinely needs.

for a senior

Reason from the failure modes: what happens when a host dies, how mixed versions arise from a partly failed push deploy, and where an external load balancer has to sit for any of this to be resilient.

for a principal

Frame it as an ownership and staffing decision, state the exact conditions that would trigger a move to Kubernetes, and show the choice is reversible because the application shape is unchanged either way.

## The question behind the question An interviewer asking this is not testing whether you know Kamal's flags. They are testing whether you choose infrastructure by requirement or by reflex. The strongest answer starts by naming what Kubernetes fundamentally *is* — a scheduler plus a reconciling control plane — and then asks whether the team has a scheduling problem. Most teams with four services on three VMs do not. They have a *release* problem: how to ship a new container without dropping requests, and how to get back to the previous one quickly. That is precisely the problem Kamal solves, and it solves it without asking you to operate a distributed system to get there. ## What Kamal actually gives you - Health-gated zero-downtime cutovers: the new container must pass its healthcheck before the proxy sends it traffic. - A rolling release across hosts, paced by `boot: limit:` and `boot: wait:`. - A proxy on each host that terminates TLS with automatic Let's Encrypt certificates (`proxy: ssl: true` with a `host`, as of Kamal 2). - Fast rollback to a container still resident on the host. - Long-lived dependencies as `accessories:`, outside the release cycle. - All of it in one `config/deploy.yml` in the application repo, driven over SSH with no agent installed. The cost side is what makes the argument: no control plane to upgrade, no node pool to patch on someone else's schedule, no CNI or ingress controller choices, no cluster bill, and an onboarding cost measured in an afternoon rather than a quarter. ## What you give up — say this before you are asked - **Scheduling and bin-packing.** You decide which service runs on which host, by editing a list of IP addresses. That is fine for six services and painful for sixty. - **Self-healing beyond restart.** A crashed container is restarted by Docker's restart policy on the same host. A *dead host* is not handled at all: its containers do not move, and its share of capacity is gone until a human replaces the machine. Kubernetes reschedules; Kamal does not. - **Traffic-level failover.** Each host runs its own proxy, so distributing traffic across hosts and removing a dead one is an external load balancer's or DNS's job, not Kamal's. - **Autoscaling.** There is no horizontal autoscaler. Capacity is provisioned by people, ahead of time, which is fine for predictable load and wrong for spiky load. - **Continuous reconciliation.** Deploys are an imperative push, so a host's state is whatever the last successful deploy left there. If a deploy fails halfway across the fleet you get mixed versions and nothing corrects that on its own — contrast with a reconciling agent that continuously drives the cluster back to the declared state. - **The platform surface.** No namespaces, no RBAC, no admission control, no built-in secret store, no service mesh, no operators. If you need hard multi-tenancy between teams, that absence is decisive. - **The access model.** Deploying means having an SSH login that can run Docker on production. That concentrates a lot of authority in one credential, which is why Kamal deploys belong in CI with a dedicated key rather than on laptops. ## Where the line actually falls Kamal is a good fit when the service count is small and stable, load is predictable, the team has no platform engineer, and a few minutes of degraded capacity after a host failure is acceptable. It stops being a good fit when you need autoscaling, when multiple teams must share infrastructure with real isolation, when workloads are numerous and short-lived enough that placement decisions matter, when you need per-request scale-out, or — the most common real reason — when the organisation *already* runs Kubernetes competently and adding a second deployment model costs more than it saves. ## The meta-point interviewers are listening for Platform choice is a staffing decision as much as a technical one. Kubernetes is not hard to install; it is expensive to *own* — upgrades, certificate rotation, CVE response, cost control, the accumulated YAML. Choosing it commits someone's ongoing time forever. A principal-level answer states the migration path too: because Kamal deploys plain OCI images with an explicit healthcheck endpoint and externalised configuration, the application is already shaped correctly for Kubernetes later. The thing you would rewrite is the deployment description, not the application — so choosing the smaller platform now is a reversible decision, and saying that out loud is what turns the answer from a preference into an argument.

  • One of the three hosts running your Kamal-deployed app dies at 3am. What actually happens?
    Its containers do not move — nothing reschedules them, and that third of your capacity is gone until someone provisions a replacement and deploys to it. Whether users notice depends entirely on what sits in front: an external load balancer or health-checked DNS has to stop sending traffic to the dead host, because each host runs its own proxy and none of them knows about the others.
  • What would make you say 'this team has outgrown Kamal'?
    Load that needs autoscaling rather than pre-provisioning; enough services that deciding placement by editing IP lists becomes a real chore; several teams needing isolation from each other on shared infrastructure; or a compliance requirement for RBAC and audited access that an SSH login cannot satisfy. Any one of those means you now have the problem a scheduler and control plane exist to solve.
  • How reversible is the decision if you start with Kamal and later need Kubernetes?
    Largely reversible, because the application is already an OCI image with configuration in the environment and a healthcheck endpoint — the same shape a Deployment and a readiness probe expect. You rewrite the deployment description, not the application. The migration cost is real but bounded, which is exactly why starting small is defensible rather than reckless.

saying these in an interview costs you the question

  • Argues Kubernetes is always the professional choice regardless of scale
  • Claims Kamal gives self-healing equivalent to a scheduler
  • Ignores that a dead host's containers are never rescheduled
  • Compares only install effort, never the ongoing cost of ownership
  • Cannot name a single thing the smaller platform gives up

context