skip to content

You must apply a config change across 60 web servers with Ansible without taking the whole fleet out at once. Which play-level keywords control that, and how do handlers behave under them?

level: seniorimportance: should knowfreq 46%

answer

  1. batching lives on the play, not the task
  2. canary first, then wider batches
  3. the whole play repeats per batch
  4. handlers flush per batch
  5. abort the run when a batch fails

basics

~20 s

Set serial on the play to run it in batches of hosts instead of all at once, optionally with max_fail_percentage to abort if a batch goes badly. Because the whole play repeats per batch, handlers flush at the end of each batch, so restarts roll host group by host group.

solid answer

~50 s

`serial` is the keyword. It is set on the play, not on a task, and it slices the host list into batches: `serial: 5`, a percentage like `serial: "25%"`, or a ramp such as `serial: [1, 5, 25%]` to canary one host first. The critical detail is that Ansible **re-runs the entire play for each batch** — so handlers flush at the end of every batch, and the restart happens batch by batch rather than fleet-wide at the very end. That is what makes a rolling restart work at all. Pair it with `max_fail_percentage: 0` so the play aborts instead of marching a bad change through all twelve batches. Around the batch, `delegate_to` runs a task somewhere else (drain the node on the load balancer, from the controller) while keeping the current host's variables, and `run_once: true` performs a task once per batch instead of per host. `throttle` limits concurrency for a single task without changing the batching.

code

yaml · 31 lines
yaml
- name: Rolling config change across the web tier
  hosts: web
  serial: [1, 5, "25%"]
  max_fail_percentage: 0
  tasks:
    - name: Drain this host at the load balancer
      ansible.builtin.uri:
        url: "https://lb.internal/api/drain/{{ inventory_hostname }}"
        method: POST
      delegate_to: localhost

    - name: Push the application config
      ansible.builtin.template:
        src: app.conf.j2
        dest: /etc/app/app.conf
      notify: restart app

    - name: Apply the restart before re-enabling traffic
      ansible.builtin.meta: flush_handlers

    - name: Return this host to the pool
      ansible.builtin.uri:
        url: "https://lb.internal/api/enable/{{ inventory_hostname }}"
        method: POST
      delegate_to: localhost

  handlers:
    - name: restart app
      ansible.builtin.service:
        name: app
        state: restarted

go deeper

for a junior

Know that serial is a play-level keyword that processes hosts in batches instead of all at once, and that it is how rolling updates are expressed in Ansible.

for a middle

Explain the mechanism: the play is re-run per batch, so handlers flush at the end of every batch, and serial accepts a count, a percentage, or a list that ramps the batch size.

for a senior

Demonstrate the production shape — canary batch of one, max_fail_percentage: 0 so a bad change stops early, delegate_to for load-balancer drain and re-enable, and flush_handlers so the restart lands before traffic returns.

for a principal

Own the sizing decision: how much of the tier may be unavailable at once, how batch order interacts with failure domains, and where an Ansible rolling change stops being appropriate compared with replacing instances outright.

## The default is all at once By default Ansible walks tasks in order and runs each task across **all** hosts in the play (up to `forks`) before moving to the next. For a config change that ends in a service restart, that means every host restarts at roughly the same moment — a fleet-wide outage delivered efficiently. ## serial: batching the play `serial` is a play keyword that splits the host list into batches. The play runs completely for batch one, then completely for batch two, and so on. ```yaml - hosts: web serial: 5 max_fail_percentage: 0 tasks: - name: Push the config ansible.builtin.template: src: app.conf.j2 dest: /etc/app/app.conf notify: restart app handlers: - name: restart app ansible.builtin.service: name: app state: restarted ``` Accepted forms: - `serial: 5` — fixed batch size. - `serial: "25%"` — proportional to the number of hosts. - `serial: [1, 5, "25%"]` — a ramp: one host first as a canary, then five, then quarters. The last value repeats for the remaining batches. ## Why handlers are the crux Because the **entire play** is re-executed per batch, the end-of-play handler flush happens at the end of **each batch**. This is what turns "change config, notify restart" into a genuine rolling restart: batch one is reconfigured and restarted, then batch two, and so on. Without `serial`, the same playbook queues the handler on every host and restarts all sixty at the end of the single play. Two consequences follow: - Anything that must happen once for the whole run — a database migration, a notification — must be kept out of the batched play, or guarded, because a batched play's tasks execute again for every batch. - Facts gathered and variables set with `set_fact` inside the play are re-derived per batch, since the play genuinely runs again. ## Failing safely: max_fail_percentage By default Ansible keeps going as long as at least one host in the batch survives; a batch is only fatal when every host in it has failed. That is far too permissive for a rolling change. `max_fail_percentage: 0` aborts the whole play if any host in a batch fails, which is normally what you want: stop after the canary breaks, not after 60 machines are broken. A non-zero value tolerates a stated fraction per batch. ## Getting hosts in and out of the pool The usual rolling-update shape brackets the change with load-balancer work, and that work must run somewhere other than the host being drained: ```yaml - name: Drain this host at the load balancer community.general.haproxy: state: disabled host: "{{ inventory_hostname }}" backend: app delegate_to: lb01.example.com ``` `delegate_to` runs the task on a different host while the loop variable and the current host's variables still refer to the host being processed — which is exactly why `inventory_hostname` above names the machine being drained, not the load balancer. `delegate_to: localhost` (or the `local_action` shorthand) runs on the controller, the standard way to call an API. `run_once: true` executes a task a single time per batch on the first host, rather than once per host. ## Ordering and strategy `order` on the play controls how hosts are picked into batches — `inventory` (default), `sorted`, `reverse_sorted`, `shuffle` — which matters when your inventory happens to group all of one availability zone together and every batch would otherwise land in the same failure domain. `strategy` is a different axis. The default `linear` keeps hosts in lockstep: every host finishes task N before any starts task N+1. `strategy: free` lets each host race ahead through the whole play independently, which is faster but destroys the lockstep guarantee — and with it any assumption that hosts restart together. For a rolling change you want `linear` inside batches and `serial` between them, not `free`. `throttle: 2` caps how many hosts run a *single task* at once inside the current batch, useful when one particular step hammers a shared resource without needing the whole play batched more tightly. ## The tradeoff to say out loud Small batches mean a long run: sixty hosts at `serial: 1` is sixty full plays. Large batches finish fast and risk more capacity at once. The honest answer sizes the batch against how much capacity the service can lose and still serve traffic — plus a canary batch of one so a broken config takes out a single host, and `max_fail_percentage: 0` so the run stops there.

  • Why does serial make handlers restart services batch by batch rather than all at the end?
    Because serial re-runs the entire play for each batch, and handlers flush at the end of a play. Batch one therefore gets its config change and its restart before batch two starts. Without serial there is one play, so all notified handlers fire once at the very end and every host restarts together.
  • What does max_fail_percentage: 0 change compared with the default?
    By default Ansible continues as long as at least one host in a batch is still alive, so a broken change can march through every batch. `max_fail_percentage: 0` aborts the whole play as soon as any host in a batch fails, which stops the rollout at the canary instead of after the fleet is damaged.
  • How does delegate_to differ from simply changing the play's host pattern?
    `delegate_to` runs one task on another host while keeping the current host's variables in scope, so `inventory_hostname` still names the host being processed — which is what lets you tell a load balancer to drain *this* machine. Changing the host pattern would switch the whole play's context and lose the per-host loop you are batching over.
  • When would strategy: free be the wrong choice for this rollout?
    Almost always here. `free` lets each host run ahead through the play independently, so hosts no longer move in lockstep and the ordering guarantees a rolling change depends on disappear. It is for long, independent, per-host work where finishing fast matters more than coordination — not for a change that must land one batch at a time.

saying these in an interview costs you the question

  • Puts serial on a task instead of on the play
  • Expects handlers to fire once at the end of the whole run
  • Assumes a failing host aborts the play by default
  • Uses strategy: free to make a rolling restart faster
  • Runs the load-balancer drain on the host being drained

context