In an Ansible play, how do you let a task fail without aborting the run, and what does block/rescue/always give you that ignore_errors does not?
answer
- continue versus recover
- still reported failed
- redefine failure, do not hide it
- rescue clears the failed state
- always is the cleanup slot
basics
~20 signore_errors continues past a failed task but still marks it failed and offers no recovery path. block/rescue/always groups tasks, runs the rescue section when any task in the block fails, and always runs cleanup either way, so the host is marked recovered rather than failed.
solid answer
~50 sThree tools, three different jobs. `ignore_errors: true` runs on and the task is still reported failed — a blunt instrument that hides real errors alongside expected ones. `failed_when` redefines what failure *means* for that task, so `failed_when: result.rc not in [0, 1]` keeps the status honest instead of suppressing it. `block/rescue/always` is structured error handling: tasks in the `block` run in order, and if any of them fails, the remaining block tasks are abandoned and the `rescue` tasks run. If rescue completes successfully the host is no longer considered failed and the play continues; inside rescue you can inspect `ansible_failed_task` and `ansible_failed_result`. `always` runs in both paths, which is where cleanup belongs — remove the maintenance flag, re-add the node to the load balancer. One gotcha: a task carrying `ignore_errors: true` does **not** trigger rescue, because as far as the block is concerned nothing failed.
code
yaml · 21 lines- name: Deploy with rollback and guaranteed cleanup
block:
- name: Install the new release
ansible.builtin.unarchive:
src: "/tmp/app-{{ release }}.tgz"
dest: /opt/app
remote_src: true
- name: Smoke test the new release
ansible.builtin.uri:
url: http://localhost:8080/health
rescue:
- name: Show what failed
ansible.builtin.debug:
msg: "failed task: {{ ansible_failed_task.name }}"
- name: Roll back
ansible.builtin.command: /usr/local/bin/rollback
changed_when: true
always:
- name: Return the node to the load balancer
ansible.builtin.command: /usr/local/bin/lb-add
changed_when: truego deeper
Know the three keywords exist: ignore_errors continues past a failure, failed_when changes what counts as a failure, and block groups tasks with rescue and always sections.
Explain the control flow precisely — remaining block tasks are abandoned, rescue runs, a successful rescue clears the failed state, and always runs on both paths.
Show the judgment: prefer failed_when to suppression, put availability-restoring cleanup in always, and know that ignore_errors inside a block silently prevents rescue from ever firing.
Own the estate-wide convention for partial failure — whether a failed host should abort the run, how max_fail_percentage interacts with per-host rescue, and what evidence a run must leave behind after an automated rollback.
## The default: a failed host leaves the play When a task fails on a host, Ansible removes that host from the rest of the play. Other hosts continue; the failed one runs no further tasks and, importantly, none of its pending handlers. That default is usually right — you do not want to keep applying a half-broken change — but it gives you no way to react. ## ignore_errors: the blunt tool ```yaml - name: Best-effort cache warm ansible.builtin.command: /usr/local/bin/warm-cache ignore_errors: true ``` The host survives the failure and the play continues. Two things people forget: the task is still **reported as failed** in the output (it just is not fatal), and the suppression is total — a typo in the path, a permissions problem and the expected condition are all swallowed identically. `ignore_unreachable: true` is the separate keyword for a host you cannot connect to at all; `ignore_errors` does not cover that case. ## failed_when: define failure correctly instead of hiding it Often the task is not really failing — your definition of failure is just wrong for that command. `failed_when` takes a raw Jinja expression (like `when`) and replaces the module's own verdict: ```yaml - name: Query the migration tool ansible.builtin.command: /usr/local/bin/migrate --check register: migrate failed_when: migrate.rc not in [0, 1] changed_when: migrate.rc == 1 ``` This is the better answer whenever the non-zero exit is a documented, meaningful outcome. It keeps the recap truthful, whereas `ignore_errors` makes the recap lie in the other direction. ## block / rescue / always `block` groups tasks so keywords apply to the group; `rescue` and `always` turn that group into structured error handling. ```yaml - name: Deploy with rollback block: - name: Take the node out of the pool ansible.builtin.command: /usr/local/bin/lb-drain changed_when: true - name: Install the new release ansible.builtin.unarchive: src: "/tmp/app-{{ release }}.tgz" dest: /opt/app remote_src: true - name: Smoke test ansible.builtin.uri: url: http://localhost:8080/health rescue: - name: Roll back to the previous release ansible.builtin.command: /usr/local/bin/rollback changed_when: true always: - name: Return the node to the pool ansible.builtin.command: /usr/local/bin/lb-add changed_when: true ``` Semantics worth stating precisely: - If a block task fails, the **remaining block tasks are skipped** and control moves to `rescue`. - If `rescue` completes without failing, the host is **no longer failed** — the play continues normally for it. This is the property `ignore_errors` cannot give you: recovery, not suppression. - If a rescue task itself fails, the host fails for real. - `always` runs on both paths, including the success path where rescue never ran. It is the cleanup slot. - Inside `rescue`, `ansible_failed_task` and `ansible_failed_result` describe what went wrong, so you can log or branch on it. - A task inside the block with `ignore_errors: true` never triggers rescue, because the block sees no failure. This surprises people who mix the two. ## Keywords on a block Because a block is a task container, keywords set on it apply to every task inside: `when`, `become`, `tags`, `vars`, `delegate_to`. Guarding six related tasks with one `when` on a block is both shorter and less error-prone than repeating it six times. Blocks can be nested, though deeply nested error handling is a smell. ## Choosing between them in an interview The answer an interviewer is listening for is a decision rule, not a feature list: - The command's non-zero exit is a **legitimate outcome** → `failed_when`. - A genuinely optional, best-effort step whose failure means nothing → `ignore_errors`, sparingly, and with a comment saying why. - The failure needs a **reaction** — roll back, notify, mark the host — or there is state that must be undone whatever happens → `block/rescue/always`. One more operational note: `always` is the right home for anything that restores availability (returning a node to the load balancer, clearing a maintenance flag), precisely because it runs on the failure path too. Putting that step as a plain task after the block means a failure skips it and leaves the node drained — a classic 3am incident with a very boring cause.
- If the rescue section completes successfully, is the host still considered failed for the rest of the play?No. A rescue that runs to completion without failing clears the failure, and the host continues with the rest of the play normally, including later tasks and its handlers. If a task inside the rescue itself fails, the host fails for real and drops out as it would have without the rescue.
- A task inside a block has ignore_errors: true and it fails. Does the rescue section run?No. `ignore_errors` suppresses the failure before the block sees it, so from the block's point of view nothing went wrong and rescue is never entered. Mixing the two usually indicates confusion about intent — either the failure matters, in which case let it reach rescue, or it does not, in which case the block was not needed.
- Where would you put the step that returns a node to the load balancer, and why?In the `always` section. It runs on both the success and the failure path, so a failed deploy still restores the node instead of leaving it drained. A plain task after the block is skipped when the block fails, which is exactly how fleets end up with silently removed capacity after a bad run.
saying these in an interview costs you the question
- Treats ignore_errors as the general-purpose error handler
- Thinks ignore_errors makes the task report ok
- Expects always to run only when the block succeeded
- Believes rescue triggers for a task marked ignore_errors
- Puts the load-balancer re-add after the block instead of in always