skip to content

An Ansible play sets `gather_facts: no` to speed up a run, and a later task referencing `ansible_facts['distribution']` fails with an undefined variable. Why does that happen, and what are your options?

level: seniorimportance: should knowfreq 48%

answer

  1. an implicit first task, nothing more
  2. the setup module produces them
  3. narrow the subset before disabling it
  4. cache trades freshness for speed
  5. cached facts rank below play vars

basics

~20 s

Facts are produced by the setup module, which Ansible runs implicitly at the start of a play. Disabling gathering skips it, so no ansible_facts exist. Options are to re-enable gathering, run setup explicitly for the hosts that need it, narrow it with gather_subset, or enable fact caching.

solid answer

~50 s

`gather_facts` controls one implicit task at the top of the play: Ansible runs the `setup` module on every host, which is what populates `ansible_facts`. Turn it off and that dictionary is simply not there, so any reference to it is undefined. There are four sensible responses. Re-enable gathering, if the facts are genuinely needed. Keep it off for the play and run `ansible.builtin.setup` as an explicit task only where you need it — the same cost, applied selectively. Keep gathering on but narrow it with `gather_subset` (for example `min`, or `!hardware`), since hardware and network discovery is where most of the time goes. Or configure fact caching in `ansible.cfg` — `fact_caching` with a `jsonfile` or `redis` backend plus a `fact_caching_timeout` — so facts come from the cache and gathering is skipped for hosts whose entry is fresh. Caching buys speed at the price of staleness, which is the tradeoff worth naming out loud.

code

yaml · 12 lines
yaml
- hosts: web
  gather_facts: no
  tasks:
    - name: Gather only the cheap subset, and only where needed
      ansible.builtin.setup:
        gather_subset:
          - min
      when: os_specific_branch | default(false)

    - name: Guarded read of a fact that may not exist
      ansible.builtin.debug:
        msg: "{{ ansible_facts.get('default_ipv4', {}).get('address', 'no default route') }}"

go deeper

for a junior

Know that facts come from an automatic first task in the play and that gather_facts: no removes it, so ansible_facts is empty and any reference to it fails.

for a middle

Explain the mechanism — the setup module and the ansible_facts dictionary — and name the levers: gather_subset, running setup as an explicit task, and fact caching in ansible.cfg.

for a senior

Show the tradeoff you actually make on a large fleet: where gathering time goes, when caching is acceptable, how you bound staleness, and when a branch should read inventory rather than a fact at all.

for a principal

Own the fleet-wide policy — the gathering mode, cache backend and timeout as a standard, and the rule for which decisions are allowed to depend on discovered state versus declared inventory.

## What gathering actually is There is nothing magical about facts. At the start of a play Ansible inserts an invisible first task that runs `ansible.builtin.setup` on every host in the play. That module connects, runs Python on the target, and returns a large dictionary describing the machine: OS distribution and version, interfaces and addresses, mounts, memory, CPU, virtualisation, environment. Ansible stores it as `ansible_facts` for the host, and (unless `inject_facts_as_vars` is disabled) also injects the top-level `ansible_*` names such as `ansible_distribution`. `gather_facts: no` removes that implicit task. Nothing else changes — the connection still works, inventory variables are still there — but `ansible_facts` is empty, so `ansible_facts['distribution']` is an undefined key and the task fails for that host. It is worth being precise about *why* people disable it. On a play targeting a few hundred hosts, gathering is a full round trip and a non-trivial Python run per host before any real work starts; on a play that only pushes one file it can dominate the wall-clock time. The instinct is right; the blanket `no` is the blunt version of it. ## Option 1 — gather, but only what you need ```yaml - hosts: web gather_facts: yes gather_subset: - '!all' - '!min' - network ``` `gather_subset` selects which fact collectors run: `all`, `min`, `hardware`, `network`, `virtual`, and negations with `!`. Hardware discovery (disks, mounts) is usually the expensive part, and most plays only want the OS family. `gather_timeout` caps a collector that hangs. This is often the cheapest fix and preserves normal semantics. ## Option 2 — gather explicitly, where needed ```yaml - hosts: web gather_facts: no tasks: - name: Facts only for hosts that need them ansible.builtin.setup: gather_subset: min when: needs_os_branching | default(false) ``` The implicit task is just a task; you can run the module yourself, later, conditionally, or against a delegated host. This keeps the play fast for the common path while making the dependency explicit and reviewable. ## Option 3 — cache the facts In `ansible.cfg`: ```ini [defaults] gathering = smart fact_caching = jsonfile fact_caching_connection = /var/tmp/ansible_facts fact_caching_timeout = 86400 ``` `gathering = smart` gathers only for hosts with no fresh cache entry; `implicit` always gathers unless the play says no; `explicit` never gathers unless the play says yes. With a cache configured, a second run within the timeout skips gathering entirely and reads facts from disk or Redis. `ansible-playbook --flush-cache` discards it when you need a clean read. The cost is staleness. A host that was resized, re-addressed or upgraded still reports its old facts until the entry expires, and a play that branches on `ansible_facts['distribution_major_version']` will happily do the wrong thing. Two practices contain that: keep `fact_caching_timeout` short enough to bound the damage, and flush the cache in any pipeline whose whole purpose is to react to a host change. There is also a precedence wrinkle: values restored from the cache are treated as host facts, which sit *below* play vars in the precedence ladder, whereas a `set_fact` executed in the current run sits near the top. The same name can therefore resolve differently depending on whether it was just computed or read back from cache. ## Option 4 — do not depend on the fact Sometimes the honest answer is that the branch should not exist. If the play only needs to know whether it is on a Debian- or RedHat-family host, and your inventory already knows that because the groups are `debian` and `el`, use `group_names` or an inventory variable — no gathering, no cache, and the value is under version control. Facts are for what only the machine can tell you; inventory is for what you already decided. ## Defensive reading Even with gathering on, not every fact exists on every platform — `ansible_facts['lsb']` is absent on hosts without the LSB tools, and `ansible_default_ipv4` is an empty dictionary on a host with no default route. Production playbooks guard these: ```yaml "{{ ansible_facts.get('default_ipv4', {}).get('address', 'unknown') }}" ``` Preferring the `ansible_facts` dictionary over the injected `ansible_*` names is also more robust, because a play variable of the same name silently shadows the injected form and `inject_facts_as_vars = false` removes it altogether.

  • Gathering is on but a play still fails reading ansible_facts['lsb'] on some hosts. Why?
    Because facts are whatever the target can report. The `lsb` subtree only appears on hosts with the LSB tooling installed, just as `ansible_default_ipv4` is empty on a host with no default route. Gathering guarantees the dictionary exists, not that any particular key does — read optional facts through `.get()` with a fallback.
  • What is the practical risk of enabling fact caching in a deployment pipeline?
    Acting on a stale picture of the fleet. A host that was resized, re-addressed or upgraded keeps reporting its previous facts until the cache entry expires, so a play branching on distribution version or memory can make the wrong decision confidently. Keep fact_caching_timeout short and flush the cache with --flush-cache in pipelines that exist to react to host changes.
  • Which fact subsets are usually worth excluding for speed?
    Hardware discovery is normally the expensive one — enumerating disks, mounts and CPU detail — followed by full network collection on hosts with many interfaces. Setting gather_subset to min, or to all with !hardware, keeps the OS and platform facts most plays actually branch on while cutting the bulk of the gathering time.

saying these in an interview costs you the question

  • Thinks facts appear automatically regardless of gather_facts
  • Believes disabling gathering breaks the connection itself
  • Treats cached facts as always current
  • Cannot name the setup module as the fact source
  • Assumes every fact key exists on every host

context