skip to content

Templating and Variables

Making one dashboard serve every environment and service with variables instead of copy-paste. A practical interview favorite: variable-driven dashboards are the difference between 5 dashboards and 500.

on this pageshow

questions

6

A Grafana dashboard needs a drop-down at the top that lists the environments actually present in the monitoring data, rather than a hardcoded list. Which Grafana variable type would you use, where does it get its options from, and what do the variable refresh settings 'Never', 'On dashboard load' and 'On time range change' control?

level: juniorimportance: must knowfreq 62%

answer

  1. Query / Custom / Constant / Text box / Data source / Interval / Ad hoc
  2. options = rows of a metadata query
  3. Never = frozen in dashboard JSON
  4. On time range change = time-scoped label lookups
  5. selection lives in URL as var-<name>=

basics

~20 s

Use a Query variable: Grafana runs a metadata query against the data source (e.g. a label-values lookup) and turns each returned row into a drop-down option. Refresh decides when that query re-runs — never (options are read from saved dashboard JSON), on every dashboard load, or additionally whenever the time range changes.

solid answer

~60 s

A **Query** variable is the right type: you pick a data source and write a query whose *result rows become the option list* (for a metrics store that is typically a label-values lookup; for SQL it is a `SELECT DISTINCT`). The other types exist for other jobs: **Custom** (a hand-typed static list), **Constant** (a hidden fixed value, useful for dashboards-as-code), **Text box** (free typing), **Data source** (the drop-down picks which data source the panels query), **Interval** (a list of durations), and **Ad hoc filters** (auto-applied key/value filters). Refresh applies to query variables only: - **Never** — options are not re-queried; the values saved in the dashboard JSON are used. Fast, but goes stale. - **On dashboard load** — the query runs each time the dashboard opens. - **On time range change** — also re-runs whenever the picker moves, which matters because most metadata queries are time-scoped: an environment that stopped reporting yesterday should disappear from a 'last 15 minutes' view. Also set: multi-value, Include All, and Hide (label/variable) for variables that are plumbing rather than UI.

go deeper

for a junior

Name the query variable, say its options come from a query against the data source, and recite the three refresh modes with one sentence each.

for a middle

Add why 'on time range change' exists (time-scoped metadata queries) and the load-time cost of refreshing several variables per dashboard open.

for a senior

Frame refresh as a staleness-versus-cost decision, mention high-cardinality label lookups dominating dashboard load time, and note hide/constant variables as the dashboards-as-code seam.

for a principal

Set a house standard: which variables are query-backed versus provisioned constants, cardinality caps on lookups, and a data-source variable so one dashboard serves every identical stack.

## What a variable is A Grafana *template variable* is a named placeholder defined once at dashboard level and referenced inside panel queries, panel titles, links and even other variables. Its current value lives in the dashboard's URL as `var-<name>=<value>`, which is why a filtered dashboard can be shared as a link and why the browser back button moves through prior selections. Variables are what turn one dashboard into N dashboards without copying panels. ## The types, and what each is for **Query** — the only type whose options come from live data. You choose a data source and write a *metadata* query, not a graphing query: a label-values lookup on a metrics store, a tag-keys/tag-values call, or `SELECT DISTINCT env FROM ...` on SQL. Whatever the query returns is flattened into a list of options; most data sources return a text/value pair so the drop-down can show a friendly name while the query substitutes an id. A regex field can post-filter or capture part of each returned string, and a sort option makes the order deterministic instead of data-order. **Custom** — a comma-separated list you type (`prod,staging,dev`, or `Prod : production` for label/value pairs). Correct when the set is genuinely fixed and you do not want a query on every load. **Constant** — one value, normally hidden. Its purpose is dashboards-as-code: a provisioned dashboard can carry a cluster name or a job label as a constant rather than being edited per environment. **Text box** — free-form input; good for a trace id or a user-typed pattern. **Data source** — the options are the configured data sources of a chosen plugin type. Panels then point at `${ds}`, so the same dashboard can be aimed at the EU or US metrics cluster. This is what makes one dashboard reusable across identical stacks. **Interval** — a list of durations (`1m,5m,1h`) plus an optional `auto` entry, used inside range functions. **Ad hoc filters** — a special type that does not get referenced by name at all; the filters it produces are injected into every query for its data source. Grafana also ships **built-in/global** variables you never define: `$__interval`, `$__rate_interval`, `$__range`, `$__from`/`$__to`, `$__timeFilter` for SQL sources, plus `$__dashboard`, `$__user`, `$__org`. ## Refresh semantics, precisely Only query variables have a refresh setting because only they run a query. - **Never** means the option list is whatever was persisted in the dashboard JSON at save time. It is the fastest and the most misleading: a new service will not appear until somebody re-saves the dashboard. Reasonable for a slow-moving, expensive lookup; wrong for anything churny like pod names. - **On dashboard load** re-runs the query at open time. This is the sane default for most variables. - **On time range change** re-runs on open *and* every time the time picker moves or auto-refresh fires a range shift. You want this whenever the metadata query is itself time-bounded — label-value lookups usually are — otherwise a dashboard opened on 'last 7 days' and then narrowed to 'last 5 minutes' still offers hosts that have been dead for days, and selecting one yields empty panels that look like a broken dashboard rather than an absent host. The cost is real: every refresh is an extra query per variable at load, and with chained variables the cascade multiplies. High-cardinality label lookups over long ranges are among the most expensive things a dashboard does, and they run *before* any panel renders, so they show up as slow dashboard load rather than as a slow panel. ## Presentation options that ride along **Hide** can hide the label or the whole control — use it for variables that exist only to feed other variables or to carry a provisioned constant. **Multi-value** allows selecting several options; **Include All** adds an `All` entry. Both change what gets substituted into queries, so enabling them is a query-compatibility decision, not a cosmetic one. **Selection options → current value** is what gets saved as the dashboard's default for the next visitor. ## How to answer in an interview Name the type (query variable), say the options come from a metadata query against the data source, and then show you understand the trade-off in refresh: staleness versus per-load query cost, and the specific reason `on time range change` exists (time-scoped metadata). Mentioning that the selection is URL-encoded as `var-name=` and therefore shareable is a cheap credibility point.

  • Why would you ever set refresh to 'Never'?
    When the lookup is expensive or the option set is effectively static — a list of clusters or regions that changes a few times a year. The option list is then read from the saved dashboard JSON and costs nothing at load. The price is that new values only appear when someone re-saves the dashboard, so it is a poor fit for anything with churn like pods, containers or build ids.
  • A user reports that a host they know is reporting does not appear in the drop-down. Where do you look?
    Check the variable's refresh setting first — 'Never' means you are seeing a snapshot from the last dashboard save. Then check whether the metadata query is time-scoped and whether the current time range covers the host's data. Finally check the variable's regex/filter field and its data source, since a capture-group regex can silently exclude values that do not match.
  • What does the 'Hide' option on a variable do, and when is it right?
    It hides the label, or the whole control, from the dashboard's variable bar. It is right for variables that exist as plumbing — a constant injected by provisioning, or an intermediate query variable that only feeds a chained child — so the top of the dashboard shows only the controls a human is meant to touch.

saying these in an interview costs you the question

  • Thinking a query variable's query is a graphing query — it is a metadata/label lookup whose rows become options.
  • Assuming options refresh automatically; with refresh 'Never' they are frozen in the saved JSON.
  • Believing the refresh setting applies to custom or constant variables.
  • Using a Text box where a query variable belongs, so typos silently produce empty panels.
  • Not knowing the selected value is in the URL, then wondering why a shared link shows different data.

context

open as a page

A Grafana dashboard variable named `service` is multi-value with 'Include All' enabled. Explain what Grafana actually substitutes into a panel's query text when the user selects two services or 'All', and how format modifiers such as `${service:regex}`, `${service:pipe}`, `${service:csv}`, `${service:sqlstring}` and `${service:raw}` change that.

level: middleimportance: must knowfreq 52%

basics

~20 s

Interpolation is textual: Grafana replaces $service in the query string before sending it, and for multi-value it joins the selected values using a format appropriate to the data source (regex alternation, a pipe list, a CSV, quoted SQL strings). All expands to every option unless you set a custom all-value. :raw disables escaping entirely.

open as a page

Grafana dashboards offer a template variable of type 'Interval'. How do you configure its list of durations and its `auto` entry, how does the selected value reach a panel query, and when should the resolution be a user-visible choice rather than left to the duration Grafana computes for each panel?

level: middleimportance: should knowfreq 36%

basics

~20 s

You give it a literal comma-separated list of durations, optionally plus an auto entry computed from the dashboard time range divided by a configured step count and floored by a minimum. It interpolates by name, appears in the URL, and is one value shared by all panels.

open as a page

Grafana panels and rows have a 'Repeat options' setting that points at a template variable. Explain what determines how many copies appear, what the repeat direction and 'max per row' controls do, and what the practical limitations of repeats are.

level: middleimportance: should knowfreq 38%

basics

~20 s

A repeated panel or row is cloned once per currently selected value of the chosen variable, so the variable must be multi-value or have Include All. Direction controls horizontal versus vertical layout and max-per-row caps the horizontal run. Only the source panel exists in the dashboard JSON; clones are generated at render time and cannot be edited individually.

open as a page

Grafana offers a variable type called 'Ad hoc filters'. Explain how it reaches a panel's query differently from a normal query variable, what it needs from the data source to work, and the situations where it produces surprising results.

level: seniorimportance: should knowfreq 28%

basics

~20 s

An ad-hoc filter variable is never referenced with $name. It holds a list of key/operator/value filters that Grafana injects automatically into every query sent to its chosen data source on that dashboard. It needs the data source plugin to support both the injection and the key/value suggestion API.

open as a page

On a Grafana dashboard, a 'namespace' drop-down should restrict the options offered by a 'pod' drop-down. How do you build that dependency between two template variables, how does Grafana decide what to re-query when the parent changes, and what breaks at scale?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Chain them: the child variable's query references the parent ($namespace) in its filter. Grafana builds a dependency graph from those references and, when the parent changes, re-runs each dependent query in order, re-resolving current selections. At scale the cascade serialises, an 'All' parent expands into a huge filter, and stale child selections silently reset.

open as a page