On a Grafana dashboard, a 'namespace' drop-down should restrict the options offered by a 'pod' drop-down. How do you build that dependency between two template variables, how does Grafana decide what to re-query when the parent changes, and what breaks at scale?
answer
- child's option query references $parent
- dependency graph → topological resolve, panels wait
- cascade is serial: load time = sum of levels
- All on parent amplifies child query — custom all value
- parent change drops an invalid child selection
basics
~20 sChain them: the child variable's query references the parent ($namespace) in its filter. Grafana builds a dependency graph from those references and, when the parent changes, re-runs each dependent query in order, re-resolving current selections. At scale the cascade serialises, an 'All' parent expands into a huge filter, and stale child selections silently reset.
solid answer
~1 minYou make the child's *option query* reference the parent variable — a label-values lookup filtered by `namespace="$namespace"`, or `SELECT DISTINCT pod ... WHERE namespace IN (${namespace:sqlstring})`. Grafana parses variable references to build a dependency graph and orders variable resolution topologically, so the parent resolves first and every dependent re-runs when it changes. Panels are only queried once all variables have settled. What bites in production: - **Serialised cascade.** A three-deep chain is three round trips before the first panel renders; dashboard load time becomes the sum, not the max. - **`All` on the parent.** The child query receives the whole expansion, so a cheap lookup becomes an expensive one. A custom all value like `.*` keeps it one token. - **Selection churn.** When the parent changes, a child value that no longer exists is dropped and Grafana falls back to the first option — panels change meaning under the user, and shared URLs carrying `var-pod=` can land on an invalid selection. - **Cardinality.** A pod-level child on a large cluster returns thousands of options; the drop-down and the query both suffer. Mitigations: cap the depth, use custom all values, prefer refresh 'on time range change' so options track the window, and hide intermediate plumbing variables.
go deeper
Know that a child variable's query can reference $parent and that changing the parent refreshes the child list.
Explain the dependency graph and ordered re-query, plus the multi-value format requirement in the child's filter.
Lead with the operational costs: serial cascade at load, All amplification, selection invalidation on shared URLs, and leaf cardinality — with concrete mitigations.
Set dashboard standards: maximum chain depth, mandatory custom all values, stable-dimension-first ordering, and a preference for ad-hoc filters over hand-rolled ladders where the data source supports them.
## Building the chain A chained (dependent) variable is nothing more special than a query variable whose *option query* contains a reference to another variable. Grafana interpolates that reference exactly as it would in a panel query, so the child sees a concrete filter. For a metrics backend that looks like a label-values call scoped by `namespace="$namespace"`; for SQL, a `WHERE namespace IN (${namespace:sqlstring})`; for a search backend, an added filter clause. Because the reference is textual, everything from the interpolation rules applies: if the parent is multi-value, the child's query must accept the joined format (regex-match rather than equality, `IN` rather than `=`). ## How Grafana resolves the graph At dashboard load and on every variable change, Grafana determines which variables reference which others and processes them in dependency order. A change to the parent invalidates its descendants: their queries re-run, their option lists are rebuilt, and their current values are re-validated against the new lists. Panel queries are held until variables have settled, which is why a dashboard with a deep chain shows its variable bar before it shows any data. Cycles are not allowed — a variable cannot, directly or transitively, depend on itself — and Grafana will refuse or misbehave if you try. ## The failure modes **Load-time serialisation.** Independent variables can be fetched concurrently; chained ones cannot, by definition. Each level is a round trip to the data source, and metadata queries are frequently the *slowest* queries a dashboard issues because they scan label indexes over the whole time range. Three levels of chaining on a busy backend routinely turns a sub-second dashboard into a multi-second one, and the delay lands before any panel paints, which users read as 'the dashboard is broken' rather than 'the dashboard is loading'. **`All` amplification.** If the parent has Include All with no custom all value, selecting All expands to the full option list, and *that entire expansion* is pasted into the child's query. A child lookup that is cheap for one namespace becomes a cluster-wide scan. Setting a custom all value (`.*` for regex-matched sources, `%` for LIKE-style SQL) collapses it to a single token and is the single highest-value fix on most chained dashboards. **Selection invalidation.** When the parent changes, the child's previously selected value may not exist in the new option list. Grafana drops it and falls back to the first available option (or All, if enabled). Two consequences: a user who carefully selected a pod loses it whenever they nudge the namespace, and a shared URL that pins `var-pod=` can arrive at a dashboard where that pod is not in the (newly computed) list, producing an empty panel that looks like missing data. Design around it by chaining on stable dimensions — environment, region, service — and letting the churny dimension (pod, container, instance) be the leaf, ideally with All as its default. **Cardinality.** The leaf of a chain is usually the highest-cardinality dimension. A drop-down with 5,000 entries is unusable, the option query is expensive, and selecting All then blows up every panel query. Constrain with the variable's regex filter, or accept that the leaf exists mainly to *drill in* after the user has narrowed the parents. **Hidden coupling with time.** Metadata queries are time-scoped on most backends. If the parent refreshes on time-range change but the child refreshes only on load (or vice versa), you can end up with a child list computed for a different window than the parent — inconsistent and hard to debug. Keep refresh settings consistent down a chain. ## Design guidance - Keep chains shallow: two levels is normal, three is the practical ceiling, four means the dashboard is trying to be a query builder. - Order by stability: slow-changing, low-cardinality dimensions at the top. - Give every high-cardinality parent a custom all value. - Hide plumbing variables that exist only to filter a child. - Where the backend supports it, consider an ad-hoc filter variable instead of a hand-built chain: it pushes key/value selection into the data source's own suggestion API rather than into a ladder of dashboard queries. - Validate with the panel inspector and the browser network tab: the variable queries are visible as separate requests and their timings tell you exactly which level of the chain is expensive. ## Interview framing Show the mechanism in one sentence (child query references parent; Grafana orders resolution by the dependency graph), then spend your answer on the operational consequences: serialised load, All amplification, selection invalidation, cardinality. That is what separates someone who has built a chained dashboard from someone who has operated one.
- Why does a deeply chained dashboard feel slow even when every panel query is fast?Variable resolution happens before panels are queried, and chained variables must resolve in order, so each level is an additional serial round trip. Metadata/label lookups are often the most expensive queries on the dashboard because they scan the index across the whole time range. The user sees a blank dashboard for the sum of all levels' latency before a single panel starts loading.
- How do you stop a shared dashboard URL from landing on an invalid child selection?Prefer chaining on stable dimensions and let the churny one default to All, so the leaf value is rarely pinned. Keep refresh settings consistent down the chain so the child's option list is computed for the same time window as the parent. Accept that Grafana will drop an option that no longer exists and fall back to the first entry — so the safest shareable link pins only the stable parents.
saying these in an interview costs you the question
- Claiming chained variables refresh in parallel — dependents are inherently serial.
- Leaving Include All with default expansion on a parent, then blaming the data source for slow child lookups.
- Building four- or five-deep chains as a query builder substitute.
- Assuming a child keeps its selected value when the parent changes.
- Mixing refresh modes down a chain so parent and child are computed over different time windows.