skip to content

Grafana

The visualization and alerting layer that sits in front of everything else — Prometheus, Loki, SQL databases, cloud APIs — through one data-source plugin model. Interviews focus on dashboards people can actually read, templated variables, unified alerting, and keeping dashboards in version control.

on this pageshow

questions

page 1 of 2

In Grafana, a single query returns several labelled series over time. Walk through how you would choose between the Time series, Stat, Gauge, Table and Heatmap panels, and explain what each one does to the query result before it renders.

level: juniorimportance: must knowfreq 58%

answer

  1. panels render data frames, not queries
  2. stat/gauge = reduce calculation per series
  3. gauge needs explicit min/max
  4. table = one frame per query until joined
  5. heatmap = distribution, pre-bucketed or calculated

basics

~20 s

Time series plots every point over time. Stat and Gauge reduce each series to one number (last, mean, max); Gauge adds a bounded min/max scale. Table shows the raw rows. Heatmap shows distribution. Choose by trend vs single value vs distribution.

solid answer

~50 s

Every Grafana panel renders **data frames** — typed columns, normally a time field plus labelled numeric fields — so the panel choice is a rendering decision, not a query decision. - **Time series**: needs a time field; draws one line/bar/point series per numeric field. Default for trends. - **Stat**: applies a *reduce calculation* (Last non-null by default, or mean/max/min) to each series and shows one big number per series — five series means five tiles unless you filter or reduce first. - **Gauge**: same reduction, but rendered against a bounded scale, so it is only meaningful when you set Min/Max and thresholds (percentages, SLO burn, disk used). - **Table**: renders frames as rows; multiple queries stay separate frames unless you join them with a transformation. - **Heatmap**: shows a distribution per time bucket, from pre-bucketed histogram data or by bucketing values itself. Rule of thumb: trend → time series; current value → stat/gauge; raw detail → table; distribution/latency spread → heatmap.

go deeper

for a junior

Name the five panels and what each is for, and know that Stat and Gauge collapse a series to one number using a calculation such as Last non-null.

for a middle

Add the data-frame framing: panels render typed frames, so wrong output usually means wrong data shape, fixed in the query or with a transformation.

for a senior

Talk about honesty of the visualisation — sparkline behind a stat, heatmap instead of an averaged latency line, explicit gauge bounds — and about series-count effects on the browser.

for a principal

Frame panel choice as a communication contract for the dashboard's audience: what decision does each panel support, and which panels are noise that will be ignored during an incident.

## What a panel actually receives A Grafana data source never returns "a graph". It returns one or more **data frames**: a table with typed columns (fields), each field carrying a name, a type (time, number, string, boolean), labels, and display configuration. A typical metrics response is a frame with a `time` field and one numeric field per series, where the series identity lives in the field's labels (`{instance="a", job="api"}`). A panel is a *renderer over frames*. Switching panel type never changes what was queried or how much data was fetched — it changes how the same frames become pixels. That single fact answers most of the follow-ups an interviewer asks here. ## Time series The general-purpose panel. It requires a time field; each numeric field becomes a series drawn as lines, bars or points. Per-series appearance (fill, width, axis placement, stacking, connect-nulls, soft min/max) is field configuration, and per-series exceptions are field overrides. If your frame has no time field, the panel refuses to render and asks you to convert the data (usually with a transformation). If your query returns thousands of series, the panel will still try to draw them — the legend and the browser are the bottleneck, not the query. ## Stat Stat reduces. Its **Calculation** option (Last *, Last non-null, Mean, Max, Min, Total, Count, Range, Delta…) collapses each series to a single value, and it renders one tile per series. The two things candidates miss: (1) "Last *" includes nulls while "Last non-null" skips them, which changes what a flapping metric shows; (2) if you did not expect N tiles, the fix is upstream — aggregate in the query, or use a Reduce/Filter transformation — not a panel option. Stat can also draw a sparkline behind the number (Graph mode: Area), which is the miniature trend that makes a single number honest. ## Gauge Gauge uses the same reduction as Stat but paints it on a bounded arc. It is only meaningful when the value has a real ceiling: percent used, quota consumed, error budget burned. Without an explicit Min/Max, Grafana derives the bounds from the data, and the gauge silently rescales as the data moves — a classic source of dashboards that always look half-full. Gauges pair with thresholds; unbounded quantities (request rate, byte counters) belong in a stat or a time series instead. ## Table Table renders frames literally: fields become columns, rows become rows. It is the right panel for inventory-shaped data, top-N lists, and last-value-per-label summaries. Important behaviours: multiple queries produce multiple frames and the table shows a frame selector rather than merging them — joining requires a *Join by field* or *Merge* transformation. Cells can be styled (coloured background from thresholds, gauge cells, sparkline cells, data links), and column display is controlled by field config plus an Organize fields transformation for renaming/hiding/reordering. ## Heatmap Heatmap answers "how were the values distributed", not "what was the value". It has two modes: consume data that is already bucketed (classic histogram series, one series per bucket boundary) or calculate buckets itself from raw values. The Y axis is the value bucket, the X axis is a time bucket, and cell colour encodes count. It is the standard latency-distribution panel and is far more honest than an averaged latency line, because it shows the tail rather than hiding it. ## How to choose out loud State the question the viewer is asking: - "Is it getting worse?" → Time series. - "Is it OK right now?" → Stat (with sparkline) or Gauge if bounded. - "Which ones are bad?" → Table sorted by the offending field. - "How is it spread?" → Heatmap (or histogram for a non-time distribution). ## Traps worth naming - Panel type does not reduce query cost; a Stat over a 30-day range still fetches 30 days. - Bar chart and Pie chart expect a categorical string field, not a time field; feeding them time-series frames produces nonsense. - Multiple series in Stat/Gauge is a data-shape problem, solved by the query or a transformation. - Legends, tooltips and value mappings are display concerns and never change the underlying numbers.

  • Your Stat panel shows twelve tiles instead of one number. Why, and what are your options?
    Stat renders one tile per series, and the query returned twelve series. The options are, in order of preference: aggregate in the query so the backend returns one series; add a Reduce or Filter-by-value transformation to collapse or select series; or, if all twelve are genuinely interesting, switch to a Bar gauge or Table which are designed for many values. Changing panel options alone will not fix it.
  • What happens if you point a Time series panel at a frame with no time field?
    It cannot render and shows a message that the data is missing a time field. Grafana requires a field typed as time to build the X axis — a string or numeric column that merely looks like a timestamp does not qualify. The fix is to make the data source type the column as time, or to convert it with a transformation such as Convert field type.

The query is the recording; the panel is the playback device. Choosing Stat over Time series is like reading the last frame of a video instead of watching it — the tape is identical, you just look at less of it.

saying these in an interview costs you the question

  • Believing Stat or Gauge queries less data than Time series
  • Using a Gauge without setting Min and Max, so the scale silently follows the data
  • Expecting a Table panel to automatically merge the results of two queries
  • Reaching for a heatmap when the question is 'what is the value now' rather than 'how is it distributed'
  • Treating panel choice as a performance lever instead of a presentation choice

context

open as a page

Describe what a data source is in Grafana, and trace what happens from the moment a dashboard panel runs its query until the result is rendered — including where the credentials live and what actually comes back over the wire.

level: juniorimportance: must knowfreq 52%

basics

~20 s

A data source is a saved connection (type, URL, auth, options) plus the plugin that knows how to query that system. The browser asks the Grafana server, which attaches the stored credentials, calls the backing system, and returns typed data frames the panel renders.

open as a page

A Grafana dashboard needs a drop-down at the top that lists the environments actually present in the monitoring data, rather than a hardcoded list. Which Grafana variable type would you use, where does it get its options from, and what do the variable refresh settings 'Never', 'On dashboard load' and 'On time range change' control?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Use a Query variable: Grafana runs a metadata query against the data source (e.g. a label-values lookup) and turns each returned row into a drop-down option. Refresh decides when that query re-runs — never (options are read from saved dashboard JSON), on every dashboard load, or additionally whenever the time range changes.

open as a page

In Grafana's unified alerting, a Grafana-managed alert rule is defined as one or more data-source queries plus expressions rather than as a single threshold on a graph. Walk through how such a rule is evaluated and how it turns into individual firing alerts.

level: middleimportance: must knowfreq 50%

basics

~20 s

Each query runs and returns labelled series. Server-side expressions then chain on those results — reduce a series to one number, do math, apply a threshold — and one expression is marked the rule's condition. Every distinct label set that satisfies it becomes its own alert instance.

open as a page

Explain the states a Grafana-managed alert instance moves through — including its pending period — and what happens when the underlying query returns no data or fails outright.

level: middleimportance: must knowfreq 45%

basics

~20 s

An instance is Normal until the condition breaches, then Pending until it has breached continuously for the pending period, then Alerting, which notifies. Missing data and query failures are separate states whose handling is configurable per rule: treat as alerting, as normal, as a dedicated no-data alert, or keep the last state.

open as a page

Explain how a Grafana dashboard's time range and refresh setting reach an individual panel's query, including what the $__interval and $__rate_interval macros resolve to and how the Max data points and Min interval query options affect them.

level: middleimportance: must knowfreq 48%

basics

~20 s

The dashboard time range (from/to, also in the URL) is sent with every panel query. Grafana computes an interval = range ÷ max data points, floored by Min interval, and exposes it as $__interval; $__rate_interval widens it to guarantee enough samples. Auto-refresh re-runs every panel.

open as a page

In Grafana, you want a log line shown from a Loki data source to become a clickable link that opens the matching trace in Tempo. How is that link configured, where does the configuration live, and what has to be true of the log line itself?

level: middleimportance: must knowfreq 32%

basics

~20 s

On the Loki data source you define a derived field: a name, a matcher that pulls the trace id out of the log (a regex capture group, or a label / structured-metadata key), and an internal link naming the Tempo data source by UID. The log itself must carry the trace id.

open as a page

Grafana can load dashboards from YAML provider files in its provisioning directory. Explain what that provider actually does at runtime, and what changes for an engineer who then opens one of those dashboards in the UI and tries to save an edit.

level: middleimportance: must knowfreq 40%

basics

~20 s

A file provider tells Grafana to scan a directory of dashboard JSON files and load them into its database, rescanning periodically so file changes apply without a restart. Provisioned dashboards are read-only in the UI by default: saving is refused unless the provider sets allowUiUpdates, and removing the file removes the dashboard.

open as a page

You export a Grafana dashboard's JSON and commit it so the same file can be deployed to dev, staging and production. Which parts of the JSON model determine whether it works unchanged in all three, and what must be removed or parameterised first?

level: middleimportance: must knowfreq 34%

basics

~20 s

Keep a stable uid, drop the numeric id, and stop hardcoding per-instance data source UIDs in panels — either pin the same data source UID in every environment or drive panels from a data-source template variable. Also expect schemaVersion migration and version churn to add diff noise.

open as a page

A Grafana dashboard variable named `service` is multi-value with 'Include All' enabled. Explain what Grafana actually substitutes into a panel's query text when the user selects two services or 'All', and how format modifiers such as `${service:regex}`, `${service:pipe}`, `${service:csv}`, `${service:sqlstring}` and `${service:raw}` change that.

level: middleimportance: must knowfreq 52%

basics

~20 s

Interpolation is textual: Grafana replaces $service in the query string before sending it, and for multi-value it joins the selected values using a format appropriate to the data source (regex alternation, a pipe list, a CSV, quoted SQL strings). All expands to every option unless you set a custom all-value. :raw disables escaping entirely.

open as a page

Once a Grafana alert is firing, how does it reach a particular destination? Explain the role of notification policies and contact points, and how grouping and timing settings shape what actually arrives.

level: seniorimportance: must knowfreq 45%

basics

~20 s

Firing alerts enter a tree of notification policies. Each alert descends the tree matching on its labels; the first matching child wins unless that child is set to continue. The matched policy names a contact point and controls grouping labels, initial wait, batching interval and repeat interval.

open as a page

You are connecting Grafana to a metrics and logs backend shared by several teams. What authentication options does a Grafana data source offer, how are its secrets stored, and what is the fundamental authorization gap you have to design around?

level: seniorimportance: must knowfreq 42%

basics

~20 s

Options include basic auth, bearer tokens, custom headers (such as a tenant header), TLS client certificates, cloud-provider signing, and forwarding the user's OAuth identity. Secrets are stored encrypted server-side and never returned by the API. The gap: a data source is one shared identity, so dashboard permissions do not limit data access.

open as a page

During a planned maintenance window a team wants a set of Grafana alerts to stop notifying. Compare a silence, a mute timing on a notification policy, and pausing the alert rule — what each one actually stops, and when you would choose each.

level: middleimportance: should knowfreq 30%

basics

~20 s

A silence is an ad-hoc, time-bounded label matcher that suppresses notification while evaluation continues. A mute timing is a reusable recurring schedule attached to a notification policy. Pausing a rule stops evaluation entirely, so no state is tracked and nothing can be detected.

open as a page

In a Grafana panel, how do thresholds, value mappings, units and field overrides relate to one another, and what is the relationship between a panel threshold and an alert?

level: middleimportance: should knowfreq 42%

basics

~20 s

Field defaults (unit, decimals, min/max, thresholds, mappings) apply to every field; overrides re-apply properties to fields matched by name, regex, type or query, and win over defaults. Thresholds and mappings are display-only — they colour and rename values, they never trigger anything.

open as a page

Explain what Grafana panel transformations are, where in the query lifecycle they execute, and when you should push the same work into the query instead of doing it as a transformation.

level: middleimportance: should knowfreq 46%

basics

~20 s

Transformations are an ordered, per-panel pipeline that reshapes the data frames a data source already returned — join, reduce, filter, rename, group. They run in the browser after the query, so they cost rendering time and never reduce what was fetched.

open as a page

Grafana data-source plugins may ship a backend component or be frontend-only. What does the backend component do, and which Grafana capabilities stop working when a data source does not have one?

level: middleimportance: should knowfreq 32%

basics

~20 s

A frontend-only plugin builds the request in the browser and Grafana's proxy forwards it. A backend plugin runs as a server-side process that executes queries without a browser — required for anything evaluated headlessly: alert-rule evaluation, caching, snapshots-without-a-viewer and streaming.

open as a page

Grafana offers special entries in the data-source picker named -- Mixed --, -- Dashboard -- and -- Grafana --. Explain what each one does and what the limits are of combining several sources in a single panel.

level: middleimportance: should knowfreq 30%

basics

~20 s

Mixed lets each query row choose its own data source; Grafana runs them separately and returns the frames side by side with no correlation. Dashboard reuses another panel's already-fetched result instead of querying. Grafana is the built-in source for test data and dashboard-scoped annotations.

open as a page

A relational query renders fine in Grafana's table view, but the Time series panel reports that the data is missing a time field. Explain what Grafana needs from a SQL data source to draw a time series, and how that differs from what a metrics store or a document store returns.

level: middleimportance: should knowfreq 36%

basics

~20 s

Every source returns data frames. Metrics stores are typed, so time and labels are known. SQL is schema-on-read: Grafana needs a column typed as time, a numeric value column, and the time-filter macro so the dashboard's range reaches the WHERE clause.

open as a page

Grafana's Explore has a split view that can hold panes querying different data sources at once. Walk through using it to investigate a latency spike across a metrics, a logs and a trace backend, and say what has to line up for that workflow to hold together.

level: middleimportance: should knowfreq 26%

basics

~20 s

Start in Explore with the metric, split the view to open a second pane, keep the time ranges synced, then follow links between panes: metric exemplar or manual query to a trace, span back to logs. It holds together only if the time ranges match, the identifiers are shared, and the user can query every data source involved.

open as a page

Grafana dashboards offer a template variable of type 'Interval'. How do you configure its list of durations and its `auto` entry, how does the selected value reach a panel query, and when should the resolution be a user-visible choice rather than left to the duration Grafana computes for each panel?

level: middleimportance: should knowfreq 36%

basics

~20 s

You give it a literal comma-separated list of durations, optionally plus an auto entry computed from the dashboard time range divided by a configured step count and floored by a minimum. It interpolates by name, appears in the URL, and is one value shared by all panels.

open as a page

Grafana panels and rows have a 'Repeat options' setting that points at a template variable. Explain what determines how many copies appear, what the repeat direction and 'max per row' controls do, and what the practical limitations of repeats are.

level: middleimportance: should knowfreq 38%

basics

~20 s

A repeated panel or row is cloned once per currently selected value of the chosen variable, so the variable must be multi-value or have Include All. Direction controls horizontal versus vertical layout and max-per-row caps the horizontal run. Only the source panel exists in the dashboard JSON; clones are generated at render time and cannot be edited individually.

open as a page

Grafana-managed alert rules are organised into evaluation groups inside a folder. What does a group actually control at evaluation time, and how should you use grouping when a folder holds many rules?

level: seniorimportance: should knowfreq 32%

basics

~20 s

A group sets the evaluation interval shared by its rules and evaluates them in order, one after another, at each tick. It controls scheduling and ordering only — it has nothing to do with how notifications are grouped or routed.

open as a page

A Grafana dashboard with dozens of panels takes over thirty seconds to load and is putting visible load on the metrics backend every time someone opens it. How do you diagnose it, and what levers do you have inside the dashboard itself?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Measure first with the panel inspector and the browser network tab: which panels are slow, and how many requests fire. Then cut the query count (fewer panels, collapsed rows, fewer repeats, reuse one query across panels) and cut per-query cost (max data points, min interval, narrower range, backend pre-aggregation).

open as a page

Explain what an exemplar attached to a Prometheus or Mimir time series is, everything that has to be enabled end to end for one to appear as a clickable dot on a Grafana time series panel, and why those clicks sometimes lead to a trace that no longer exists.

level: seniorimportance: should knowfreq 28%

basics

~30 s

An exemplar is a single sampled observation attached to a metric point, carrying labels such as a trace id. The instrumentation must record it, the exposition format must carry it, the metrics store must have exemplar storage enabled, and the Grafana data source needs an exemplar link to the trace store with the panel's exemplars option on. Clicks dead-end when the trace was never sampled or has already been dropped.

open as a page

You are looking at a span inside Grafana's trace view and want to jump to that request's log lines. How is that jump configured on the Tempo data source, and why does it so often come back empty?

level: seniorimportance: should knowfreq 30%

basics

~20 s

On the Tempo data source's trace-to-logs section: pick the logs data source, map span/resource attributes to log labels (tags), set a time-window shift around the span, and optionally filter by trace or span id. It comes back empty mostly from attribute-to-label name mismatches or too narrow a time window.

open as a page

Grafana's unified alerting can be delivered from files in the provisioning directory alongside dashboards and data sources. What rules does that delivery mechanism impose — object identity, what is replaced versus merged, and what happens to those objects in the UI?

level: seniorimportance: should knowfreq 24%

basics

~20 s

Alerting provisioning files declare rule groups, contact points, notification policies, mute timings and templates. Every rule needs a stable uid or reloads recreate it and lose its state. Provisioned objects are read-only in the UI. The notification policy tree is a single root object, so provisioning it replaces the whole tree rather than merging.

open as a page

You need to provision a Grafana data source that requires a password or API token, without committing the secret to the repository that holds the provisioning files. What mechanisms does Grafana give you, and what happens to that secret afterwards?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Put the secret under secureJsonData in the datasource provisioning file, but reference it indirectly: Grafana interpolates environment variables and can read a value from a mounted file. Grafana encrypts secureJsonData in its database and never returns it through the API, so rotation means changing the source and reloading provisioning.

open as a page

Grafana offers a variable type called 'Ad hoc filters'. Explain how it reaches a panel's query differently from a normal query variable, what it needs from the data source to work, and the situations where it produces surprising results.

level: seniorimportance: should knowfreq 28%

basics

~20 s

An ad-hoc filter variable is never referenced with $name. It holds a list of key/operator/value filters that Grafana injects automatically into every query sent to its chosen data source on that dashboard. It needs the data source plugin to support both the injection and the key/value suggestion API.

open as a page

On a Grafana dashboard, a 'namespace' drop-down should restrict the options offered by a 'pod' drop-down. How do you build that dependency between two template variables, how does Grafana decide what to re-query when the parent changes, and what breaks at scale?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Chain them: the child variable's query references the parent ($namespace) in its filter. Grafana builds a dependency graph from those references and, when the parent changes, re-runs each dependent query in order, re-resolving current selections. At scale the cascade serialises, an 'All' parent expands into a huge filter, and stale child selections silently reset.

open as a page

Grafana lets you write alert rules that Grafana itself evaluates, or rules that are pushed down to be evaluated by the metrics backend. What does each choice change operationally, and how would you decide for a large multi-team estate?

level: principalimportance: should knowfreq 28%

basics

~20 s

Grafana-managed rules are evaluated by Grafana, can chain expressions and combine data sources, and depend on Grafana's own availability and state store. Data-source-managed rules live in the backend's ruler, evaluate where the data is, scale with it, and are limited to that backend's query language.

open as a page

showing 1–30 of 32