A Grafana dashboard with dozens of panels takes over thirty seconds to load and is putting visible load on the metrics backend every time someone opens it. How do you diagnose it, and what levers do you have inside the dashboard itself?
answer
- one request per panel — count them in the network tab
- panel inspector: request, response size, backend timing
- collapsed rows do not query until expanded
- repeat over N values = N queries
- max data points ≈ panel width; push aggregation to recording rules
basics
~20 sMeasure first with the panel inspector and the browser network tab: which panels are slow, and how many requests fire. Then cut the query count (fewer panels, collapsed rows, fewer repeats, reuse one query across panels) and cut per-query cost (max data points, min interval, narrower range, backend pre-aggregation).
solid answer
~60 s**Diagnose before tuning.** Each panel issues its own query request, so open the browser network tab on a cold load to count requests and see which ones dominate, and use each panel's **Query inspector** for the exact request, the response size and the backend timing. Separate slow-because-many from slow-because-heavy — and note whether the browser is still busy after the responses land, which points at rendering or transformations rather than the backend. **Cut the number of queries**: fewer panels per dashboard (split overview and drill-down, linked by data links); put detail panels in **collapsed rows**, which do not query until expanded; check **repeats**, since one repeated panel over a twenty-value list is twenty queries; remember annotation queries also run; reuse one panel's result in others via the dashboard's own shared-query source. **Cut the cost of each query**: set Max data points sensibly, set Min interval, avoid queries that return thousands of series, shorten the default range, slow or remove auto-refresh, and push aggregation into the backend with recording rules or pre-aggregated tables.
go deeper
Know that each panel runs its own query, and that fewer panels, a narrower time range and a slower refresh mean less load.
Use the panel inspector and the network tab to identify the offending panels, and apply max data points, min interval and collapsed rows deliberately.
Separate query count from query cost from render cost, quantify the multiplier from repeats, refresh and viewers, and push expensive aggregation into the backend as the durable fix.
Set an estate-level policy — panel budgets, default ranges, allowed refresh intervals, recording rules as shared infrastructure — and decide which dashboards are incident-grade and get the tight budget.
## Step 1: get evidence A Grafana dashboard is not one request. Every panel issues its own query request when it becomes visible, so the first question is always *how many requests, and how heavy is each*. - **Browser network tab, hard reload.** Count the query requests, look at the waterfall (are they serialised behind a connection limit?), and note total transferred bytes. This distinguishes "sixty cheap queries" from "three monstrous ones". - **Panel inspector → Query tab.** Shows the exact request sent, the raw response, the response size and the backend-reported timing for that panel. Run it on the two or three worst offenders from the waterfall. - **Panel inspector → Data / Stats.** Shows how many frames, fields and rows came back. A panel returning several thousand series is usually the real problem even when each request looks fast. - **After the responses complete**, is the tab still frozen? Then the cost is client-side: series count, table row count, or a heavy transformation pipeline — not the backend. - **Backend side**, check its own slow-query logging and its request-rate metrics, and correlate the spike with dashboard opens. If the backend degrades only when this dashboard is open, you have your answer without argument. ## Step 2: reduce the number of queries - **Panel count.** The most effective lever is fewer panels. Split into an overview dashboard with a handful of decision-supporting panels, and drill-down dashboards reached by data links. Most 60-panel dashboards are three dashboards that were never separated. - **Collapsed rows.** Panels inside a collapsed row are not queried until the row is expanded. Making detail rows collapsed-by-default converts a 60-query load into a 12-query load for the common case, and costs nothing when someone genuinely needs the detail. This is the single highest-value, lowest-risk change on most big dashboards. - **Lazy loading.** Grafana only queries panels as they scroll into the viewport, so a tall dashboard is already partly protected — but everything above the fold, and everything a wall display shows, still fires. - **Repeats.** A repeated panel or repeated row expands into one copy per value, each with its own query. A repeat over a list that grew from five to fifty services multiplies your load by ten with no edit to the dashboard. Cap or narrow whatever drives the repeat, or replace the repeat with one panel showing all series. - **Annotations.** Every annotation query on the dashboard runs on every load and refresh, and they are easy to forget because they draw almost nothing. - **Reuse a result.** Grafana's built-in dashboard data source lets a panel consume another panel's already-fetched result instead of issuing its own query — ideal for a graph and a table of the same data, or several stat tiles derived from one query plus transformations. - **Mixed sources multiply requests**, since each source in the panel is a separate request. ## Step 3: reduce the cost of each query - **Max data points.** Left unbounded or set absurdly high, a panel asks for far more points than it has pixels. Setting it to roughly the panel width removes transfer and render cost with zero visual loss. - **Min interval.** Floors the step at the collection resolution so wide ranges aggregate instead of streaming raw points. - **Series cardinality.** A query returning thousands of series is expensive on the backend, on the wire, and catastrophic in the browser. Aggregate in the query, or show a top-N. - **Default time range.** A dashboard saved at 30 days makes every open expensive. Save it at the range people actually use and let them zoom out. - **Refresh.** Auto-refresh multiplies everything by viewers ÷ interval, forever. Manual refresh, or a refresh no finer than the collection interval, is usually correct; reserve fast refresh for genuinely live wall displays with few panels. - **Push aggregation down.** Recording rules, continuous aggregates, summary tables — computing an expensive expression once in the backend beats computing it per viewer per refresh. This is the fix that scales; everything else is rationing. - **Query caching**, where the deployment offers it, deduplicates identical queries across viewers — very effective for wall dashboards where many people watch the same thing, and useless for dashboards where everyone picks different filters. ## Step 4: state the trade-offs Each lever costs something. Collapsed rows add a click during an incident. Coarser intervals hide short spikes. Longer refresh means staler numbers. Splitting dashboards adds navigation. A senior answer names the trade-off and ties it to the dashboard's purpose: an incident-response dashboard earns fast refresh and few panels; an exploration dashboard earns manual refresh and rich detail; a wall display earns caching and a fixed narrow range. Finally, close the loop by re-measuring the same way you diagnosed — request count and cold-load time — rather than declaring victory because it felt faster.
- How does putting panels in a collapsed row differ from simply moving them further down the dashboard?Both help, but differently. Panels below the fold are lazily loaded, so they query when the viewer scrolls to them — which on a wall display or a maximised window may be immediately. Panels inside a collapsed row do not query at all until the row is expanded, regardless of scrolling, so it is a firmer guarantee and it survives a large monitor.
- The requests finish in two seconds but the tab is unresponsive for twenty. Where do you look?That is client-side cost, not backend cost. Check the panel inspector's data stats for the number of series and rows returned, then the transformation pipeline, then the panel type — thousands of series in a time series panel, a table with tens of thousands of rows, or a heavy join in the browser will all do it. The fix is to aggregate or top-N in the query so less data reaches the browser.
saying these in an interview costs you the question
- Tuning before measuring — changing max data points without knowing which panels are slow
- Assuming the dashboard makes a single request for all panels
- Believing transformations reduce backend load
- Raising the auto-refresh rate to make a slow dashboard 'feel' fresher
- Ignoring repeats and annotation queries when counting where the requests come from
- Treating query caching as a fix for dashboards where every viewer selects different filters