skip to content

Grafana

The visualization and alerting layer that sits in front of everything else — Prometheus, Loki, SQL databases, cloud APIs — through one data-source plugin model. Interviews focus on dashboards people can actually read, templated variables, unified alerting, and keeping dashboards in version control.

on this pageshow

questions

page 2 of 2

For a platform running several Grafana instances across environments and teams, compare delivering dashboards, data sources and folder access as files on disk, via the Terraform Grafana provider, and via a Kubernetes operator with custom resources. What decides the choice, and where does each one hurt?

level: principalimportance: should knowfreq 28%

basics

~20 s

Files are simplest but need filesystem access, express no permissions and have no drift detection. Terraform drives the HTTP API, so it reaches hosted instances and can manage folders, teams and permissions, at the cost of state ownership and destroy blast radius. An operator reconciles custom resources continuously and fits GitOps and many instances, at the cost of a controller and CRD versioning.

open as a page

You own the observability platform and want engineers to be able to move from a metric spike, to one slow request, to that request's logs, for any service in the fleet. What conventions and policies have to be true across teams for that navigation to work reliably, and where do you accept that it will not?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Fleet-wide you need one service-identity convention mapped consistently into every store, the trace id present and exactly matchable in logs, exemplars on the key latency metrics, a sampling policy that keeps the traces those exemplars name, and retention windows aligned. Accept broken navigation outside trace retention and for unsampled, unremarkable requests.

open as a page

showing 31–32 of 32