skip to content

For a platform running several Grafana instances across environments and teams, compare delivering dashboards, data sources and folder access as files on disk, via the Terraform Grafana provider, and via a Kubernetes operator with custom resources. What decides the choice, and where does each one hurt?

level: principalimportance: should knowfreq 28%

answer

  1. Files: no state, no drift detection, no permissions, needs filesystem
  2. Terraform: API-driven, covers folders/teams/permissions, plan = drift, state + destroy risk
  3. Operator: continuous reconcile, instance selectors, GitOps, CRD/controller cost
  4. Common split: content by files/operator, structure by Terraform
  5. Two-zone estate: code-owned read-only folders vs sandbox folders

basics

~20 s

Files are simplest but need filesystem access, express no permissions and have no drift detection. Terraform drives the HTTP API, so it reaches hosted instances and can manage folders, teams and permissions, at the cost of state ownership and destroy blast radius. An operator reconciles custom resources continuously and fits GitOps and many instances, at the cost of a controller and CRD versioning.

solid answer

~60 s

**File provisioning** is a loader from disk into each instance's database: trivial to reason about, no extra component, dashboards rescanned on an interval. Limits — it needs filesystem access (so it does not work against a hosted Grafana), it has no reconciliation loop for data sources, it detects no drift, and in the open-source path it cannot express **folder permissions, teams or service accounts** at all. **Terraform provider** talks to the Grafana **HTTP API**, so it works anywhere, including hosted. It covers the objects files cannot: folders with permissions, teams, service accounts and tokens, org-level settings, alerting. It gives real plan/apply drift detection. Costs: a state file that must be owned and locked per instance, a destroy blast radius, and applies that happen on your CI cadence rather than continuously. **Grafana Operator** reconciles custom resources in Kubernetes, matching dashboards to instances by selector and pulling JSON from ConfigMaps, URLs or git. It fits GitOps and many instances well; costs are the controller itself, CRD version churn, and reconciliation debugging. Decide on: where the instances run, whether permissions must be code, who may edit in the UI, and how you want drift handled.

go deeper

for a junior

Not expected beyond knowing that files, Terraform and an operator are three ways to deliver the same objects.

for a middle

Contrast filesystem versus API delivery and note that permissions and teams need the API path.

for a senior

Weigh drift handling, state ownership, secret flow and per-instance scaling, and propose the common content-versus-structure split.

for a principal

Refuse a universal answer: state the deciding questions, pick for the situation, define the two-zone editing policy and name what each choice gives up.

## The axes that actually decide it Rather than a feature table, judge the three by five properties. **1. Reachability.** File provisioning requires you to place files on the instance's filesystem. That rules it out for a hosted Grafana and makes it awkward anywhere the instance is not yours to mount volumes into. Terraform and the operator both go through the HTTP API, so they work against anything you can authenticate to. **2. Object coverage.** This is the most under-appreciated difference. File provisioning covers dashboards, data sources, alerting, plugins — the content. It does **not** cover organisational structure in the open-source path: teams, users, service accounts, and folder permissions are API objects. If your requirement is "team X can edit dashboards in folder X and only view folder Y," file provisioning cannot express it and you need the API-driven path (with the caveat that fine-grained role-based access control is an Enterprise/Cloud capability, while basic folder permissions by role, team or user exist more broadly). **3. Drift handling.** Files push desired state on an interval (dashboards) or on apply/reload (data sources) but never *report* divergence. Terraform's plan is an explicit diff against recorded state and is the strongest story for "tell me what someone changed by hand." An operator continuously reconciles, which is the strongest story for "put it back automatically." Choose according to whether you want to be told or to be corrected. **4. State ownership and blast radius.** Terraform's power comes from its state file, and so do its failure modes: state must be remote, locked and backed up; a lost or mismatched state produces plans that want to recreate everything; a mis-targeted destroy removes real dashboards and folders. The operator keeps desired state in Kubernetes objects, so the failure mode shifts to CRD schema changes and controller bugs, which are usually less catastrophic but harder to debug when reconciliation silently fails. Files have no state at all — which is exactly why they cannot detect or repair drift. **5. Cadence and operator experience.** Files update within a rescan interval and are extremely easy for anyone to reason about. Terraform applies when CI runs. The operator applies continuously. For a small team with one or two self-hosted instances, files plus a CI job that validates JSON is often the correct, boring answer, and adding Terraform buys little. For dozens of instances or a hosted deployment, the API-driven paths pay for themselves quickly. ## Mixing them, deliberately The common mature setup is not a single choice: **files or operator for content** (dashboards, data sources) and **Terraform for structure** (folders, permissions, teams, service accounts, org settings), because content changes constantly and cheaply while structure changes rarely and matters. If you do mix, make ownership explicit per object type and never let two mechanisms manage the same object — two writers over one dashboard uid means the last apply wins and every review is a lie. ## The UI-editing question Whichever mechanism you pick, decide who may edit in the UI, because the mechanisms all express it: provisioned dashboards are read-only unless `allowUiUpdates` is set; provisioned alerting objects carry provenance that blocks UI edits until it is disabled; Terraform-managed objects are editable in the UI but the next plan shows the drift and the next apply reverts it. A workable policy is a **two-zone estate**: code-owned folders where the UI is read-only and every change goes through review, and sandbox folders where teams own their dashboards freely and the platform makes no promises. The failure state to avoid is a single estate where nobody can tell which zone they are in. ## Scale considerations - **Many instances.** Files mean the same content baked into or mounted onto every instance; the operator's instance selectors handle this natively; Terraform needs a workspace or provider alias per instance. - **Secrets.** All three ultimately write `secureJsonData`; files reference mounted secrets or environment variables, Terraform pulls from its own secret sources and records that the field is set (never the value, since the API will not return it), the operator reads Kubernetes Secrets. The secret store, not Grafana, remains the source of truth in all three. - **Identity.** Every mechanism depends on stable `uid`s — dashboards, data sources, alert rules. Pin them, whichever path you choose; portability of links, exemplar targets and derived fields depends on it. - **Review quality.** Raw dashboard JSON diffs are close to unreadable. Whatever the delivery mechanism, generating dashboards from a higher-level definition, or at least normalising churn fields in CI, is what makes as-code delivery worth anything to reviewers. ## Answer shape A strong answer refuses to name a universal winner and instead states the deciding questions: are the instances reachable by filesystem or only by API; must permissions be code; do you want drift reported or auto-corrected; who owns state; and how many instances are there. Then it picks for the described situation and names what it is giving up.

  • You must guarantee that a specific folder's dashboards are only editable through review, while another team keeps full UI freedom. How do you implement that?
    Split the estate into zones. Code-owned folders are populated by provisioning with allowUiUpdates off, so the UI refuses saves and the repository is the only way in; folder permissions, managed through the API path (Terraform or the operator), grant that team view access rather than edit. Sandbox folders are not provisioned at all, and the team has editor permission there with the platform making no durability or portability promises. The essential part is that a user can tell which zone they are in, so a refused save is expected rather than a surprise.
  • What is the strongest argument against introducing Terraform for a single self-hosted Grafana instance whose dashboards are already file-provisioned?
    It adds a state file to own, lock, back up and reconcile for benefits that instance may not need: drift detection matters only if people can change things by hand, and the UI is already read-only for provisioned dashboards. If permissions, teams and service accounts do not need to be code, Terraform's unique coverage is unused and you have taken on a destroy blast radius plus another tool in the release path. The right trigger to adopt it is a concrete requirement — hosted instances, permissions as code, or many instances — not the general appeal of having state.

saying these in an interview costs you the question

  • Claiming file provisioning can manage folder permissions, teams or service accounts
  • Letting two mechanisms manage the same dashboard uid, so applies fight and reviews mislead
  • Adopting Terraform or an operator with no requirement it uniquely satisfies
  • Ignoring Terraform state ownership, locking and destroy blast radius
  • Treating raw dashboard JSON diffs as reviewable without normalising churn fields or generating from a higher-level source

context