skip to content

Observability & Monitoring

Everything that tells you what production is doing: metric stores, dashboards, APM and distributed tracing, and log aggregation. Interviewers use this area to check that you can answer 'how would you know it broke, and how would you find out why' with named tools rather than hand-waving.

on this pageshow

explore

questions

197 · 15 sections

What does a Prometheus alerting rule contain, and what does its `for` duration change about when it fires?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A Prometheus alerting rule pairs a PromQL expression with an optional for duration, plus labels and annotations. The for duration requires the expression to hold continuously across that window, so the alert sits in pending before it becomes firing.

open as a page

In Prometheus's text exposition format, what does one metric line contain, and what does the server keep from the `# HELP` and `# TYPE` lines?

level: juniorimportance: must knowfreq 58%
basics
~20 s

One line carries a metric name, an optional label set in braces, a float64 value and an optional millisecond timestamp. The HELP comment describes the family and the TYPE comment declares counter, gauge, histogram, summary or untyped.

open as a page

In PromQL, what does an instant vector selector return versus a range vector selector, and which functions require each?

level: juniorimportance: must knowfreq 82%
basics
~20 s

An instant vector selector returns one sample per series at the query timestamp; a range vector selector adds a bracketed duration and returns every sample in that window. Functions like rate() only accept the bracketed range form.

open as a page

In Prometheus, what happens on each scrape of one target, and what do scrape_interval and scrape_timeout bound?

level: juniorimportance: must knowfreq 84%
basics
~20 s

Each cycle Prometheus sends one HTTP GET to the target's metrics path, parses the body, and stores every sample under that scrape's start timestamp. scrape_interval sets how often it repeats; scrape_timeout bounds one request. Prometheus also records an up sample per scrape.

open as a page

How does Prometheus find its scrape targets, and what does a service-discovery mechanism hand it besides a list of addresses?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Prometheus builds its target list from the discovery blocks in each scrape job: a static list, files on disk, or a platform API such as Kubernetes or EC2. Each target arrives carrying read-only metadata labels describing its origin.

open as a page

In Grafana, a single query returns several labelled series over time. Walk through how you would choose between the Time series, Stat, Gauge, Table and Heatmap panels, and explain what each one does to the query result before it renders.

level: juniorimportance: must knowfreq 58%
basics
~20 s

Time series plots every point over time. Stat and Gauge reduce each series to one number (last, mean, max); Gauge adds a bounded min/max scale. Table shows the raw rows. Heatmap shows distribution. Choose by trend vs single value vs distribution.

open as a page

Describe what a data source is in Grafana, and trace what happens from the moment a dashboard panel runs its query until the result is rendered — including where the credentials live and what actually comes back over the wire.

level: juniorimportance: must knowfreq 52%
basics
~20 s

A data source is a saved connection (type, URL, auth, options) plus the plugin that knows how to query that system. The browser asks the Grafana server, which attaches the stored credentials, calls the backing system, and returns typed data frames the panel renders.

open as a page

A Grafana dashboard needs a drop-down at the top that lists the environments actually present in the monitoring data, rather than a hardcoded list. Which Grafana variable type would you use, where does it get its options from, and what do the variable refresh settings 'Never', 'On dashboard load' and 'On time range change' control?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Use a Query variable: Grafana runs a metadata query against the data source (e.g. a label-values lookup) and turns each returned row into a drop-down option. Refresh decides when that query re-runs — never (options are read from saved dashboard JSON), on every dashboard load, or additionally whenever the time range changes.

open as a page

In Grafana's unified alerting, a Grafana-managed alert rule is defined as one or more data-source queries plus expressions rather than as a single threshold on a graph. Walk through how such a rule is evaluated and how it turns into individual firing alerts.

level: middleimportance: must knowfreq 50%
basics
~20 s

Each query runs and returns labelled series. Server-side expressions then chain on those results — reduce a series to one number, do math, apply a threshold — and one expression is marked the rule's condition. Every distinct label set that satisfies it becomes its own alert instance.

open as a page

Explain the states a Grafana-managed alert instance moves through — including its pending period — and what happens when the underlying query returns no data or fails outright.

level: middleimportance: must knowfreq 45%
basics
~20 s

An instance is Normal until the condition breaches, then Pending until it has breached continuously for the pending period, then Alerting, which notifies. Missing data and query failures are separate states whose handling is configurable per rule: treat as alerting, as normal, as a dedicated no-data alert, or keep the last state.

open as a page

In Zabbix, what is an item, and how do passive and active agent checks differ?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A Zabbix item is one configured metric on a host, named by a key. In a passive check the Zabbix server connects to the agent and asks for a value; in an active check the agent connects out and sends.

open as a page

In Zabbix, what does a trigger expression evaluate, and what does a recovery expression add?

level: middleimportance: must knowfreq 62%
basics
~20 s

A Zabbix trigger is a boolean expression over an item's stored history, built from functions like min or avg over a window rather than a threshold on the latest value. A separate recovery expression clears the problem on a looser condition.

open as a page

In Zabbix, how do templates and low-level discovery keep hundreds of hosts configured?

level: middleimportance: should knowfreq 55%
basics
~20 s

A Zabbix template is a reusable container of items, triggers and discovery rules linked to hosts; editing it changes every linked host at once. Low-level discovery expands prototypes into one item or trigger per entity found on each host.

open as a page

In Zabbix, where does the server bottleneck as hosts are added, and what does a proxy change?

level: seniorimportance: nice to knowfreq 33%
basics
~20 s

A Zabbix server is sized by new values per second, not host count. Collection saturates first: pollers are a fixed pool, so checks are delayed and queue up. A proxy moves collection off the server and reaches networks it cannot address.

open as a page

What does the Datadog Agent do on a host, and how do infrastructure metrics, custom metrics and traces each reach it?

level: middleimportance: must knowfreq 68%
basics
~20 s

The Datadog Agent is one host process with several separate intake paths: it collects host metrics itself, runs integration checks that poll local services, accepts custom metrics pushed to its DogStatsD listener, and receives spans from in-process tracing libraries.

open as a page

In Datadog, what must be true of a service's telemetry for its metrics, traces and logs to link up?

level: seniorimportance: must knowfreq 62%
basics
~20 s

Every signal must carry the same env, service and version tags, set on the Agent, the tracing library and the log records. Log records must also carry the active trace and span identifiers to be joined to one request.

open as a page

How do you decide which Datadog monitors notify a human, and keep one failure from becoming a notification storm?

level: seniorimportance: should knowfreq 48%
basics
~10 s

A Datadog monitor only reaches a person if its notification message addresses a recipient handle; without one it records state silently. Storms are controlled by deliberate query grouping, composite monitors, and scheduled downtimes.

open as a page

How would you halve a Datadog bill dominated by custom metrics and indexed logs without losing incident coverage?

level: principalimportance: should knowfreq 54%
basics
~20 s

Attack the two billed units directly: custom metrics, counted per unique metric-name-and-tag-value combination, and indexed logs, billed separately from ingestion. Bound what tags may contain, aggregate before submitting, and decide after ingest which logs are worth indexing.

open as a page

In New Relic, why does it matter that events, metrics, logs and spans all land in one database?

level: juniorimportance: must knowfreq 66%
basics
~20 s

New Relic writes every signal - events, metrics, logs and spans - as timestamped, attributed records in one store, NRDB, and reads them all with one language, NRQL. Correlating a log with a span becomes a query, not an export.

open as a page

In New Relic, how does a NRQL query select, filter, facet and bucket results over time?

level: middleimportance: must knowfreq 74%
basics
~20 s

NRQL aggregates records at query time: SELECT chooses the aggregation, FROM the data type, WHERE filters, FACET splits results by an attribute's values, SINCE and UNTIL set the window, and TIMESERIES cuts that window into buckets.

open as a page

In a New Relic distributed trace, how do you attribute latency when only some services are instrumented?

level: seniorimportance: should knowfreq 51%
basics
~20 s

New Relic attributes time only to the spans it received. Work in an uninstrumented hop shows up as unexplained time inside its caller's span, and the platform cannot say whether that was network, queueing, or the missing service's own work.

open as a page

What is an entity in New Relic, and how does the platform assign telemetry to one?

level: seniorimportance: nice to knowfreq 24%
basics
~20 s

An entity in New Relic is anything the platform monitors and can name - a service, host, container, database or browser app - identified by an entity GUID. Incoming telemetry is matched to one by the identifying attributes it carries.

open as a page

What does the Dynatrace OneAgent do to a running application, and what does it still not see?

level: middleimportance: must knowfreq 58%
basics
~20 s

Dynatrace's OneAgent installs once per host, discovers the processes running there, and injects instrumentation into supported runtimes so traces, code-level timings and host metrics appear with no code change. It cannot supply business meaning or cover unsupported runtimes.

open as a page

How does Dynatrace build its live topology map, and why does automated root cause need one?

level: middleimportance: should knowfreq 41%
basics
~20 s

Dynatrace's agents continuously report hosts, processes, containers and the connections between them, and the platform keeps that as a time-versioned entity graph. Automated root cause needs it because causation is a walk over dependency edges, not a correlation of charts.

open as a page

What is a Dynatrace automated root-cause verdict actually worth on call, and when is it wrong?

level: seniorimportance: should knowfreq 34%
basics
~20 s

A Dynatrace verdict is a ranked hypothesis, not a diagnosis. From dependency edges, timing and change events it can say where a degradation started and what sits downstream. It stays blind to correctness bugs and causes outside the monitored estate.

open as a page

An application can export OTLP telemetry straight to a vendor backend. What does running an OpenTelemetry Collector in between buy you, and when is direct export from the SDK good enough?

level: juniorimportance: must knowfreq 55%
basics
~20 s

The Collector is a separate process that receives, processes and re-exports telemetry. It decouples apps from backends: batching, retry, enrichment, redaction, translation and routing become config in one place instead of code in every service.

open as a page

What does the W3C `traceparent` HTTP header carry, field by field, and what should a receiving service do with its sampled flag and with the accompanying `tracestate` header?

level: juniorimportance: must knowfreq 68%
basics
~20 s

traceparent is four hyphen-separated fields: version, 32-hex trace id, 16-hex span id of the caller, and 2-hex flags whose lowest bit means sampled. The receiver continues the same trace id, parents its span on that span id, and forwards tracestate unchanged apart from its own entry.

open as a page

What fields make up an OpenTelemetry span, and what do its `kind` and `status` fields actually mean to a backend that receives it?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A span carries identity (trace id, span id, parent span id), a name, a kind, start and end timestamps, attributes, events, links and a status. Kind tells the backend the span's role in a call (server, client, producer, consumer, internal); status is Unset, Ok or Error.

open as a page

OpenTelemetry offers a zero-code (agent-based) way to instrument an application and a code-based way using its API. What does each actually capture, what can neither capture, and how do the two combine inside one process?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Zero-code agents patch libraries at load time and emit spans for inbound requests, outbound calls, database and messaging clients with no source changes. Code-based instrumentation adds spans and attributes only your code knows. Both feed one SDK and nest.

open as a page

Your services are instrumented with the OpenTelemetry SDKs and the telemetry has to reach an observability backend. Make the case for exporting in OTLP (the OpenTelemetry Protocol) rather than using a backend-specific exporter inside the process, say when a native exporter is still the right call, and explain how you would choose between OTLP's gRPC and HTTP transports.

level: middleimportance: must knowfreq 55%
basics
~20 s

OTLP is the one exporter every OpenTelemetry SDK ships and it carries traces, metrics and logs, so the destination stays a config change. A native exporter compiles a vendor's protocol into the process, making a backend switch a redeploy. Use gRPC in-cluster, HTTP where proxies or browsers break it.

open as a page

In Jaeger, what path does a span take from an instrumented process into storage, and where can spans be dropped along it?

level: middleimportance: must knowfreq 60%
basics
~20 s

A span leaves the instrumented process through a bounded in-process reporter queue, reaches a Jaeger collector, and is written to storage, optionally via Kafka and an ingester. Every hop buffers, and every buffer drops spans when it fills.

open as a page

In Jaeger's trace waterfall, how do you tell a genuinely slow service from one merely waiting on a downstream call?

level: middleimportance: must knowfreq 66%
basics
~20 s

Compare a span's duration with the time its children cover. A parent's bar encloses its children, so the slow hop is the one with large self time — duration minus child coverage — not simply the longest bar on the screen.

open as a page

In Jaeger, what does serving sampling configuration centrally to clients solve, and what can it not solve?

level: seniorimportance: should knowfreq 42%
basics
~20 s

Jaeger clients poll the backend for their sampling rate, so per-service and per-operation rates change without a redeploy. Central configuration cannot pick traces by outcome, cannot bound spans per trace, and does nothing for a client that never fetches it.

open as a page

What access patterns must a Jaeger trace store support, and how do its Cassandra and Elasticsearch backends meet them?

level: seniorimportance: nice to knowfreq 26%
basics
~20 s

A Jaeger store must absorb write-heavy append-only ingest, answer point lookups by trace identifier, and serve ad-hoc search by service, operation, tag and duration. Cassandra needs extra index tables for that search; a search engine indexes everything at ingest instead.

open as a page

What does a Sentry SDK capture when an unhandled exception escapes, and what makes the stack trace readable in minified production code?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A Sentry SDK sends the exception type, message and stack trace plus the surrounding scope: request, user, tags, release and a trail of breadcrumbs. Minified or compiled frames stay unreadable until matching source maps or symbol files are uploaded.

open as a page

In Sentry, what does associating events with a release and its commit range buy you, and what does resolving an issue in the next release mean?

level: middleimportance: should knowfreq 50%
basics
~20 s

A release stamps events with the deployed version, so Sentry shows which deploy first produced an issue, marks a reappearance as a regression, and points at suspect commits. Resolving in the next release reopens the issue only if it recurs afterwards.

open as a page

In Sentry, how is an event's fingerprint derived to group it into an issue, and what goes wrong at each extreme?

level: seniorimportance: should knowfreq 54%
basics
~20 s

Sentry groups events by a fingerprint computed from the event itself -- normally the stack trace's in-app frames and the exception type, falling back to the exception value or message when there is no trace. Same fingerprint, same issue.

open as a page

In Sentry, what actually reduces the event volume a project ingests, and why should an issue alert fire on a new or regressed issue rather than on every event?

level: seniorimportance: nice to knowfreq 30%
basics
~20 s

Volume is cut in the SDK -- an error sample rate, ignore lists, a before-send hook that drops the event -- and at ingest by inbound filters, client-key rate limits and spike protection. Alert on new or regressed issues, not event counts.

open as a page

In an Elastic Stack, how do you choose between parsing logs in Filebeat, in Logstash, or in an Elasticsearch ingest pipeline?

level: middleimportance: must knowfreq 71%
basics
~20 s

Parse where you can afford the CPU. Filebeat ships cheaply on the host; a Logstash tier adds heavy enrichment and multi-destination routing but costs a tier to run; an Elasticsearch ingest pipeline parses inside the cluster.

open as a page

In a Logstash pipeline config, what do input, filter and output do, and what does a failed grok match cost?

level: middleimportance: should knowfreq 58%
basics
~20 s

Inputs receive events, filters transform them, outputs write them, in that order. A grok pattern that fails to match does not drop the event: Logstash tags it _grokparsefailure and ships it unparsed, after burning more CPU than a successful match.

open as a page

Several teams ship logs into one Elastic Stack cluster. How do you make a common field schema stick, and what breaks when two services send one field name with different types?

level: principalimportance: should knowfreq 41%
basics
~20 s

Agree a common vocabulary such as the Elastic Common Schema and enforce it in the shared shipping tier, not in documentation. When two services disagree on a field's type, one index rejects documents and searches across several go quietly wrong.

open as a page

What do Logstash's persistent queue and dead-letter queue each protect, and how does backpressure reach Filebeat?

level: seniorimportance: nice to knowfreq 27%
basics
~20 s

Logstash's persistent queue writes in-flight events to disk before filtering, so a restart or a downstream outage does not lose them. The dead-letter queue holds events Elasticsearch will never accept. Backpressure fills the queue and stalls Filebeat.

open as a page

What does a Kibana data view (formerly index pattern) bind together, and what changes once you nominate a time field?

level: juniorimportance: must knowfreq 64%
basics
~20 s

A Kibana data view names the indices a search may touch, usually a wildcard, and exposes their merged field list to Discover and dashboards. Nominating a time field is what makes the global time picker filter queries built on it.

open as a page

In Kibana, how does a filter pill differ from a KQL query typed in the query bar, and what pushes you off KQL?

level: middleimportance: should knowfreq 51%
basics
~20 s

A KQL query is one text expression for the whole search; a filter pill is a discrete condition you can negate, disable or pin. KQL has no regex, fuzziness or boosting; for those, switch the bar to Lucene syntax.

open as a page

In Kibana, what actually has to travel with a dashboard for it to work in another space or another cluster?

level: seniorimportance: nice to knowfreq 22%
basics
~20 s

A Kibana dashboard is a saved object that points at other saved objects by id: its panels and the data views behind them. Export it without those references and it imports cleanly but renders nothing.

open as a page

In Grafana Loki, what is a stream, what goes into the index, and what does not?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Loki indexes only the label set attached to each log line, never the line's text. A stream is one unique combination of label values; its lines are batched into compressed chunks kept in object storage.

open as a page

Why does adding a high-cardinality label to a Grafana Loki stream destroy its cost advantage?

level: seniorimportance: must knowfreq 64%
basics
~20 s

Each distinct label value creates a separate Loki stream with its own chunks. A high-cardinality label multiplies streams, so the index swells, ingesters hold thousands of barely-filled chunks in memory, and object storage fills with tiny, poorly-compressed files.

open as a page

In Loki's LogQL, why can only the stream selector make a query cheap?

level: middleimportance: should knowfreq 58%
basics
~20 s

The stream selector decides which chunks are read from object storage. Every filter, parser and metric stage after it runs over lines already fetched and decompressed, so they change what you see, not how much is read.

open as a page

Before Promtail ships a line to Loki, what must it decide about labels, multi-line entries and timestamps?

level: seniorimportance: nice to knowfreq 32%
basics
~20 s

The agent picks the stream labels through discovery and relabelling, joins continuation lines into one entry, and sets each entry's timestamp. All three are expensive to revisit: labels and timestamps are baked into chunks that are immutable once flushed.

open as a page

In Graylog, how does a raw log line reaching an input become a structured message, and where should field extraction happen?

level: middleimportance: must knowfreq 62%
basics
~20 s

A Graylog input binds a protocol to a port and its codec decodes each payload into a message with source, message and timestamp fields. Extractors on that input parse further; pipeline rules do so later, across inputs.

open as a page

In Graylog, how do streams and their connected pipelines decide a message's fate, and what happens if it matches no stream rule?

level: seniorimportance: should knowfreq 55%
basics
~20 s

Every Graylog message is tested against all stream rules and can join several streams, each with its own index set. Connected pipelines then run stage by stage and may enrich, re-route or drop it. Unmatched messages stay in the default stream.

open as a page

In Graylog, how would you give one team alerting on and read access to only its own logs?

level: seniorimportance: nice to knowfreq 30%
basics
~20 s

Route that team's messages into their own Graylog stream, grant read on just that stream through a role or an entity share, and define event definitions whose search is scoped to it. Alerting and access both hang off the stream boundary.

open as a page

In Splunk, what is decided about an event at index time versus search time, and which of those decisions are permanent?

level: middleimportance: must knowfreq 74%
basics
~20 s

At index time Splunk fixes an event's index, timestamp, host, source and sourcetype. Fields, aliases and lookups are worked out at search time and can be changed at any moment. The index-time decisions stick: correcting one means re-ingesting the data.

open as a page

In Splunk, how is an SPL search composed as a pipeline, and what separates a stage that streams from one that must see every result?

level: middleimportance: must knowfreq 68%
basics
~20 s

An SPL search starts with a retrieving stage that pulls events from Splunk indexes over a time range, then pipes them onward. Streaming stages act on one event at a time; transforming stages must gather the whole result set first.

open as a page

In Splunk, how do you choose between a universal forwarder, a heavy forwarder and a network input to get data in?

level: seniorimportance: should knowfreq 55%
basics
~20 s

A universal forwarder is a light agent shipping a host's raw data. A heavy forwarder parses events first, so it can mask, drop and route them before indexing. A network input accepts pushed data with no agent and no source-side buffer.

open as a page

In a large Splunk estate, when is precomputing an expensive search into a summary index or an accelerated report the wrong call?

level: principalimportance: nice to knowfreq 36%
basics
~20 s

A summary index stores a scheduled aggregation's results as small events you search instead of raw data; report acceleration has Splunk maintain equivalent precomputed data for a qualifying saved search. Both buy speed with freshness, storage and recurring load.

open as a page

What does each conventional log level from TRACE to FATAL mean, and what test separates WARN from ERROR?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Log levels are a contract about who must act. ERROR means something failed and a human should look; WARN means something recovered but degraded; INFO records state changes; DEBUG and TRACE are developer detail, off by default in production.

open as a page

What does a log backend that full-text indexes every line at ingest buy you, and what does it cost?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A full-text index built at ingest makes any word in any line searchable without declaring it first, and makes aggregations fast. You pay twice: CPU on the write path for every record, and index bytes stored alongside the logs.

open as a page

What does structured logging give a log consumer that a formatted message string cannot?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Structured logging emits each line as named, typed fields instead of one formatted sentence. Ingest no longer has to guess where each value starts and ends, and queries can filter, compare and aggregate on a field rather than searching raw text.

open as a page

What identifies one time series in a dimensional metrics system, and how does adding a metric label change the total count?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A time series is the metric name plus its complete label set; change any label value and it is a different series. Adding a label multiplies the series count by that label's distinct values rather than adding to it.

open as a page

What makes a metric a counter rather than a gauge, and why is a counter read as a rate over a window?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A counter only ever increases and returns to zero when its publishing process restarts; a gauge is a level that moves either way. Because a counter's absolute value depends on process uptime, you read its rate over a window.

open as a page