Observability & Monitoring
Everything that tells you what production is doing: metric stores, dashboards, APM and distributed tracing, and log aggregation. Interviewers use this area to check that you can answer 'how would you know it broke, and how would you find out why' with named tools rather than hand-waving.
on this pageshowhide
explore
- Prometheus25 questions
- Metric Types and Data Model4 questions
- Scraping and Exporters4 questions
- Service Discovery4 questions
- PromQL5 questions
- Alerting and Alertmanager4 questions
- Storage and Operations4 questions
- Grafana32 questions
- Dashboards and Panels5 questions
- Data Sources5 questions
- Templating and Variables6 questions
- Unified Alerting6 questions
- Provisioning and Dashboards-as-Code5 questions
- LGTM Ecosystem Integration5 questions
- Zabbix4 questions
- Datadog4 questions
- New Relic4 questions
- Dynatrace3 questions
- OpenTelemetry34 questions
- Data Model & Signals6 questions
- Context & Propagation5 questions
- Instrumentation6 questions
- SDK & Resource Config6 questions
- Collector & Pipelines6 questions
- Backends & Ecosystem5 questions
- Jaeger4 questions
- Sentry4 questions
- Elastic Stack4 questions
- Kibana3 questions
- Loki4 questions
- Graylog3 questions
- Splunk4 questions
- Observability Foundations65 questions
- Telemetry Signals11 questions
- Metrics Concepts19 questions
- Logging Concepts19 questions
- Tracing and Sampling16 questions
- Backend Developerrole
- Cyber Security Expertrole
- Data Engineerrole
- DevOps / SRE Engineerrole
- DevSecOps Engineerrole
- Elasticsearchskill
- Forward Deployed Engineerrole
- Full Stack Developerrole
- Java Backend Developerrole
- Java SDETrole
- Kotlin Backend Developerrole
- Kubernetesskill
- MLOps Engineerrole
- Network Engineerrole
- PostgreSQL DBArole
- QA Engineerrole
- Software Architectrole
questions
197 · 15 sectionsWhat does a Prometheus alerting rule contain, and what does its `for` duration change about when it fires?
basics
~20 sA Prometheus alerting rule pairs a PromQL expression with an optional for duration, plus labels and annotations. The for duration requires the expression to hold continuously across that window, so the alert sits in pending before it becomes firing.
In Prometheus's text exposition format, what does one metric line contain, and what does the server keep from the `# HELP` and `# TYPE` lines?
basics
~20 sOne line carries a metric name, an optional label set in braces, a float64 value and an optional millisecond timestamp. The HELP comment describes the family and the TYPE comment declares counter, gauge, histogram, summary or untyped.
In PromQL, what does an instant vector selector return versus a range vector selector, and which functions require each?
basics
~20 sAn instant vector selector returns one sample per series at the query timestamp; a range vector selector adds a bracketed duration and returns every sample in that window. Functions like rate() only accept the bracketed range form.
In Prometheus, what happens on each scrape of one target, and what do scrape_interval and scrape_timeout bound?
basics
~20 sEach cycle Prometheus sends one HTTP GET to the target's metrics path, parses the body, and stores every sample under that scrape's start timestamp. scrape_interval sets how often it repeats; scrape_timeout bounds one request. Prometheus also records an up sample per scrape.
How does Prometheus find its scrape targets, and what does a service-discovery mechanism hand it besides a list of addresses?
basics
~20 sPrometheus builds its target list from the discovery blocks in each scrape job: a static list, files on disk, or a platform API such as Kubernetes or EC2. Each target arrives carrying read-only metadata labels describing its origin.
In Grafana, a single query returns several labelled series over time. Walk through how you would choose between the Time series, Stat, Gauge, Table and Heatmap panels, and explain what each one does to the query result before it renders.
basics
~20 sTime series plots every point over time. Stat and Gauge reduce each series to one number (last, mean, max); Gauge adds a bounded min/max scale. Table shows the raw rows. Heatmap shows distribution. Choose by trend vs single value vs distribution.
Describe what a data source is in Grafana, and trace what happens from the moment a dashboard panel runs its query until the result is rendered — including where the credentials live and what actually comes back over the wire.
basics
~20 sA data source is a saved connection (type, URL, auth, options) plus the plugin that knows how to query that system. The browser asks the Grafana server, which attaches the stored credentials, calls the backing system, and returns typed data frames the panel renders.
A Grafana dashboard needs a drop-down at the top that lists the environments actually present in the monitoring data, rather than a hardcoded list. Which Grafana variable type would you use, where does it get its options from, and what do the variable refresh settings 'Never', 'On dashboard load' and 'On time range change' control?
basics
~20 sUse a Query variable: Grafana runs a metadata query against the data source (e.g. a label-values lookup) and turns each returned row into a drop-down option. Refresh decides when that query re-runs — never (options are read from saved dashboard JSON), on every dashboard load, or additionally whenever the time range changes.
In Grafana's unified alerting, a Grafana-managed alert rule is defined as one or more data-source queries plus expressions rather than as a single threshold on a graph. Walk through how such a rule is evaluated and how it turns into individual firing alerts.
basics
~20 sEach query runs and returns labelled series. Server-side expressions then chain on those results — reduce a series to one number, do math, apply a threshold — and one expression is marked the rule's condition. Every distinct label set that satisfies it becomes its own alert instance.
Explain the states a Grafana-managed alert instance moves through — including its pending period — and what happens when the underlying query returns no data or fails outright.
basics
~20 sAn instance is Normal until the condition breaches, then Pending until it has breached continuously for the pending period, then Alerting, which notifies. Missing data and query failures are separate states whose handling is configurable per rule: treat as alerting, as normal, as a dedicated no-data alert, or keep the last state.
In Zabbix, what is an item, and how do passive and active agent checks differ?
basics
~20 sA Zabbix item is one configured metric on a host, named by a key. In a passive check the Zabbix server connects to the agent and asks for a value; in an active check the agent connects out and sends.
In Zabbix, what does a trigger expression evaluate, and what does a recovery expression add?
basics
~20 sA Zabbix trigger is a boolean expression over an item's stored history, built from functions like min or avg over a window rather than a threshold on the latest value. A separate recovery expression clears the problem on a looser condition.
In Zabbix, how do templates and low-level discovery keep hundreds of hosts configured?
basics
~20 sA Zabbix template is a reusable container of items, triggers and discovery rules linked to hosts; editing it changes every linked host at once. Low-level discovery expands prototypes into one item or trigger per entity found on each host.
In Zabbix, where does the server bottleneck as hosts are added, and what does a proxy change?
basics
~20 sA Zabbix server is sized by new values per second, not host count. Collection saturates first: pollers are a fixed pool, so checks are delayed and queue up. A proxy moves collection off the server and reaches networks it cannot address.
What does the Datadog Agent do on a host, and how do infrastructure metrics, custom metrics and traces each reach it?
basics
~20 sThe Datadog Agent is one host process with several separate intake paths: it collects host metrics itself, runs integration checks that poll local services, accepts custom metrics pushed to its DogStatsD listener, and receives spans from in-process tracing libraries.
In Datadog, what must be true of a service's telemetry for its metrics, traces and logs to link up?
basics
~20 sEvery signal must carry the same env, service and version tags, set on the Agent, the tracing library and the log records. Log records must also carry the active trace and span identifiers to be joined to one request.
How do you decide which Datadog monitors notify a human, and keep one failure from becoming a notification storm?
basics
~10 sA Datadog monitor only reaches a person if its notification message addresses a recipient handle; without one it records state silently. Storms are controlled by deliberate query grouping, composite monitors, and scheduled downtimes.
How would you halve a Datadog bill dominated by custom metrics and indexed logs without losing incident coverage?
basics
~20 sAttack the two billed units directly: custom metrics, counted per unique metric-name-and-tag-value combination, and indexed logs, billed separately from ingestion. Bound what tags may contain, aggregate before submitting, and decide after ingest which logs are worth indexing.
In New Relic, why does it matter that events, metrics, logs and spans all land in one database?
basics
~20 sNew Relic writes every signal - events, metrics, logs and spans - as timestamped, attributed records in one store, NRDB, and reads them all with one language, NRQL. Correlating a log with a span becomes a query, not an export.
In New Relic, how does a NRQL query select, filter, facet and bucket results over time?
basics
~20 sNRQL aggregates records at query time: SELECT chooses the aggregation, FROM the data type, WHERE filters, FACET splits results by an attribute's values, SINCE and UNTIL set the window, and TIMESERIES cuts that window into buckets.
In a New Relic distributed trace, how do you attribute latency when only some services are instrumented?
basics
~20 sNew Relic attributes time only to the spans it received. Work in an uninstrumented hop shows up as unexplained time inside its caller's span, and the platform cannot say whether that was network, queueing, or the missing service's own work.
What is an entity in New Relic, and how does the platform assign telemetry to one?
basics
~20 sAn entity in New Relic is anything the platform monitors and can name - a service, host, container, database or browser app - identified by an entity GUID. Incoming telemetry is matched to one by the identifying attributes it carries.
What does the Dynatrace OneAgent do to a running application, and what does it still not see?
basics
~20 sDynatrace's OneAgent installs once per host, discovers the processes running there, and injects instrumentation into supported runtimes so traces, code-level timings and host metrics appear with no code change. It cannot supply business meaning or cover unsupported runtimes.
How does Dynatrace build its live topology map, and why does automated root cause need one?
basics
~20 sDynatrace's agents continuously report hosts, processes, containers and the connections between them, and the platform keeps that as a time-versioned entity graph. Automated root cause needs it because causation is a walk over dependency edges, not a correlation of charts.
What is a Dynatrace automated root-cause verdict actually worth on call, and when is it wrong?
basics
~20 sA Dynatrace verdict is a ranked hypothesis, not a diagnosis. From dependency edges, timing and change events it can say where a degradation started and what sits downstream. It stays blind to correctness bugs and causes outside the monitored estate.
An application can export OTLP telemetry straight to a vendor backend. What does running an OpenTelemetry Collector in between buy you, and when is direct export from the SDK good enough?
basics
~20 sThe Collector is a separate process that receives, processes and re-exports telemetry. It decouples apps from backends: batching, retry, enrichment, redaction, translation and routing become config in one place instead of code in every service.
What does the W3C `traceparent` HTTP header carry, field by field, and what should a receiving service do with its sampled flag and with the accompanying `tracestate` header?
basics
~20 straceparent is four hyphen-separated fields: version, 32-hex trace id, 16-hex span id of the caller, and 2-hex flags whose lowest bit means sampled. The receiver continues the same trace id, parents its span on that span id, and forwards tracestate unchanged apart from its own entry.
What fields make up an OpenTelemetry span, and what do its `kind` and `status` fields actually mean to a backend that receives it?
basics
~20 sA span carries identity (trace id, span id, parent span id), a name, a kind, start and end timestamps, attributes, events, links and a status. Kind tells the backend the span's role in a call (server, client, producer, consumer, internal); status is Unset, Ok or Error.
OpenTelemetry offers a zero-code (agent-based) way to instrument an application and a code-based way using its API. What does each actually capture, what can neither capture, and how do the two combine inside one process?
basics
~20 sZero-code agents patch libraries at load time and emit spans for inbound requests, outbound calls, database and messaging clients with no source changes. Code-based instrumentation adds spans and attributes only your code knows. Both feed one SDK and nest.
Your services are instrumented with the OpenTelemetry SDKs and the telemetry has to reach an observability backend. Make the case for exporting in OTLP (the OpenTelemetry Protocol) rather than using a backend-specific exporter inside the process, say when a native exporter is still the right call, and explain how you would choose between OTLP's gRPC and HTTP transports.
basics
~20 sOTLP is the one exporter every OpenTelemetry SDK ships and it carries traces, metrics and logs, so the destination stays a config change. A native exporter compiles a vendor's protocol into the process, making a backend switch a redeploy. Use gRPC in-cluster, HTTP where proxies or browsers break it.
In Jaeger, what path does a span take from an instrumented process into storage, and where can spans be dropped along it?
basics
~20 sA span leaves the instrumented process through a bounded in-process reporter queue, reaches a Jaeger collector, and is written to storage, optionally via Kafka and an ingester. Every hop buffers, and every buffer drops spans when it fills.
In Jaeger's trace waterfall, how do you tell a genuinely slow service from one merely waiting on a downstream call?
basics
~20 sCompare a span's duration with the time its children cover. A parent's bar encloses its children, so the slow hop is the one with large self time — duration minus child coverage — not simply the longest bar on the screen.
In Jaeger, what does serving sampling configuration centrally to clients solve, and what can it not solve?
basics
~20 sJaeger clients poll the backend for their sampling rate, so per-service and per-operation rates change without a redeploy. Central configuration cannot pick traces by outcome, cannot bound spans per trace, and does nothing for a client that never fetches it.
What access patterns must a Jaeger trace store support, and how do its Cassandra and Elasticsearch backends meet them?
basics
~20 sA Jaeger store must absorb write-heavy append-only ingest, answer point lookups by trace identifier, and serve ad-hoc search by service, operation, tag and duration. Cassandra needs extra index tables for that search; a search engine indexes everything at ingest instead.
What does a Sentry SDK capture when an unhandled exception escapes, and what makes the stack trace readable in minified production code?
basics
~20 sA Sentry SDK sends the exception type, message and stack trace plus the surrounding scope: request, user, tags, release and a trail of breadcrumbs. Minified or compiled frames stay unreadable until matching source maps or symbol files are uploaded.
In Sentry, what does associating events with a release and its commit range buy you, and what does resolving an issue in the next release mean?
basics
~20 sA release stamps events with the deployed version, so Sentry shows which deploy first produced an issue, marks a reappearance as a regression, and points at suspect commits. Resolving in the next release reopens the issue only if it recurs afterwards.
In Sentry, how is an event's fingerprint derived to group it into an issue, and what goes wrong at each extreme?
basics
~20 sSentry groups events by a fingerprint computed from the event itself -- normally the stack trace's in-app frames and the exception type, falling back to the exception value or message when there is no trace. Same fingerprint, same issue.
In Sentry, what actually reduces the event volume a project ingests, and why should an issue alert fire on a new or regressed issue rather than on every event?
basics
~20 sVolume is cut in the SDK -- an error sample rate, ignore lists, a before-send hook that drops the event -- and at ingest by inbound filters, client-key rate limits and spike protection. Alert on new or regressed issues, not event counts.
In an Elastic Stack, how do you choose between parsing logs in Filebeat, in Logstash, or in an Elasticsearch ingest pipeline?
basics
~20 sParse where you can afford the CPU. Filebeat ships cheaply on the host; a Logstash tier adds heavy enrichment and multi-destination routing but costs a tier to run; an Elasticsearch ingest pipeline parses inside the cluster.
In a Logstash pipeline config, what do input, filter and output do, and what does a failed grok match cost?
basics
~20 sInputs receive events, filters transform them, outputs write them, in that order. A grok pattern that fails to match does not drop the event: Logstash tags it _grokparsefailure and ships it unparsed, after burning more CPU than a successful match.
Several teams ship logs into one Elastic Stack cluster. How do you make a common field schema stick, and what breaks when two services send one field name with different types?
basics
~20 sAgree a common vocabulary such as the Elastic Common Schema and enforce it in the shared shipping tier, not in documentation. When two services disagree on a field's type, one index rejects documents and searches across several go quietly wrong.
What do Logstash's persistent queue and dead-letter queue each protect, and how does backpressure reach Filebeat?
basics
~20 sLogstash's persistent queue writes in-flight events to disk before filtering, so a restart or a downstream outage does not lose them. The dead-letter queue holds events Elasticsearch will never accept. Backpressure fills the queue and stalls Filebeat.
What does a Kibana data view (formerly index pattern) bind together, and what changes once you nominate a time field?
basics
~20 sA Kibana data view names the indices a search may touch, usually a wildcard, and exposes their merged field list to Discover and dashboards. Nominating a time field is what makes the global time picker filter queries built on it.
In Kibana, how does a filter pill differ from a KQL query typed in the query bar, and what pushes you off KQL?
basics
~20 sA KQL query is one text expression for the whole search; a filter pill is a discrete condition you can negate, disable or pin. KQL has no regex, fuzziness or boosting; for those, switch the bar to Lucene syntax.
In Kibana, what actually has to travel with a dashboard for it to work in another space or another cluster?
basics
~20 sA Kibana dashboard is a saved object that points at other saved objects by id: its panels and the data views behind them. Export it without those references and it imports cleanly but renders nothing.
In Grafana Loki, what is a stream, what goes into the index, and what does not?
basics
~20 sLoki indexes only the label set attached to each log line, never the line's text. A stream is one unique combination of label values; its lines are batched into compressed chunks kept in object storage.
Why does adding a high-cardinality label to a Grafana Loki stream destroy its cost advantage?
basics
~20 sEach distinct label value creates a separate Loki stream with its own chunks. A high-cardinality label multiplies streams, so the index swells, ingesters hold thousands of barely-filled chunks in memory, and object storage fills with tiny, poorly-compressed files.
In Loki's LogQL, why can only the stream selector make a query cheap?
basics
~20 sThe stream selector decides which chunks are read from object storage. Every filter, parser and metric stage after it runs over lines already fetched and decompressed, so they change what you see, not how much is read.
Before Promtail ships a line to Loki, what must it decide about labels, multi-line entries and timestamps?
basics
~20 sThe agent picks the stream labels through discovery and relabelling, joins continuation lines into one entry, and sets each entry's timestamp. All three are expensive to revisit: labels and timestamps are baked into chunks that are immutable once flushed.
In Graylog, how does a raw log line reaching an input become a structured message, and where should field extraction happen?
basics
~20 sA Graylog input binds a protocol to a port and its codec decodes each payload into a message with source, message and timestamp fields. Extractors on that input parse further; pipeline rules do so later, across inputs.
In Graylog, how do streams and their connected pipelines decide a message's fate, and what happens if it matches no stream rule?
basics
~20 sEvery Graylog message is tested against all stream rules and can join several streams, each with its own index set. Connected pipelines then run stage by stage and may enrich, re-route or drop it. Unmatched messages stay in the default stream.
In Graylog, how would you give one team alerting on and read access to only its own logs?
basics
~20 sRoute that team's messages into their own Graylog stream, grant read on just that stream through a role or an entity share, and define event definitions whose search is scoped to it. Alerting and access both hang off the stream boundary.
In Splunk, what is decided about an event at index time versus search time, and which of those decisions are permanent?
basics
~20 sAt index time Splunk fixes an event's index, timestamp, host, source and sourcetype. Fields, aliases and lookups are worked out at search time and can be changed at any moment. The index-time decisions stick: correcting one means re-ingesting the data.
In Splunk, how is an SPL search composed as a pipeline, and what separates a stage that streams from one that must see every result?
basics
~20 sAn SPL search starts with a retrieving stage that pulls events from Splunk indexes over a time range, then pipes them onward. Streaming stages act on one event at a time; transforming stages must gather the whole result set first.
In Splunk, how do you choose between a universal forwarder, a heavy forwarder and a network input to get data in?
basics
~20 sA universal forwarder is a light agent shipping a host's raw data. A heavy forwarder parses events first, so it can mask, drop and route them before indexing. A network input accepts pushed data with no agent and no source-side buffer.
In a large Splunk estate, when is precomputing an expensive search into a summary index or an accelerated report the wrong call?
basics
~20 sA summary index stores a scheduled aggregation's results as small events you search instead of raw data; report acceleration has Splunk maintain equivalent precomputed data for a qualifying saved search. Both buy speed with freshness, storage and recurring load.
What does each conventional log level from TRACE to FATAL mean, and what test separates WARN from ERROR?
basics
~20 sLog levels are a contract about who must act. ERROR means something failed and a human should look; WARN means something recovered but degraded; INFO records state changes; DEBUG and TRACE are developer detail, off by default in production.
What does a log backend that full-text indexes every line at ingest buy you, and what does it cost?
basics
~20 sA full-text index built at ingest makes any word in any line searchable without declaring it first, and makes aggregations fast. You pay twice: CPU on the write path for every record, and index bytes stored alongside the logs.
What does structured logging give a log consumer that a formatted message string cannot?
basics
~20 sStructured logging emits each line as named, typed fields instead of one formatted sentence. Ingest no longer has to guess where each value starts and ends, and queries can filter, compare and aggregate on a field rather than searching raw text.
What identifies one time series in a dimensional metrics system, and how does adding a metric label change the total count?
basics
~20 sA time series is the metric name plus its complete label set; change any label value and it is a different series. Adding a label multiplies the series count by that label's distinct values rather than adding to it.
What makes a metric a counter rather than a gauge, and why is a counter read as a rate over a window?
basics
~20 sA counter only ever increases and returns to zero when its publishing process restarts; a gauge is a level that moves either way. Because a counter's absolute value depends on process uptime, you read its rate over a window.