skip to content

Latency on a production Atlas cluster spiked after a release — which Atlas tools do you open, and in what order?

level: seniorimportance: should knowfreq 52%

answer

  1. Four tools, each answering a different question
  2. One shows now, one shows the last day
  3. Scanned-versus-returned is the deploy-regression signal
  4. Recommendations still need verifying before you build
  5. Fast-but-frequent queries stay invisible

basics

~20 s

Start with the Real-Time Performance Panel to see what is running now, use Metrics for the shape and timing of the regression, then the Query Profiler for the slow operations in the last day, and Performance Advisor for ranked index suggestions — validating any suggestion with explain before shipping it.

solid answer

~50 s

Work outside-in. The **Real-Time Performance Panel** shows live operations and the hottest collections, and lets you kill a runaway operation — that is the triage step. The **Metrics** tab establishes when the regression started and which resource moved: opcounters, connections, disk latency, replication lag, and especially **Query Targeting**, the ratio of documents and index keys scanned to documents returned. A ratio that jumps at deploy time is the signature of a query that lost its index. The **Query Profiler** then lists slow operations from roughly the last day with duration, `keysExamined`, `docsExamined` and shape, so you can name the offending namespace and predicate. **Performance Advisor** turns that into ranked index recommendations, and also flags redundant or unused indexes and schema anti-patterns. Confirm any candidate index with `explain()` on the real query before creating it, and remember the profiler only sees operations that crossed the slow-operation threshold.

go deeper

for a junior

Be able to name the tools Atlas gives you — Metrics, Query Profiler, Performance Advisor, the Real-Time Performance Panel — and say which one suggests indexes.

for a middle

Explain what feeds each tool, especially that the profiler and advisor only see operations above the slow-operation threshold, and what the query-targeting ratio measures.

for a senior

Walk a real incident end to end: stabilise, correlate the change point with the release, isolate the operation, verify a candidate index before creating it, and leave an alert behind.

for a principal

Own the observability standard: which alert conditions are mandatory on production projects, how they route into on-call, and how index changes are reviewed rather than applied from a recommendation panel.

## The four tools and what each is for Atlas ships a layered set of observability tools, and part of the interview is knowing which answers which question. **Real-Time Performance Panel (RTPP)** — a live view of the cluster: current operations, network throughput, hottest collections, and the ability to terminate a specific operation. This is the triage instrument. If a single unbounded aggregation is pinning CPU right now, RTPP is where you see it and where you kill it. It shows the present, not the past. **Metrics** — time-series charts per node. Opcounters, query targeting, connections, disk IOPS and latency, cache usage, replication lag and oplog window. This is where you answer *when* and *what kind*: did the regression start exactly at the deploy, is it CPU or disk, did connections climb, is a secondary falling behind. Correlating the change point with a deployment timestamp is often the whole diagnosis. **Query Profiler** — slow operations captured over roughly the last day, plotted by duration, each with operation type, namespace, duration, `keysExamined`, `docsExamined`, response length and the query shape. This converts "the cluster is slow" into "this predicate on this collection is scanning hundreds of thousands of documents". **Performance Advisor** — analysis on top of those slow operations, producing ranked index recommendations with an estimated impact, plus suggestions to drop unused or redundant indexes and warnings about schema anti-patterns such as unbounded arrays or excessive collection counts. It is a recommendation engine, not a measurement. ## The order that works 1. **Is it happening right now?** Open RTPP. If yes, identify the hot namespace and, if something is clearly pathological, kill it to stabilise. Stabilise before you investigate. 2. **When did it start and what moved?** Metrics. Line the change point up against the release. Look at query targeting first for a post-deploy latency regression — a jump in scanned-to-returned is the single most diagnostic signal available, because it says the workload is now reading far more than it returns. 3. **Which operations?** Query Profiler over the window containing the change point. Sort by duration, look for shapes that were not there yesterday, and read `docsExamined` against `nreturned`. 4. **What is the fix?** Performance Advisor for the ranked index suggestion, then verify it against the actual query with `explain()` before creating anything. Index creation on a large production collection is itself an operational event, so it deserves a check rather than a click. ## Blind spots to name The most valuable thing to say at senior level is where these tools do not look. - **The profiler only sees slow operations.** Operations that never exceed the slow-operation threshold are invisible to both the Query Profiler and Performance Advisor. A query taking 40 ms but executed thousands of times per second can dominate CPU and never appear. Atlas adjusts that threshold automatically by default, and you can pin it to a fixed value when you need to see faster operations — at the cost of more logging. When the profiler is empty but the cluster is busy, this is usually why, and the answer is to look at opcounters and RTPP's hottest collections instead. - **Performance Advisor optimises the queries it saw.** It will happily recommend an index for a query that should not be running at all. It has no view of your application; a suggestion that adds a large index to fix a report nobody reads is worse than deleting the report. - **Indexes are not free.** Every added index costs write throughput, cache residency and disk. Performance Advisor's drop-index suggestions exist because teams accumulate them; treat additions with the same scrutiny. - **Not every latency problem is a query.** Disk latency saturation, connection storms from a misconfigured pool, a secondary lagging and dragging majority writes, or a tier that is simply undersized all show up in Metrics rather than the profiler. ## Closing the loop with alerts Diagnosis after the fact is the weaker half. Atlas alerting exists so the same signals page you. The conditions worth configuring on a production cluster are the ones that precede user-visible failure: **Query Targeting: Scanned Objects / Returned** above a threshold (the missing-index alarm), **Replication Oplog Window** dropping below a safe number of hours (initial sync and recovery risk), disk space percentage used, connections as a percentage of the configured limit, and system CPU. Alerts are configured per project and can be routed to email, SMS, Slack, PagerDuty, Opsgenie, Datadog, Microsoft Teams, or a generic webhook, so they land in the same on-call flow as everything else. Configuring these once, in code through the Admin API or Terraform, is what turns the four tools above from an incident-time scramble into a routine. The answer an interviewer wants is this shape: stabilise with RTPP, locate in time with Metrics, name the operation with the Query Profiler, get a candidate from Performance Advisor, verify with `explain()`, then leave behind an alert so the next occurrence announces itself.

  • Performance Advisor shows nothing, but CPU is pinned. What is going on?
    Both Performance Advisor and the Query Profiler are fed by operations that crossed the slow-operation threshold. A query that is individually fast but executed at very high rate never qualifies, yet can dominate CPU. Look at opcounters and the hottest collections in the Real-Time Performance Panel, and consider pinning a lower threshold temporarily to make those operations visible.
  • Which Atlas alert best warns you that a query has lost its index?
    Query Targeting: Scanned Objects / Returned. It fires when the cluster reads far more documents or index keys than it returns, which is exactly what a collection scan looks like in aggregate. It is the highest-signal early alarm for a deploy that changed a predicate or dropped an index, well before latency becomes user-visible.
  • Why not just create every index Performance Advisor recommends?
    Each index costs write throughput, cache space and disk, and the advisor only optimises the queries it happened to observe — including ones that should not be running. Verify the candidate with explain() on the real query, weigh it against the write path, and check the advisor's own drop-index suggestions so you are not accumulating overlapping indexes.

saying these in an interview costs you the question

  • Jumps straight to adding every suggested index
  • Assumes an empty Query Profiler means no expensive queries
  • Confuses the live performance panel with historical metrics
  • Ignores the scanned-to-returned ratio after a deployment
  • Treats index recommendations as free to apply

context