skip to content

How does monitoring differ from observability, and what must a system already be emitting for the second to work?

level: middleimportance: should knowfreq 58%

answer

  1. One is prediction, the other is exploration
  2. Known failure modes versus unanticipated questions
  3. Decided at emit time, not query time
  4. Can you answer without shipping code?

basics

~20 s

Monitoring answers questions you predicted, from views and conditions built in advance for known failure modes. Observability is the property that lets you answer questions nobody anticipated, by slicing retained per-request detail along dimensions you did not choose beforehand.

solid answer

~50 s

**Monitoring** watches for failure modes someone already thought of: conditions and views built in advance. It is excellent at "is the thing I worried about happening again?" and useless the first time a system fails in a way nobody modelled. **Observability** is a property of the system rather than a product you buy: can you answer a new question about production without shipping code first? That requires the system to have already emitted enough detail — per-request records carrying the dimensions you might later want to filter and group by, a shared identifier that links one request's records across services, and retention long enough to cover the question. Pre-aggregated counters cannot be un-aggregated later, and a dimension nobody recorded cannot be recovered. The checkable test is simple: how often does answering a production question here require a deploy?

go deeper

for a junior

Recall the one-line contrast: monitoring watches for problems someone predicted, observability is being able to ask a new question of production. Have one example of a question no pre-built view would answer.

for a middle

Explain the mechanics behind the slogan: which properties of the emitted data make arbitrary post-hoc slicing possible, and why folding events into a total is a one-way transformation that no later query can undo.

for a senior

Demonstrate it on a real system you operated: a question you could not answer, what was missing at emit time, and what you changed. Be honest that the fix usually costs a deploy and a wait, not a purchase.

for a principal

Own the emit-time budget. Detail and retention cost real money, so decide which dimensions are worth carrying on every request forever, which are worth carrying on a sampled subset, and how you keep that decision from being made by accident.

## Two different jobs **Monitoring** is the practice of watching a system for failure modes someone already thought of. Its artefacts are built in advance: a condition on a counter, a view of the four or five curves a team believes describe health, a check that runs every minute. Monitoring is judged by how quickly it tells you *that* something known has gone wrong. **Observability** is not a second set of artefacts and not a product you buy. It is a property of the running system: can you answer a question about production that nobody anticipated, without shipping code first? It is judged by how often the answer to "why is it doing that?" requires a deploy before you can even look. The two are complementary and a real system needs both. Monitoring is how you notice. Observability is how you explain. ## Known and unknown failure modes | | Monitoring | Observability | |---|---|---| | Question shape | Predicted: "is X happening again?" | Novel: "why is this subset behaving like that?" | | Built | Before the incident | During it, out of data already collected | | Fails when | The failure mode is new | The detail was never emitted, or has aged out | | Cost driver | How many things are watched | How much detail is retained per event | | Success measure | Time to notice | Time to explain, without a deploy | The failure that pages you is usually a **known unknown**: you knew it could happen, you did not know when, and someone built for it. The one that ruins an afternoon is the **unknown unknown** — a combination of a client version, one tenant's data shape and a downstream timeout that nobody would have built a view for. Pre-built views cannot cover it by construction, because covering it would have meant predicting it. ## What a system must already be emitting Observability is decided at emit time, not at query time. Four things must already be true before a novel question is answerable: 1. **Per-request records exist, with their dimensions attached.** Not merely "a request failed", but which route, which tenant, which client version, which upstream was called, how long it took and how it ended. A dimension nobody recorded cannot be recovered afterwards at any price. 2. **A shared identifier links the records of one request across services.** Without it you can find one hop and never the rest, and every cross-service question degrades into timestamp guesswork. 3. **The store can group by a field you did not choose in advance.** Pre-aggregation is a one-way transformation: once individual events are folded into a total, no query unfolds them. Arbitrary post-hoc slicing needs the underlying records to still be there. 4. **Retention covers the age of the question.** A question asked on Friday about Monday is unanswerable if records live for 24 hours, however rich they were. None of those four is a purchase. All of them are decisions taken in application code and in the collection pipeline, usually long before the incident that needs them. ## A worked example An orchestra-touring logistics platform emits roughly 18 GB of log lines a day across `tour-routing`, `freight-booking`, `visa-docs` and `venue-slotting`, and about seventy per cent of that volume comes from a single team. The monitoring is genuinely good: failure ratio and latency per endpoint, with a condition on each. Then a question arrives that nobody built for: *since the client update, are tours whose freight crosses two customs jurisdictions failing more often than single-jurisdiction ones?* Nothing pre-built answers it. It is answerable only if the per-request records for `freight-booking` already carry the jurisdiction count and the client version as fields, and only if those records still exist for the period in question. If they do, it is a five-minute query and the system was observable *for that question*. If they do not, the honest answer is "we will know in three days" — because the fix is a code change, a deploy, and then waiting for data to accumulate. That gap, not the tooling, is the entire distinction. ## Where the word becomes a slogan Two claims deserve pushback. The first is that observability simply *is* the collection of several signals: collecting all of them says nothing about whether the detail you will need is inside them, and a service that emits five fields per request is equally unobservable in every product on the market. The second is that a platform confers it: a platform can only slice what was sent to it. The useful, checkable version of the question is the one to end an answer on: **how often does answering a production question here require shipping code?** If the answer is "most of the time", the system is monitored and not observable, whatever is installed in front of it.

  • Name a concrete question your system cannot answer today, and say exactly what would have to change for it to answer it.
    A good example is "are requests from one client version failing more often on a specific route?" It becomes answerable when the client version and the route are fields on the per-request record, when those records are retained long enough to cover the period asked about, and when the store can group by a field chosen at query time. All three are emit-time and pipeline decisions, not query-time ones.
  • Does adding distributed tracing automatically make a service observable?
    No. Tracing gives you the request's structure across services, which is a large gain, but you can only filter and group by attributes that were actually attached to the spans. A trace carrying nothing but service names and durations answers "where did the time go" and cannot answer "which tenants were affected". The property comes from the detail recorded, not from the signal being present.
  • Is monitoring obsolete once a system is observable?
    No — they answer different questions. Pre-built views and conditions are how you notice a known failure quickly and without a human deciding to look. Exploration over retained detail is how you explain something new. A team with only the second finds out late; a team with only the first finds out fast and then cannot say why.

Monitoring is the set of warning lights someone chose to fit on the dashboard. Observability is whether the car records enough about itself that a mechanic can diagnose a fault nobody fitted a light for.

saying these in an interview costs you the question

  • Says observability just means collecting several signals
  • Claims a vendor platform makes a system observable by itself
  • Thinks more pre-built views add up to observability
  • Believes pre-aggregated counters can be sliced afterwards
  • Names nothing that must change at emit time
  • Treats the distinction as pure marketing with no content