A service mesh is often sold on giving you 'free' observability — uniform metrics, distributed tracing, and access logs for every service without code changes. What exactly does the mesh actually provide here, what's the catch, and what production failure mode commonly trips up teams relying on it?
answer
- per-hop metrics free, full trace needs header propagation
- golden signals: rate, errors, duration
- trace context header propagated across hops
- app must forward headers on outbound calls
- sampling rate <100% can miss the exact failing request
basics
~20 sThe mesh's sidecars automatically record how long calls took and whether they succeeded, for every service, without any app code — but they can't fully trace a request's journey unless the app helps pass along a few tracing headers.
solid answer
~50 sBecause every request passes through a sidecar on both ends, the mesh can uniformly emit L7 metrics (request rate, error rate, latency histograms) and access logs for every service with zero application instrumentation, which is a big win over each team hand-rolling its own metrics. What it can't do alone is stitch together a full distributed trace across a multi-hop call: each sidecar only sees and reports its own hop, so connecting hop A-to-B-to-C into one trace requires the application to propagate a few tracing headers from an inbound request onto any outbound calls it makes — a small but real code change the mesh cannot do for you. A common production trip-up is a service that drops those headers, silently breaking trace continuity for every request passing through it, while dashboards still look fine because per-hop metrics keep working.
go deeper
Should know the mesh gives metrics and logs automatically for every service without app changes.
Should know tracing needs the app to forward some headers, even if they can't name the header format.
Should explain the per-hop-span vs end-to-end-trace distinction precisely and describe the silent-breakage failure mode.
Should reason about sampling strategy trade-offs, incident-time tracing gaps, and how to make header-propagation compliance auditable/enforced across a large polyglot fleet.
## Why the mesh is so well placed for this A sidecar mesh sits on every network hop in the system, which puts it in a genuinely privileged position for observability: it sees every request's method, path, response code, and latency, for every service, automatically, the moment sidecar injection is turned on. This is real value: - **no team has to remember** to add a metrics library; - **no language is left uninstrumented**; - and the data is **consistent in shape** across the whole fleet because it comes out of the same proxy implementation. ## The three pillars Concretely, this breaks into three pillars. - **Metrics:** the sidecar emits per-request counters and histograms — request count by response code, request duration percentiles, bytes in/out — labeled by source and destination service, which a metrics backend like Prometheus scrapes and dashboards typically surface as the 'golden signals' (rate, errors, duration) for every service pair in the mesh. - **Access logs:** each sidecar can emit a structured log line per request it handles, giving a request-level audit trail without the app writing anything. - **Distributed tracing:** each sidecar, on seeing a request, generates or forwards trace spans representing this hop, this long, this status, and exports them to a tracing backend. ## The catch: one sidecar, one hop Here is the catch, and it's the single most important nuance in mesh observability: a sidecar only ever sees one hop of a request — it has no inherent knowledge that the request it's proxying from A to B is part of a longer chain that started at a client and continues on to C, D, E. Two separate sidecar-generated spans for A-to-B and B-to-C will only be recognized as part of the same end-to-end trace if they share a common trace ID, and propagating that trace ID across the hop from B's inbound request to B's outbound calls is something only the application code inside B can do — by reading the incoming tracing headers off the request it received and copying them onto any downstream calls it makes. The mesh cannot inject this continuity into a request's business logic layer because it has no visibility into which outbound call, if any, a given inbound request is causally driving; that causal link only exists inside the application. So distributed tracing genuinely is a **mesh-plus-app-code collaboration**: the mesh generates the spans and exports them, but the application has to carry the trace context header through. ## What goes wrong This produces a specific, common production failure mode: a service written by a team unaware of, or that forgot about, this requirement doesn't propagate the tracing headers from its inbound request onto its own outbound calls. The result isn't an error or an alert — everything keeps working, every service's per-hop metrics keep looking healthy, dashboards show normal latency and error rates. What silently breaks is **trace continuity**: instead of one coherent end-to-end trace showing a request's full path through five services, the tracing backend shows two or more disconnected trace fragments with no link between them. An on-call engineer debugging a slow end-to-end request during an incident opens the tracing UI expecting to see the whole path and instead sees the trace mysteriously stop at the service that dropped the headers, making it look like that service is the origin of a new, unrelated request rather than the middle of a chain — actively misleading during exactly the moment an incident demands accurate tracing. ## When the trace you need was never taken A second, related failure mode is sampling misconfiguration: because tracing every single request at high volume is expensive (storage and processing cost), meshes are typically configured with a trace sampling rate well below 100%. During a low-traffic-volume incident affecting a specific customer or a specific rare code path, the odds that the one problematic request was actually sampled and traced can be low, leaving engineers without a trace for the exact request they need, and having to fall back to correlating logs and metrics by timestamp instead — slower and less precise. ## Where it shows up A concrete real-world pattern: Istio's own documentation and observability guides explicitly call out that applications must propagate trace headers for tracing to work correctly across multiple hops, listing this as a required application-side responsibility distinct from anything sidecar injection provides automatically — one of the few pieces of 'you still have to write a little code' guidance in an otherwise code-free mesh adoption story.
- If a service's dashboards show completely normal latency and error rates during an incident, does that rule out that service as the cause of a broken distributed trace?No — per-hop metrics and trace continuity are separate concerns. A service can have perfectly healthy request/error/latency metrics while still silently failing to propagate tracing headers to its downstream calls, which breaks the end-to-end trace without affecting any metric a mesh dashboard would flag.
- What is the practical difference between what a mesh gives you 'for free' in metrics versus what it gives you for free in tracing?Metrics are genuinely zero-code: every sidecar independently emits request rate/error/latency data for its own hop with no application involvement needed at all. Tracing is only partially free — the mesh generates and exports spans automatically, but connecting those spans into one coherent end-to-end trace requires the application to propagate trace context headers from inbound to outbound calls, which is real code every service team must implement.
- Why might an engineer investigating a specific customer's slow request during an incident find no trace for it at all, even though tracing is enabled mesh-wide?Trace sampling is usually configured well below 100% to control storage and processing cost, so most individual requests are never traced at all. If that specific customer's request wasn't one of the sampled ones, no trace exists for it regardless of whether tracing infrastructure is otherwise working correctly, forcing the engineer to fall back on logs and metrics correlation.
Like security cameras at every doorway in a building recording who passed through each door — great for seeing each individual crossing, but reconstructing one person's full path through the building requires someone to pass along a note saying 'this is the same visit' at each doorway, or the recordings look like unrelated separate visits.
saying these in an interview costs you the question
- believes the mesh gives fully automatic distributed tracing with zero app involvement
- doesn't know the difference between per-hop metrics and end-to-end trace continuity
- can't explain why trace context headers need application-level propagation
- assumes 100% of requests get traced by default
- treats healthy dashboards as proof a service isn't breaking anything