skip to content

What access patterns must a Jaeger trace store support, and how do its Cassandra and Elasticsearch backends meet them?

level: seniorimportance: nice to knowfreq 26%

answer

  1. Three very different read shapes
  2. The trace identifier is the point lookup
  3. Ad-hoc search is the awkward one
  4. Retention differs sharply by backend
  5. Read-time assembly, never write-time joins

basics

~20 s

A Jaeger store must absorb write-heavy append-only ingest, answer point lookups by trace identifier, and serve ad-hoc search by service, operation, tag and duration. Cassandra needs extra index tables for that search; a search engine indexes everything at ingest instead.

solid answer

~50 s

Three patterns, pulling different ways. Ingest is **append-only and write-heavy**: spans arrive individually, out of order, are never updated, and are joined into a trace only at read time. Opening a trace is a **point lookup by trace identifier**. Searching is **multi-predicate**: service, operation, tag key and value, a duration range and a time window. The third pattern decides the backend. A wide-column store like Cassandra serves queries only against keys declared in advance, so Jaeger maintains additional index structures by service, by service and operation, by duration and by tag — cheap writes, but real write amplification and search limited to what was indexed. A search engine indexes each span as it lands, so expressive search comes for free while ingest costs more CPU and a heavier cluster. Retention differs too: row-level expiry versus deleting whole per-day indices.

go deeper

for a junior

Recall that Jaeger does not keep traces in the application or in memory: spans go to a real backing store, commonly Cassandra or Elasticsearch, and the UI reads them back from there rather than from the services themselves.

for a middle

Explain the three access patterns and why searching by tag and duration is harder than fetching by trace identifier. Be able to say that a trace is assembled when it is read, not when it is written.

for a senior

Show the operational side: write amplification from index structures, index overhead in a sizing calculation, and the difference between expiring rows and deleting whole per-day indices when you change retention.

for a principal

Own the cost model. Argue where retention, sampling rate and backend choice trade against each other for the estate you actually run, and be ready to justify a pluggable backend over the one your team already knows how to operate.

## The three access patterns Jaeger's storage layer looks unusual until you write down what actually reads and writes it. There are exactly three shapes, and they pull in different directions. 1. **Write-heavy, append-only ingest.** Spans arrive continuously, individually, out of order, from many processes. They are never updated and never deleted individually. Two spans of the same trace may arrive seconds apart and are not joined on the way in — a trace is a *read-time* construct. 2. **Point lookup by trace identifier.** Opening a trace in the UI means "give me every span carrying this trace ID". This is the cheapest possible query if the trace ID is the key, and ruinous if it is not. 3. **Multi-predicate search.** The search screen filters by service, operation name, tag key and value, a minimum and maximum duration, and a time range, then returns a bounded list of matching traces. This is an ad-hoc query over several independent dimensions — and it is the pattern that decides which backends are viable. Pattern 3 is the awkward one. Patterns 1 and 2 are satisfied by almost any key-value store; a system that also has to answer "spans from `order-api`, operation `reserveStock`, tag `error=true`, longer than 250 ms, in the last hour" needs real indexing. ## What that demands of the backends | | Wide-column store (Cassandra) | Search engine (Elasticsearch / OpenSearch) | |---|---|---| | Ingest | Very cheap; append-friendly writes are its strength | More expensive; every span is indexed as it lands | | Lookup by trace ID | Direct, keyed read — the natural case | A term lookup on the trace ID field | | Ad-hoc multi-field search | Not native: Jaeger writes extra index tables per searchable dimension | Native: what an inverted index exists to do | | Cost of a new search dimension | Another write on every span | Usually none — already indexed | | Retention | Time-to-live on the written rows | Delete whole indices, which are written per day | The row that matters is the third. A wide-column store answers queries against keys you declared in advance, so Jaeger cannot simply ask it for "all spans with this tag longer than that duration" — it maintains additional index structures keyed by service, by service and operation, by duration, and by tag, and writes into them alongside every span. That is real **write amplification**: one span produces several writes, storage grows faster than the raw span volume suggests, and tag search is coarser than it looks because only indexed tags are searchable. In exchange you get a backend that shrugs at ingest rates a search cluster would struggle with. A search engine inverts the trade. Search is natural and expressive because the whole document is indexed, but you pay for that indexing on every single span at ingest time, and the cluster you must operate to absorb the write rate is heavier than the equivalent wide-column ring. Jaeger writes its data into per-day indices, which makes retention a matter of deleting whole indices on a schedule rather than expiring individual rows — cheap and predictable, at the price of day-granularity retention. ## Retention, which is where the bill lands Traces are the highest-volume signal most estates produce, and nobody keeps them long. The two mechanisms differ in an operationally important way: - **Row-level expiry** (the wide-column path) is set when the span is written, so changing retention only affects data written afterwards, and expired data still occupies space until the store reclaims it. - **Index-level deletion** (the search-engine path) drops a whole day at once, which reclaims space immediately and predictably, but means your retention granularity is a day and a single oversized day is deleted whole. Either way, the sizing question is the same one: sampled traces per second, multiplied by mean spans per trace, multiplied by mean span size, multiplied by retention — and then multiplied again by whatever the index overhead is. In a seed-catalogue ordering estate where one team contributes seventy per cent of the span volume, that multiplication is also the argument for setting that team's sampling rate differently from everyone else's, because storage is the constraint the sampling policy is really being tuned against. ## The rest of the options Jaeger also ships small backends that are not production storage and should not be pitched as such: an in-memory store used by the single-binary distribution, and an embedded local key-value store useful for a single node or a development box. Neither is horizontally scalable. Beyond that, Jaeger exposes a storage interface over gRPC so a backend it does not ship can be plugged in without forking it — which is how vendor-specific and object-storage-backed trace stores attach. The judgement an interviewer is really testing: can you name the read patterns before naming a product? A candidate who says "Elasticsearch, because search" without being able to say what is being searched has memorised a deployment, not a design.

  • Why does Jaeger's wide-column backend write more than one row per span?
    Because that class of store answers queries only against keys declared up front, and Jaeger's search filters on service, operation, duration and tags. Each searchable dimension therefore needs its own index structure written alongside the span. The result is write amplification: storage grows faster than raw span volume, and only the dimensions Jaeger chose to index are searchable at all.
  • How would you size a Jaeger trace store before deploying it?
    Multiply sampled traces per second by mean spans per trace by mean span size by retention, then add index overhead, which on a search-engine backend is a large fraction of the raw figure. Measure the first two from a short full-sampling window rather than guessing. The sum almost always sends you back to revisit the sampling policy rather than buying more disk.
  • When is Jaeger's in-memory or single-node embedded storage the right choice?
    For local development, demos, and CI where a trace only needs to outlive the test that produced it. Neither is horizontally scalable and neither survives a restart in any useful way, so treating one as production storage means losing traces exactly when an incident makes you want them. For anything shared or long-lived, use a real backend.

saying these in an interview costs you the question

  • Names a backend before naming the read patterns
  • Thinks Jaeger joins spans into traces at write time
  • Assumes any tag is searchable on every backend
  • Ignores index overhead when sizing storage
  • Treats the in-memory store as production storage
  • Believes spans are updated as later spans arrive