What is a wide structured event, and which questions can it answer that a pre-aggregated metric cannot?
answer
- One row, not many lines
- Emitted once when the work finishes
- Every dimension you might slice by later
- Aggregation is a one-way door
- High-cardinality fields cost bytes, not series
basics
~20 sA wide structured event is one record emitted per unit of work — a request, a job, a message — carrying every dimension you might later slice by. Pre-aggregation discards those rows, so questions nobody anticipated can no longer be asked.
solid answer
~50 sA **wide structured event** is a single record emitted once per unit of work — a request, a job run, a consumed message — holding everything known about it when it finishes: identifiers, timings, sizes, outcome, and the situational fields only that unit of work knows, such as build version, region, tenant or feature-flag state. A pre-aggregated metric collapses many units of work into counts and bucketed distributions at emit time, keyed by a handful of dimensions chosen in advance. That answers *how much* and *how often* cheaply, but the individual rows are gone, so any grouping nobody predicted is unanswerable. Wide events keep the rows, so the grouping is chosen at query time — including on high-cardinality fields such as a customer id or a build hash that a metric series could not carry.
code
json · 16 lines{
"unit": "request",
"endpoint": "/vaults/maturation-schedule",
"duration_ms": 1873,
"status": 500,
"error_class": "page_limit_exceeded",
"build": "2026.09.03-4471",
"region": "eu-central",
"node": "vault-node-06",
"tenant": "alpine-creamery",
"cache": "miss",
"rows_scanned": 41208,
"downstream_retries": 2,
"queue_wait_ms": 611,
"flag_batch_reader": true
}go deeper
Be able to say what the record is: one row emitted when a unit of work finishes, holding many fields about that one piece of work. Contrast it with scattering several short log lines through the same handler.
Explain the mechanics — a context object accumulating fields through the handler, flushed once at the end — and why grouping on a field is possible afterwards only if the individual rows still exist.
Show that you have debugged with them: grouping by a dimension nobody predicted, correlating two fields inside the same row, and pulling the outlier row itself. Say how you keep the volume affordable without losing errors.
Own the tradeoff across an estate: which questions justify keeping rows at all, where sampling weights are mandatory so numbers stay reconstructable, and how you stop teams from answering a cheap always-on metric question with an expensive event query.
## One record per unit of work A **unit of work** is whatever the service treats as one indivisible job: an inbound HTTP request, a consumed queue message, a scheduled recompute, one step of a migration. A wide structured event is a single machine-readable record emitted **once, when that unit of work finishes**, so it can carry the outcome and the total duration alongside everything learned along the way. The word that matters is *wide*, and it means columns, not verbosity. A handler that writes eleven log lines as it goes is producing narrow records: each knows a fragment of the story and none knows the whole. A wide event is the opposite shape — one row, dozens to a few hundred fields, every field describing the same unit of work. Teams usually build one by threading a mutable context object through the handler, adding fields as facts become known, and flushing it once at the end. Fields that typically appear: - **Identity and correlation** — the request or job id, and the ids that let the row be joined to other signals. - **Timing** — total duration plus the phases that matter: queue wait, downstream call, serialization. - **Volume** — rows scanned, bytes returned, items in the batch, retries attempted. - **Outcome** — status, error class, whether a fallback fired. - **Deployment context** — build version, region, node, process instance. - **Routing and business context** — endpoint, tenant, plan tier, client version, feature-flag state. ## What pre-aggregation permanently destroys A counter or a bucketed distribution is a decision taken at instrumentation time: you pick the dimensions, the process adds up every unit of work that shares them, and it ships the totals. That is exactly why metrics are cheap enough to leave on for everything. It is also why they cannot be re-asked. | Question during an incident | Pre-aggregated series | Stored wide events | |---|---|---| | How many failed in the last hour? | Yes — the thing it exists for | Yes, by counting rows | | Did the failures cluster on one build? | Only if build was already a dimension | Yes, group by that field | | Are they all one customer? | Almost never — too many distinct values | Yes | | Are the slow ones *also* cache misses *and* retries? | No | Yes, filter on both fields | | Show me one of the bad ones | No — the rows do not exist | Yes, read the row | The fourth line is the one candidates miss. Aggregation does not merely lose resolution; it loses the **correlation between fields inside the same unit of work**. Two metric series can both rise in the same minute with nobody able to say whether the same requests were involved. In a store of rows that is one filter. ## The questions only the rows answer 1. **Grouping by a dimension nobody chose in advance** — the field is on the row whether or not you predicted needing it. 2. **Grouping by a high-cardinality dimension** — a customer id, a build hash, a device model. 3. **Correlating several fields within one unit of work** — slow *and* cache-missed *and* on the new build. 4. **Finding and reading the outlier rows themselves**, rather than a bucket that contains them. A worked example. On a cheese-ageing inventory platform, p99 on the maturation-schedule endpoint climbs from 340 ms to 1.9 s. The dashboards say latency and errors are up and nothing more. Grouping the stored events by build shows 96% of the slow rows carry a single build. Filtering to that build and grouping by tenant shows they are one large tenant whose vault count has crossed the point where a batch reader stops fitting in one page. No pre-chosen set of metric dimensions was going to reach that, and nobody would have added *tenant* and *build hash* as metric dimensions in advance. ## The cardinality asymmetry, and what it costs The reason a wide event may hold a customer id is arithmetic. Adding a field to a row costs **bytes on that row** — linear in volume. Adding a dimension to a metric **multiplies the number of stored series** by the number of distinct values, and multiplies again for the next dimension. That asymmetry, not fashion, is why high-cardinality context lives on events. The bill arrives instead as volume: one row per unit of work does not shrink as traffic grows. The usual levers are - **sampling with a retained weight** — keep one in *N* ordinary rows and store the weight so counts and rates can still be reconstructed, while keeping every error and every slow row; - **tiered retention** — full fidelity for days, a derived rollup for longer; - **field discipline** — dropping fields no query has ever touched. ## What interviewers listen for - Distinguishing *wide* (one row, many columns) from *verbose* (many rows). - Naming the cardinality asymmetry out loud rather than hand-waving at cost. - Treating aggregation as a one-way door, not as lossy-but-recoverable. - Not overclaiming: wide events do not replace always-on metrics, which answer 'how much, right now, for everything' at a per-query cost events cannot match.
- One row per unit of work does not get cheaper as traffic grows. How do you keep that affordable?Sample the ordinary rows and store the sampling weight with each kept row so counts and rates can still be reconstructed, while keeping every error and every slow row unconditionally. Then tier retention — full fidelity for a short window, a derived rollup beyond it — and periodically drop fields no query has touched.
- Why can a wide event carry a customer id when a metric series usually cannot?Cost scales differently. A field on a row costs bytes, so it grows linearly with the number of rows. A dimension on a metric multiplies the number of stored series by its number of distinct values, and multiplies again with the next dimension, so a field with millions of values becomes millions of series to store, index and query.
A pre-aggregated metric is the totals row at the bottom of a spreadsheet; a wide event is the spreadsheet row itself. Once you keep only the totals, you can never re-sort by a column nobody thought to total.
saying these in an interview costs you the question
- Says a wide event is just a log line with more text in it
- Thinks emitting more log lines per request achieves the same thing
- Believes any dimension can be added to a metric for free
- Cannot say what aggregation permanently destroys
- Claims wide events remove the need for metrics entirely