You are sending OpenTelemetry log records to a log backend such as Loki. Explain how the OTel log record maps onto that backend's storage model, what makes logs correlate with traces, and where the cardinality trap is.
answer
- record = body, severity, resource attrs, record attrs, trace/span id
- few low-cardinality stream labels + non-indexed metadata
- Loki 3.x: OTLP ingest + structured metadata tier
- trace_id as a LABEL = one stream per trace = disaster
- sampled-out trace id links to nothing — expected
basics
~20 sAn OTel log record has a body, a severity, resource and record attributes, and optional trace and span ids. A label-indexed store like Loki can only afford a handful of low-cardinality labels, so a small subset of resource attributes becomes stream labels and everything else must ride as non-indexed structured metadata or in the body. Correlation works because the trace and span ids travel on the record itself.
solid answer
~60 sThe OTel log data model is: timestamp(s), severity number/text, a body (string or structured), resource attributes (who emitted it), record attributes (about this event), and — crucially — trace id, span id and trace flags when the record was produced inside an active span. A label-indexed store splits that into two tiers. **Stream labels** define the index and must stay low cardinality: service, namespace, environment, maybe level. Everything else should land as **non-indexed structured metadata** (Loki 3.x) or inside the body, where it is searchable but not indexed. Loki 3.x accepts OTLP directly and applies exactly that split, with configurable promotion of chosen attributes to labels. **Correlation** is automatic if the logging path is trace-aware: the ids on the record let a UI pivot from a log line to the trace and back. Without them, correlation degrades to timestamp-and-service guessing. **The trap** is promoting a high-cardinality attribute — pod name on a churny deployment, user id, request id — to a label. Each distinct value creates a stream; stream explosion drives ingest and query cost far more than log volume does.
go deeper
List the log record's fields and say that trace and span ids on the record are what link logs to traces.
Explain the split between indexed stream labels and non-indexed payload, and give one concrete high-cardinality field that must not be a label.
Reason about stream/chunk economics, the collector-side allow-list, sampled-out traces, observed versus event timestamp, and per-version backend support.
Set cross-signal cardinality policy in one place, budget logs as the highest-volume signal, and define what correlation guarantees the platform actually offers.
## What an OTel log record actually contains Unlike traces and metrics, logs were not invented by OpenTelemetry — the project's job was to define a model that *existing* logging can be bridged into. A log record carries: - **Timestamp** and, separately, **observed timestamp** (when the pipeline saw it), which matters when records are read from files or a queue long after they were written. - **Severity number** (a normalised numeric scale) and **severity text** (the original level string). The numeric scale is what lets records from different logging systems be compared. - **Body** — the message; may be a plain string or structured data. - **Attributes** — key/values about this specific event. - **Resource** — key/values about the emitting entity: service name, instance, host, container, region. Shared by every record from that process. - **Trace id, span id, trace flags** — populated when the record was created while a span was active. Because of the bridging goal, the usual production shape is not 'the app calls an OTel logging API directly' but 'the app logs as it always did, and a bridge/appender or a collector receiver converts those records into the OTel model'. Either way the destination sees the same structure. ## Mapping onto a label-indexed store A store like Loki does not index the log text. It indexes **stream labels** — a small set of key/values that identify a log stream — and stores the raw lines in chunks per stream. Query performance therefore comes from selecting few streams and then grepping within them. That forces a three-way split of an OTel record: 1. **Stream labels** — a handful of stable, low-cardinality resource attributes: service name, namespace, environment, cluster, sometimes severity. These are what you filter on first. 2. **Structured metadata** — Loki 3.x introduced a non-indexed key/value tier attached to each line. This is where the remaining attributes belong: trace id, span id, pod name, request id, tenant. They are queryable and displayable without creating a stream per value. 3. **Body** — the message. Loki 3.x exposes an OTLP ingest path and does this split automatically, with configuration to choose which attributes are promoted to labels. Other backends differ: a document store indexes far more fields and its constraint is mapping explosion and shard count rather than stream count; a columnar store cares about column count and compression. The principle is the same everywhere — *decide deliberately which fields are index keys*, because the store's cost model lives there. ## Correlation with traces The payoff for putting logs through the OTel model is that trace id and span id are first-class fields on the record rather than something parsed out of a message with a regex. Two directions of navigation follow: - **Trace → logs**: given a trace id, query the log store for records carrying it. Cheap if trace id is structured metadata; ruinous if it is a stream label (one stream per trace). - **Logs → trace**: a log line displays its trace id as a link into the tracing backend. For this to work the ids must actually be populated, which requires the logging path to be aware of the active context at the moment the record is created. When they are missing, look at whether logging happens on a different thread/task than the instrumented work, whether an async boundary dropped the context, or whether the bridge simply is not wired. A related gotcha: if trace sampling drops the trace but logs are retained, a log line can carry a trace id that resolves to nothing in the tracing backend. That is expected — the sampled flag on the record tells you which case you are in — but it confuses people, so decide whether the UI should hide links for unsampled traces. ## The cardinality trap, concretely Every distinct combination of stream labels creates a stream, and each stream has its own chunk lifecycle: open chunks held in memory, flushed on size or age. Promote `pod` to a label on a deployment that rolls twice a day and you multiply streams by the number of pods ever seen in the retention window. Promote `request_id` and you have created one stream per request, which is a pathological case that will take the ingest path down. Symptoms are memory pressure and 'too many streams'/rate-limit rejections on ingest, and slow queries because the index is enormous. The cure is policy, applied in the collector before the data reaches the store: an explicit allow-list of attributes permitted to become labels, with everything else demoted to structured metadata or dropped. This is exactly the same governance problem as metric label cardinality, and it should be solved in the same place. ## Signal maturity and volume Logs were the last OTel signal to stabilise, so backend support and SDK ergonomics lag traces and metrics; check per-version rather than assuming. Volume is also the differentiator: logs are usually the largest signal by bytes, which makes the collector-side decisions — batching, compression, attribute pruning, sampling of high-volume debug records — matter more here than for traces. ## Interview framing Describe the record's fields, then the index-versus-payload split the destination forces, then correlation via the ids on the record, then the cardinality trap with a concrete example and the collector-side fix. Mention that logs stabilised last so support varies.
- A log line carries a trace id, but the tracing backend has no such trace. Is that a bug?Usually not — it means the trace was not sampled for export while the log was retained. The record's trace flags tell you whether the trace was sampled, so the UI can decide whether to render a link. It becomes a bug only if the sampling decision and the log-retention decision were meant to be consistent, which is a pipeline design choice rather than a defect in propagation.
- Why not just index every attribute so all logs are searchable by any field?Because the index is the store's cost centre. In a stream-label model each distinct label combination is a stream with its own chunk lifecycle, so an unbounded attribute produces unbounded streams, memory pressure and ingest rejections. Non-indexed structured metadata keeps the field searchable within a selected stream at a fraction of the cost, which is the right tier for anything with high cardinality.
- Where should attribute pruning happen — the SDK or the pipeline?Pruning belongs in the pipeline so the policy is centrally owned and changeable without redeploying services; the SDK should avoid emitting genuinely useless fields in the first place. Doing it only at the SDK means every team can bypass the policy, and doing it only at the store means you already paid the network and ingest cost.
saying these in an interview costs you the question
- Promoting trace id or request id to a stream label so 'logs are searchable by trace'.
- Assuming correlation works from timestamps and service name alone without the ids on the record.
- Treating the OTel logs signal as equally mature and equally supported as traces.
- Ignoring the observed-timestamp field when records are read from files or queues long after they were written.
- Doing attribute pruning only in the application, so pipeline-side policy cannot be enforced.