skip to content

What fields make up an OpenTelemetry span, and what do its `kind` and `status` fields actually mean to a backend that receives it?

level: juniorimportance: must knowfreq 58%

answer

  1. 16-byte trace id, 8-byte span id, parent span id
  2. Client and server each create their own span — ids are not shared
  3. Kind: INTERNAL / SERVER / CLIENT / PRODUCER / CONSUMER drives the service graph
  4. Status UNSET is the default; OK is an explicit assertion
  5. Resource → InstrumentationScope → Span envelope

basics

~20 s

A span carries identity (trace id, span id, parent span id), a name, a kind, start and end timestamps, attributes, events, links and a status. Kind tells the backend the span's role in a call (server, client, producer, consumer, internal); status is Unset, Ok or Error.

solid answer

~50 s

A span is one timed operation. Its identity fields are a 16-byte `trace_id`, an 8-byte `span_id`, the `parent_span_id` (empty for a root), plus trace flags and tracestate carried from the incoming context. Its content is `name` (a low-cardinality operation label, not a URL with an id in it), `start_time_unix_nano` and `end_time_unix_nano`, `attributes` as typed key/values, `events` (timestamped annotations inside the span), `links` (references to other spans) and `status`. `kind` is one of INTERNAL, SERVER, CLIENT, PRODUCER, CONSUMER, and it is how a backend knows which spans are remote-call boundaries — that is what makes service topology and correct client-versus-server latency possible. `status` is `UNSET` by default, `ERROR` when the operation failed, and `OK` only when application code explicitly asserts success; instrumentation should normally leave it unset rather than mark everything OK. Spans are nested under an instrumentation scope and a resource, which say what library emitted them and what entity produced them.

code

text · 7 lines
text
service A                        service B
span(id=aa11, kind=CLIENT)  ---->  span(id=bb22, kind=SERVER, parent=aa11)
  start 10.000                       start 10.004
  end   10.061                       end   10.058
  duration 61ms                      duration 54ms
                       difference = network + queueing, visible only because
                       both sides emit their own span

go deeper

for a junior

List the fields confidently — ids, name, kind, timestamps, attributes, events, links, status — and give the five span kinds.

for a middle

Explain why kind matters to topology and latency, and get the UNSET-versus-OK convention and the exception-event trap right.

for a senior

Add the two-spans-per-call consequence (client/server latency delta, clock skew), naming cardinality, and the resource/scope envelope's operational uses.

for a principal

Frame span shape as an interface contract: names and attribute keys are consumed by dashboards and alerts across teams, so semantic conventions and low-cardinality naming are governance, not style.

## What a span represents A span is a single named operation with a start and an end: a request handled, a database call made, a message consumed, a chunk of local work. A trace is the set of spans sharing a `trace_id`, connected into a tree by `parent_span_id`. Everything else in the model hangs off those two ideas. ## Identity fields - `trace_id` — 16 bytes, globally unique, identical across every span of the trace. All-zero is invalid. - `span_id` — 8 bytes, unique within the trace. All-zero is invalid. - `parent_span_id` — the parent's span id; empty on a root span. - `trace_flags` — an 8-bit field whose bit 0 is the sampled flag. - `trace_state` — vendor-specific key/value state travelling with the trace. How those values reach a remote process is the propagation concern and is out of scope here; what matters for the data model is that a receiving service creates its **own** span with a **new** span id whose parent is the caller's span. OpenTelemetry does not share one span id between client and server the way some older systems did — you get two spans per remote call, and their duration difference is exactly the network and queuing time the client experienced but the server did not. ## Descriptive fields **`name`** should describe a class of operation, not an instance: `GET /orders/{id}`, not `GET /orders/8891`. Names are what backends group by, so embedding identifiers destroys the grouping and inflates the backend's index. **Timestamps** are `start_time_unix_nano` and `end_time_unix_nano`, absolute nanoseconds since the Unix epoch. Duration is derived, not stored. Since the two endpoints of a remote call are timestamped by different machines, clock skew is visible in the data and a child can appear to start before its parent — backends usually clamp this for display, but it is a real artifact. **`attributes`** are typed key/value pairs: string, bool, int64, double, byte array, or a homogeneous array of those. Semantic conventions define standard keys so that different libraries describe the same thing the same way. Spans tolerate high-cardinality attributes — a user id or an order id on a span is fine and often exactly what makes an investigation possible — which is emphatically not true of metrics. **`events`** are timestamped annotations inside the span's lifetime, each with a name and attributes: an exception thrown, a cache miss, a retry attempt. **`links`** reference other spans without a parent-child relationship. Both have dedicated treatment in their own right; the point here is that they are span fields, and that the model records how many of each were dropped when limits were hit (`dropped_attributes_count`, `dropped_events_count`, `dropped_links_count`) so silent truncation is at least visible. ## SpanKind, and why it is load-bearing `kind` takes one of five values: - `SERVER` — handling an inbound synchronous request. Usually the root of the service's work. - `CLIENT` — making an outbound synchronous call. - `PRODUCER` — enqueuing a message for asynchronous consumption. - `CONSUMER` — processing such a message. - `INTERNAL` — everything else; the default. This is not decoration. A backend uses it to reconstruct the service graph (CLIENT→SERVER pairs are the edges), to compute latency correctly on each side of a call, and to know that a PRODUCER/CONSUMER pair may be separated by an arbitrary delay rather than nested in time. Instrumentation that marks every span INTERNAL still produces a valid trace and a useless topology. ## Status, and the Unset/Ok subtlety `status` has a `code` of `UNSET`, `OK` or `ERROR`, plus a `message` that is meaningful only for errors. The convention that trips people up: **`UNSET` is the correct default and instrumentation libraries should not set `OK`**. `OK` means an application has explicitly decided the operation succeeded, overriding any inference. A backend therefore treats UNSET as "no failure reported", not as "unknown, probably broken". A second trap: recording an exception as an event does **not** set the status. A caught-and-handled exception is a recorded event on a span whose status stays unset; an operation that actually failed needs the status set explicitly. And what counts as an error depends on the kind — a 404 is usually not a server-side error but may be a client-side one. ## The envelope: resource and scope Spans are not shipped bare. In the OTLP structure they nest as **Resource → InstrumentationScope → Span**. The *resource* identifies the entity producing telemetry (service name, instance, host, region) and is shared by everything that entity emits; the *scope* identifies the instrumentation that created the span (library name, version, schema URL), which is how you tell a framework's automatic spans from your hand-written ones and how you disable a noisy library without touching the rest. ## What a good answer emphasises Identity, timing, description, and the two fields interviewers actually probe: kind, because topology depends on it, and status, because the Unset/Ok distinction and the exception-event trap are where real instrumentation goes wrong.

  • Why shouldn't a client library set span status to OK on every successful call?
    Because OK is defined as an explicit application-level assertion that overrides inference, not as "no exception was thrown". If instrumentation stamps OK everywhere, a backend or an operator can no longer distinguish spans where someone deliberately judged the outcome from spans where nothing was evaluated. The correct default is UNSET, with ERROR set when the operation genuinely failed.
  • What is wrong with naming a span `GET /orders/8891`?
    Span names are a grouping key. Embedding an identifier makes every request its own operation name, so aggregate views, latency percentiles per operation and any backend index over names explode in cardinality and stop being useful. The identifier belongs in an attribute, where high cardinality is expected and searchable; the name should be the route template, `GET /orders/{id}`.

saying these in an interview costs you the question

  • Believing the client and server share one span id for a remote call
  • Leaving every span as INTERNAL and then wondering why the service map is empty
  • Assuming a recorded exception event automatically sets span status to ERROR
  • Putting request identifiers in the span name instead of an attribute
  • Thinking the span stores a duration field rather than start and end timestamps

context