skip to content

Why add OpenLLMetry's @workflow and @task decorators when init already traces model calls?

level: middleimportance: must knowfreq 66%

answer

  1. auto spans are per library call, not per operation
  2. your business grouping exists only in code
  3. one parent span, everything nests under it
  4. workflow, task, agent, tool
  5. stable names become your latency series

basics

~20 s

Auto-instrumentation only sees individual library calls. The decorators create named parent spans for your own multi-step logic, so a retrieval plus three model calls appears as one named operation you can time, compare and filter on rather than four unrelated spans.

solid answer

~40 s

`Traceloop.init()` gives you one span per instrumented library call — a model call, a vector-store query — and nothing that says those belonged to the same business operation, because that grouping exists only in your code. The decorators from `traceloop.sdk.decorators` fix that: `@workflow` marks a top-level multi-step operation, `@task` marks a step inside it, and `@agent` and `@tool` are the agent-flavoured equivalents for an autonomous loop and for a function the model can invoke. Decorating a function opens a span for its duration, so every auto-instrumented call made inside nests under it. Each takes an optional `name` argument, defaulting to the function name, and that name is what you group and compare latency by in the backend. The result is a trace shaped like your application rather than like your dependency list.

code

python · 15 lines
python
from traceloop.sdk import Traceloop
from traceloop.sdk.decorators import task, workflow

Traceloop.init(app_name="support-bot")

@task(name="retrieve_docs")
def retrieve(question: str) -> list[str]:
    return ["policy.md", "faq.md"]

@workflow(name="answer_question")
def answer(question: str) -> str:
    docs = retrieve(question)
    return f"{question} answered from {len(docs)} sources"

answer("Where is my order?")

go deeper

for a junior

Know the four decorator names and that putting @workflow on a function creates a span that the model calls inside it nest under. Be able to write the import and apply one.

for a middle

Explain why auto-instrumentation alone leaves a flat list of library calls and how a parent span restores the shape of your application. Cover naming, nesting, and the difference between workflow/task and agent/tool.

for a senior

Talk about granularity and cost: which functions deserve spans, what span volume does to ingest bills, and how you debug an empty workflow span caused by work escaping the active context onto another thread.

for a principal

Own the vocabulary. Decide the workflow names teams use so latency and cost roll up per product capability across services, and treat those names as a stable contract that survives refactors rather than a per-team free-for-all.

## The gap the decorators fill Auto-instrumentation is honest but shallow: it records what a *library* did. A single user request in a RAG service might produce an embedding call, a vector-store query, a re-rank call and a generation call — four spans with no notion that they were one "answer_question" operation. Nothing in the OpenAI SDK knows your business steps, so nothing can name them for you. The decorators are how you supply that structure: ``` from traceloop.sdk.decorators import workflow, task, agent, tool ``` Decorating a function starts a span when it is entered and ends it when it returns, and because OpenTelemetry propagates the active context, every auto-instrumented call inside becomes a child of that span. You get a tree instead of a list. ## The four decorators They behave the same mechanically; the difference is the semantic label recorded on the span (the SDK writes a `traceloop.span.kind` attribute), which is what backends use to render and filter the trace. - **`@workflow`** — the outermost unit of work you would name in a product conversation: `answer_question`, `summarize_ticket`. One per request, typically. - **`@task`** — a step inside a workflow: `retrieve_docs`, `rerank`, `format_answer`. - **`@agent`** — an autonomous loop that decides its own next step, rather than a fixed pipeline. - **`@tool`** — a function the model can choose to call. Marking it as a tool makes "which tool did the agent pick, how often, and how long did it take" a query rather than a code read. Each accepts an optional `name`; omit it and the function name is used. Prefer explicit, stable names — they become the identity you compare across deploys, and renaming the Python function should not silently split your latency history in two. ## What it buys you operationally Once the workflow span exists, several things become possible that were not before: 1. **End-to-end latency per operation.** The workflow span's duration is the number users feel. The sum of model-call durations is not, because it excludes your own retrieval, parsing and retries. 2. **Attribution of cost and failure.** Token usage on child spans rolls up under a named parent, so you can say which product feature is expensive, not just which model is. 3. **Comparison across versions.** Filtering by workflow name gives a stable series to compare before and after a prompt or model change. 4. **Agent legibility.** For a tool-calling loop, the trace shows the sequence of tool spans the model chose, which is usually the fastest way to explain a wrong answer. ## Nesting and granularity Nest freely: a `@workflow` calling several `@task`s, one of which calls an `@agent` that calls `@tool`s. The practical limit is judgment, not the SDK — a span per trivial helper produces traces nobody reads and inflates ingest volume. A good rule is to decorate a function when you would want its duration on a dashboard or when it is a unit you might replace. Async code is handled: recent versions of the SDK detect coroutine functions and manage the span across the await rather than closing it at the first suspension. Generators and streaming responses are the case to think about, because the interesting work happens while the caller iterates, not when the function returns. ## What the decorators are not They do not enable tracing. Without `Traceloop.init()` there is no pipeline to export into. They do not add evaluation scores — judging whether the output was any good is a separate concern from recording that it happened. And they are not a replacement for the auto-instrumented LLM spans: the workflow span records that a step happened and how long it took; the model span underneath records which model, which parameters and how many tokens. ## Failure mode to know If a decorated function spawns work onto another thread or an executor without propagating context, the child spans can land outside the workflow span and appear as orphan traces. That is a context-propagation problem in the general tracing sense rather than an OpenLLMetry quirk, but it is the usual reason someone says "my workflow span is empty".

  • What distinguishes @agent and @tool from @workflow and @task in practice?
    Mechanically nothing — all four open a span around the function. The difference is the semantic kind recorded on the span, which drives how a backend renders and filters it. Use @agent for a loop that decides its own next step and @tool for a function the model can choose to call, so you can ask which tools an agent picked and how often, rather than reading code to find out.
  • How should you name decorated functions if you want to compare latency across releases?
    Pass an explicit, stable name argument rather than relying on the function name. The name is the grouping key in the backend, so renaming or refactoring the Python function would otherwise split one metric series into two and make a before/after comparison look like a regression appearing from nowhere.
  • Your @workflow span shows up with no children even though the function calls a model. What would you check?
    Most often the model call runs on another thread, in an executor, or in a task started without the active context, so its span attaches to a different parent or none. Check that the work happens inside the decorated call rather than being scheduled out of it. Also confirm the provider library is actually instrumented, since an uninstrumented client produces no child span at all.
  • Do the decorators work on async functions?
    Yes — recent SDK versions detect coroutine functions and keep the span open across awaits instead of ending it at the first suspension point. The case that still needs thought is streaming: if a function returns a generator, the interesting work happens while the caller iterates, so the span duration may not represent what you assume.

saying these in an interview costs you the question

  • Claiming decorators are required for any tracing at all
  • Decorating every helper function and drowning traces in noise
  • Treating summed model-call time as end-to-end latency
  • Assuming @tool computes a score or quality metric
  • Renaming decorated functions freely without realizing dashboards key on the name

context