How do Helicone's session headers group an agent's many calls into one trace?
answer
- flat rows lose the run's shape
- one id shared by every call
- a name makes runs findable
- a slash path builds the tree
- nothing is inferred, you propagate it
basics
~20 sSending the same Helicone-Session-Id on every call of one run ties those requests together, Helicone-Session-Name labels the run, and Helicone-Session-Path places each call in a hierarchy so a multi-step agent reads as a tree instead of scattered rows.
solid answer
~50 sBy default Helicone logs each LLM call as an independent row, which is fine for a single completion and useless for an agent that made fourteen calls to answer one question. The session headers restore that structure. You generate one identifier per run and send it as `Helicone-Session-Id` on every call belonging to the run; `Helicone-Session-Name` gives the run a human-readable label so sessions are findable; and `Helicone-Session-Path` describes where the call sits in the run's hierarchy, using a slash-delimited path such as `/plan` or `/plan/search`. Helicone then renders the run as a single session you can walk through in order, with the whole run's cost and latency aggregated. The important operational detail is that nothing is inferred: the headers are per-request, so your code must thread the session id and the current path through every layer, including nested tool calls and any background steps. A call that forgets the header simply falls outside the session, which is the usual reason a trace looks incomplete.
code
python · 27 linesimport os
import uuid
from openai import OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url="https://oai.helicone.ai/v1",
default_headers={"Helicone-Auth": f"Bearer {os.environ['HELICONE_API_KEY']}"},
)
session_id = str(uuid.uuid4())
def agent_call(prompt: str, path: str):
return client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
extra_headers={
"Helicone-Session-Id": session_id,
"Helicone-Session-Name": "support-agent",
"Helicone-Session-Path": path,
},
)
agent_call("Plan how to answer the ticket.", "/plan")
agent_call("Search the knowledge base.", "/plan/search")go deeper
Know that the same Helicone-Session-Id on several calls groups them into one run, and that a session name makes that run findable in the dashboard.
Explain what Helicone-Session-Path adds — hierarchy, so the run reads as a tree — and that all three headers are per-request rather than inferred.
Show how you propagate the session id through nested tool calls, retries and cross-service hops, and use whole-run cost and call count to catch agent loops.
Treat run correlation as a design requirement for agent systems: decide the identifier's lifetime and where it is minted, so tracing survives queues, service boundaries and retries.
## Why a flat log is not enough One prompt and one completion fits comfortably in a table row. An agent does not. A single user question can fan out into a planning call, several tool-argument calls, a retrieval step, a synthesis call and a self-check, possibly across multiple services. Logged flatly, those are unrelated rows that happen to share a timestamp range — you cannot tell which belong together, you cannot total the run's real cost, and you cannot replay the reasoning. Helicone's answer is a set of request headers that let the caller declare the structure the proxy cannot infer. ## The three headers - **`Helicone-Session-Id`** — the correlation key. Any requests sharing this value belong to the same session. Generate one identifier per run (a UUID is typical) at the entry point. - **`Helicone-Session-Name`** — a human label for the run, such as `support-agent` or `weekly-report`. Sessions are otherwise identified by an opaque id, and a name is what makes them findable and comparable. - **`Helicone-Session-Path`** — the position of this call in the run, expressed as a slash-delimited path. A top-level step might be `/plan`; a call made while executing that step might be `/plan/search`. The nesting implied by the path is what turns a list into a tree. All three are ordinary request headers, so they are set per call — typically as extra headers on the individual request, not as client defaults, since the path (and often the id) changes between calls. ## Threading them through the code This is where implementations go wrong. There is no ambient context: if a header is not on the request, the call is not in the session. In practice that means: - Create the session id once, at the boundary where the run begins, and put it somewhere the whole run can read — a request-scoped context object, a contextvar, an explicit parameter. - Maintain a current path as you descend, appending a segment when you enter a sub-step, so calls made inside that step carry the deeper path. - Propagate across process boundaries yourself. If a step is handed to a worker or another service, the session id must travel with the job payload or it will start its own orphan trace. - Cover retries and background work. A retried call that loses the header creates a row that looks unrelated to the run it belongs to. ## What you get With the headers in place, a session view gives you the things flat rows cannot: - **Whole-run cost.** The number that matters for an agent is cost per completed task, not per call — and it is often startling the first time you see it, because the planning and tool-argument calls outnumber the visible ones. - **Ordered replay.** Walk the calls in sequence to see where a run went wrong: a bad plan, a tool that returned junk, a synthesis step that ignored its context. - **Shape comparison.** Two runs of the same agent that differ wildly in call count usually indicate a loop that is not terminating, which is far easier to see as a tree than as rows. - **Latency decomposition.** Which step dominates wall-clock time for the user. ## Sessions versus properties versus user id These three metadata mechanisms are complementary and get confused: - A **session** groups the calls of one run. High cardinality is expected and correct here — one value per run. - A **custom property** (`Helicone-Property-*`) is a low-cardinality descriptive dimension for aggregation: feature, environment, version. - **`Helicone-User-Id`** identifies who the run was for, and is what per-user cost views key on. A well-instrumented agent call carries all three: a user id, a couple of properties, and the session trio. Reaching for a property when you meant a session is the common mistake, and it produces a dimension with one member per run that no dashboard can use. ## Mode notes Grouping is a property of the log record rather than something interception provides, so sessions work under async logging too — the identifiers are part of the payload you send instead of headers on a proxied request. Either way the discipline is the same: the structure exists only because your code declared it, so treat session propagation as part of the agent's plumbing rather than an afterthought bolted on during an incident.
- Why does a session sometimes show only half of an agent's calls?Because the missing calls did not carry the header. Session membership is declared per request, so any path that loses the identifier — a nested tool call, a retry, a step handed to a background worker or another service — logs as an unrelated row. The fix is to propagate the session id through a request-scoped context and across job payloads, not to expect the proxy to infer it.
- Should the session id be a custom property instead?No. Custom properties are aggregation dimensions and want low cardinality; a session id is unique per run, so as a property it would produce one group per run and be useless for grouping, while missing the dedicated session view entirely. Use `Helicone-Session-Id` for run grouping and keep properties for categorical context such as feature or environment.
- What does Helicone-Session-Path give you that the session id alone does not?Ordering by hierarchy rather than just membership. The id says these calls belong together; the path says this retrieval happened inside the planning step, which happened inside the run. That nesting is what lets you read an agent's execution as a tree, attribute cost to a sub-step, and spot a sub-step that is looping.
saying these in an interview costs you the question
- Assuming Helicone infers which calls belong to one run
- Setting the session id once as a client default for all traffic
- Using a custom property to carry the run identifier
- Forgetting to propagate the session id into background workers
- Thinking the path header is only a cosmetic label