How do you use httptrace.ClientTrace to split one outbound call into DNS, connect, TLS and first-byte phases?
answer
- hooks, not a wrapper this time
- it rides in the request context
- one function field per phase
- an empty phase can be the answer
- GotFirstResponseByte separates them from us
basics
~10 sBuild a httptrace.ClientTrace with hook functions that stamp time.Now, attach it to the request's context with httptrace.WithClientTrace and req.WithContext, then send the request normally. The hooks fire as the transport reaches each phase.
solid answer
~40 sI build a `&httptrace.ClientTrace{}` whose hooks record timestamps — `DNSStart`/`DNSDone`, `ConnectStart`/`ConnectDone`, `TLSHandshakeStart`/`TLSHandshakeDone`, `GetConn`/`GotConn`, `WroteRequest` and `GotFirstResponseByte` — then attach it with `req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace))` and send it through the ordinary client. Subtracting the marks gives the breakdown, and `GotFirstResponseByte` minus the request write is the upstream's own think time. Two caveats decide how you use it: the hooks may run on different goroutines and some can fire after the call has finished, so shared state needs synchronising or a per-request struct; and on a healthy client most calls reuse an idle connection, so `GotConn` reports `Reused: true` and no DNS, dial or handshake hooks fire at all. That absence is itself the finding. I treat it as a diagnostic I turn on for a sample or an investigation, not per-request instrumentation for everything.
code
go · 13 linesvar dnsStart, connStart, wrote, firstByte time.Time
var reused bool
trace := &httptrace.ClientTrace{
DNSStart: func(httptrace.DNSStartInfo) { dnsStart = time.Now() },
ConnectStart: func(network, addr string) { connStart = time.Now() },
GotConn: func(info httptrace.GotConnInfo) { reused = info.Reused },
WroteRequest: func(httptrace.WroteRequestInfo) { wrote = time.Now() },
GotFirstResponseByte: func() { firstByte = time.Now() },
}
req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace))
resp, err := client.Do(req)go deeper
Know that net/http can report the phases of an outbound call through callbacks, and that they are attached per request rather than configured once on the client.
Be able to name the phase hooks and how they are attached through the request context, and explain what the gap between writing the request and the first response byte represents.
Use the breakdown to assign blame: pool wait, DNS, handshake or upstream think time, and read connection reuse correctly instead of mistaking missing hooks for a broken trace.
Decide the posture: a sampled or on-demand diagnostic rather than always-on instrumentation, weighing hook cost and metric cardinality against how often the phase split actually changes a decision.
## The problem it solves A transport-level timer tells you an upstream call took 900ms. It does not tell you whether that was a slow DNS resolver, a saturated connection pool, a TLS handshake against a distant region, or an upstream that simply took 900ms to think. `net/http/httptrace` fills that gap by letting you register callbacks that `net/http`'s transport invokes as it moves through the phases of one request. ## Attaching a trace A `ClientTrace` is a struct of optional function fields; you set the ones you care about and leave the rest nil. It travels **in the request's context**, which is what makes it per-request rather than global: ```go trace := &httptrace.ClientTrace{ DNSStart: func(httptrace.DNSStartInfo) { ... }, GotFirstResponseByte: func() { ... }, } req = req.WithContext(httptrace.WithClientTrace(req.Context(), trace)) resp, err := client.Do(req) ``` Nothing else changes: the same `http.Client`, the same transport. ## The hooks that make up a breakdown - `GetConn(hostPort string)` — the transport wants a connection. The gap to `GotConn` is time spent waiting for the pool, and it is the hook that exposes a per-host connection limit throttling you. - `GotConn(GotConnInfo)` — a connection is in hand. `Reused`, `WasIdle` and `IdleTime` tell you whether anything was dialled at all. - `DNSStart` / `DNSDone` — name resolution, when it happened. - `ConnectStart(network, addr string)` / `ConnectDone(network, addr string, err error)` — the TCP dial, once per address tried. - `TLSHandshakeStart` / `TLSHandshakeDone(tls.ConnectionState, error)` — the handshake. - `WroteRequest(WroteRequestInfo)` — the request, body included, is on the wire. - `GotFirstResponseByte()` — the first byte of the response arrived. Subtracting the `WroteRequest` mark gives the upstream's own latency, cleanly separated from everything the network cost you. ## The two caveats that matter in practice **Concurrency.** The documentation is explicit that hooks may be called from different goroutines, and that some may fire after the request has completed or failed. Writing into shared variables from them without synchronisation is a data race the race detector will find. The usual shape is a per-request struct whose fields are only read after `Do` returns, plus a mutex or atomic values for anything that could arrive late. **Reuse.** On a warm client with keep-alive working, the overwhelming majority of calls take an idle connection: `GotConn` fires with `Reused: true` and there is no `DNSStart`, no `ConnectStart` and no handshake to report. New engineers read the empty phases as broken instrumentation; they are the healthy case. Conversely, seeing dial and handshake hooks fire on *every* call is a real finding — a client constructed per request, `DisableKeepAlives` set, an upstream closing connections, or a per-host limit forcing new ones. ## Where it fits, and where it does not `httptrace` observes the internals of the standard transport. A custom `RoundTripper` that does not delegate to it will fire nothing, and there is no coverage of reading the response body — the timeline ends at the first response byte, so a slow streaming payload is still invisible here. Cost and cardinality both argue against enabling it for every request in a hot path: the hooks allocate a little, and turning six timestamps per call into metrics multiplies your series count. The realistic posture is to keep it behind a flag or a sampling decision, or to enable it on a debug endpoint against one upstream while an investigation is live. It answers a question you ask occasionally and precisely, which is exactly the shape of a diagnostic rather than a metric. ## Reading the result For the engineer who has to say where the milliseconds went, the breakdown resolves the argument quickly. Time concentrated between `GetConn` and `GotConn` is a pool problem you own. Time in DNS is infrastructure. Time in the handshake with `Reused: false` on most calls is a connection-reuse problem in your own client configuration. Time between `WroteRequest` and `GotFirstResponseByte` belongs to the upstream, and no amount of tuning on your side will move it.
- Most calls fire no DNS or connect hooks at all. What does that tell you?That those requests reused an idle keep-alive connection, which `GotConn` confirms with `Reused: true`. It is the healthy case, not broken instrumentation. The alarming pattern is the opposite: dial and handshake hooks on every call, which points at a client built per request, keep-alives disabled, a per-host limit, or an upstream closing connections.
- Why can writing timestamps into shared variables from these hooks be a race?The hooks are documented as possibly running on different goroutines, and some may fire after the request has completed or failed. Unsynchronised writes are then a genuine data race that `-race` will report. Use a per-request struct read only after `Do` returns, and guard anything that might arrive late.
- Would you enable this for every outbound request in production?No. The hooks allocate, and six timestamps per call turned into metrics multiplies series count for little routine value, since the phase split only matters while you are investigating. Keep it behind a sample rate or a debug flag scoped to one upstream, and rely on the plain transport-level duration the rest of the time.
saying these in an interview costs you the question
- Setting the trace on the client or transport instead of the request context
- Assuming missing DNS and connect hooks mean the trace is not working
- Writing to shared variables from hooks with no synchronisation
- Expecting the timeline to cover reading the response body
- Enabling full phase tracing on every request in a hot path