A Go http.Client returns stopped after 10 redirects for some crawled URLs. How do you find the loop?
answer
- ten is a budget, not a bug
- unwrap it rather than matching text
- Do still hands you the last response
- the client threw the hops away
- CheckRedirect is where the chain is visible
basics
~20 sThat message is the default redirect policy giving up after ten consecutive hops. Client.Do returns a *url.Error alongside the last response, whose body is already closed. Install a CheckRedirect that logs each hop's req.URL to see the cycle.
solid answer
~50 sThe message comes from the built-in policy that applies when `CheckRedirect` is nil: stop after ten consecutive redirects. `Do` returns the error wrapped in a `*url.Error`, which you unwrap with `errors.As`, and it also returns the most recent response with its body already closed, so status and headers are still readable. Nothing in that error tells you the chain, because the client discarded the hops. To see them, set a `CheckRedirect` that logs `len(via)` and `req.URL` for every hop, or dumps the outgoing request with `httputil.DumpRequestOut(req, false)` when you need headers too. The chain almost always shows a cycle: two hosts bouncing a scheme or a trailing slash back and forth, or a site that keeps redirecting to a login page because the client has no cookie jar. Raising the limit is not the fix — cap it lower and report the cycle.
code
go · 13 linesclient := &http.Client{
CheckRedirect: func(req *http.Request, via []*http.Request) error {
if len(via) >= 5 {
return fmt.Errorf("redirect chain too long from %s", via[0].URL)
}
dump, err := httputil.DumpRequestOut(req, false)
if err != nil {
return err
}
log.Printf("hop %d:\n%s", len(via), dump)
return nil
},
}go deeper
Know that ten consecutive redirects is the built-in limit and that exceeding it produces an error rather than a response. Recognise that the limit is a symptom of a cycle, not the problem itself.
Explain that the error is wrapped in a *url.Error and unwrapped with errors.As, that the last response comes back with its body closed, and that CheckRedirect is the only place the hops are visible.
Walk the diagnosis: log or dump each hop, read the chain, and name the usual culprits — scheme and host canonicalisation fighting, or a cookie-driven login wall with no jar. Argue for capping lower rather than raising the limit.
Decide what a redirect cycle means for the product: a finding the crawler reports and moves past, or a hard failure. Set the cap and the cookie posture once, so every crawl behaves the same rather than per-team.
## What the error actually is When `http.Client.CheckRedirect` is nil the client uses its default policy: follow at most ten consecutive redirects, then stop with an error carrying that text. Because it is produced by the redirect check, it reaches you the way every `CheckRedirect` error does — wrapped in a `*url.Error`: ``` var uerr *url.Error if errors.As(err, &uerr) { // uerr.Op, uerr.URL, uerr.Err } ``` String-matching on the message is fragile; unwrap instead. Two other facts matter for a crawler: - **`Do` returns a response as well as the error.** On this path the client hands back the most recent response with its body already closed. You can still read its status and headers — including the `Location` that was about to be followed — which is genuinely useful for a report. Code that assumes a nil response when `err != nil` throws that away. - **The error carries no chain.** The intermediate responses were consumed and closed inside the client. Whatever you want to know about the hops must be recorded while they happen. ## Recording the chain `CheckRedirect` is the only vantage point. It is called before each hop with the upcoming request and the requests already made: ``` CheckRedirect: func(req *http.Request, via []*http.Request) error { log.Printf("hop %d -> %s", len(via), req.URL) return nil } ``` When you need more than the URL — which headers survived a host change, what the client decided to send — dump the outgoing request: ``` dump, err := httputil.DumpRequestOut(req, false) ``` `DumpRequestOut` renders a client-side request the way it will go on the wire, headers included; pass `false` so it does not touch the body. (`httputil.DumpRequest` is the server-side counterpart, for a request you received.) Printing one dump per hop makes the cycle obvious in a way a bare error never will. For a crawler, do not log — collect. Build a per-call slice of URLs inside the closure and attach it to the result, so the report says exactly which four URLs the site cycles between. ## What the chain usually shows - **Scheme and host canonicalisation fighting.** `http://example.com` redirects to `https://example.com`, which redirects to `https://www.example.com`, which redirects back to the apex over http. Two rules written by two teams, each correct alone. - **A cookie-driven login wall.** The site redirects to a login page, sets a cookie, and redirects back — forever, because the Go client has no `Jar` and never returns the cookie. This one looks like a broken site and is actually a missing three lines of client configuration. - **Trailing-slash ping-pong.** One layer appends a slash, another strips it. - **Request-shape sniffing.** The site behaves differently for the default `User-Agent` or a missing `Accept`, sending a crawler somewhere a browser never goes. The first two are far and away the most common, and the dump distinguishes them instantly: in the cookie case you will see `Set-Cookie` being offered on each hop and no `Cookie` going back. ## The policy a crawler should actually run Ten is generous for a link checker. A legitimate chain is one to three hops. So: 1. **Cap lower.** Return an error from `CheckRedirect` once `len(via)` passes your own limit — three or five. You fail faster and you have the chain in hand. 2. **Make the loop a result, not a crash.** A cycle is a finding about the site, exactly the kind of thing the tool exists to report. Record the URLs and carry on to the next link. 3. **Decide about cookies deliberately.** If the targets are trusted, a jar removes a whole class of false loops. If they are not, no jar plus a low cap is the safer posture. 4. **Do not raise the limit.** Going from ten to fifty turns a fast failure into a slow one and hides the fact that a site is cycling. The only defensible reason is a specific known target with a genuinely long legitimate chain. ## Related failures that are not this A chain that is long but not cyclic produces the same error, and the fix is entirely different — usually a stale seed URL that should be updated to the canonical destination. A chain that stalls rather than looping is a timeout problem, not a redirect problem: the redirect budget counts hops, not seconds, and ten slow hops can each be within their deadline while the whole call is far too slow. Reading the recorded chain tells you which of the three you have in a few seconds; guessing from the error text does not.
- Why would a browser load these URLs fine while the crawler loops?Usually cookies. The site redirects to a login or consent page, sets a cookie and redirects back; a browser returns the cookie and settles, while a Go client with a nil Jar never does and cycles until the budget runs out. Less often it is request-shape sniffing on User-Agent or Accept sending a crawler down a different path.
- Would raising the limit above ten ever be the right call?Rarely. A legitimate chain is one to three hops, so a higher cap mostly converts a fast failure into a slow one and hides that a site is cycling. The defensible case is one known target with a genuinely long chain, pinned to that host in a dedicated client rather than raised globally.
- How do you record the chain per request rather than just logging it?Build the client's CheckRedirect as a closure over a per-call slice: append req.URL on each hop and return nil until your own cap. Since a shared client runs the callback on many goroutines, either construct a small throwaway client per crawl target or key the collection by something request-scoped and guard it with a mutex.
saying these in an interview costs you the question
- Raises the redirect limit instead of finding the loop
- Assumes Do returns a nil response when the limit is hit
- String-matches the error text instead of unwrapping it
- Looks for a redirect limit field on http.Transport
- Blames the remote site without logging the chain