skip to content

Your security reviewer wants every caller-facing Go error cut to one fixed string; support wants detail. How do you set the policy?

level: principalimportance: nice to knowfreq 30%

answer

  1. posture, not per-endpoint judgement
  2. allowlist beats blocklist
  3. what does support actually need?
  4. code, caller's own input, correlation id
  5. enforce in one helper with a test

basics

~20 s

Write one default-deny rule with a named allowlist of caller-safe facts: a stable error code, values the caller supplied, and a correlation id. Enforce it in a single boundary helper with a test, and record who may grant exceptions.

solid answer

~50 s

I would not settle it per endpoint; I would set a posture. Default deny: nothing from a wrapped Go error chain crosses to a caller except an allowlisted set — a stable machine-readable code, values the caller itself supplied, and a correlation id. That gives the reviewer what they want, because the allowlist is short and auditable rather than a blocklist of things to strip. Support gets more than they have today, because the correlation id resolves to the full chain plus request context in the logs, which beats a sentence a customer copy-pasted. I enforce it structurally — one boundary helper, plus a test asserting captured response bodies — so it survives staffing changes, and I write down who signs exceptions and where they are recorded. Then I revisit for external partners, where a documented stable code set is worth more than any prose.

go deeper

for a junior

Be ready to state the safe default: callers get a fixed message and a reference id, operators get the full wrapped chain in the log. You are not expected to set the policy, only to follow it and know why it exists.

for a middle

Explain the mechanism that makes a policy enforceable: one boundary helper, one place where an error becomes a response, and a test on the captured body. Know why a blocklist of substrings does not hold up.

for a senior

Show that you would take the support objection seriously and pay it off with a correlation id, retention and a fast lookup, rather than winning the argument and watching the rule erode under on-call pressure.

for a principal

Own the posture end to end: default deny with a named allowlist, an explicit exception owner and record, different treatment for partner APIs, and a stated trigger for revisiting it. Be able to defend the triage cost to support and the residual disclosure to the reviewer in the same conversation.

## Why this is a posture, not a bug fix "How much internal detail may a caller see" is not decided by whoever writes the next handler. It is a standing rule with three interested parties — the security reviewer who can veto, the support and on-call organisation who pay for terseness, and the API consumers who have to build against whatever you emit. In Go the question is unusually sharp, because a wrapped error chain is one flat string with no marking of which parts are internal, so the *only* place the distinction can exist is a rule you write and enforce. ## Default deny, with a named allowlist Blocklists lose. "Strip connection strings, strip SQL, strip file paths" needs updating for every new dependency, every new wrap site, and every wording change in a driver. An allowlist is stable and short: 1. **A stable error code** — a small, documented, machine-readable set (`not_found`, `invalid_argument`, `rate_limited`, `internal`). Codes are better than prose for both audiences: consumers can branch on them, and they disclose nothing about implementation. 2. **Values the caller supplied** — echoing back an id or a field name they sent reveals nothing new and makes the message actionable. 3. **A correlation id** — the join key between the terse response and the full chain in your logs. Everything else — the wrapped chain, driver wording, host names, internal identifiers you resolved on the caller's behalf, timings that hint at internal topology — stays on the operator side by default. "Default" is the operative word: the posture must name who can grant an exception, and exceptions should be written down next to the endpoint rather than living in someone's memory of a review. ## Answering support honestly The worst version of this argument is one where you tell support to accept less. Do not. Show them that a correlation id buys them the **whole** chain, the request fields, and the surrounding log context, instead of one truncated sentence pasted from a customer ticket. Then remove the friction that makes the id feel like a downgrade: it must appear in the response body, in the log line, and ideally in a response header, and there must be a one-click search that turns it into the incident. If that lookup is slow or gated behind an access request, the posture will erode, because engineers under pressure will re-add detail to the message. Retention matters too — a five-minute log retention makes the id useless for anything a customer reports the next day. ## Enforcement is the actual decision A policy nobody can violate is worth more than a policy everyone agrees with: - **One helper** that turns an error into a response, so the disclosure decision exists in one readable function per service. - **A test** that forces a failure and asserts on the captured response body: fixed text and id present, driver and query text absent. This is what a security reviewer can sign off, and it is what makes a regression a red build instead of an incident. - **A review rule** narrow enough to be mechanical: handlers never pass error text to the response writer. A style-guide paragraph and a review checklist are the weakest options, because both depend on attention. Prefer the version that fails loudly. ## Where the posture legitimately differs - **Internal service-to-service calls** are not automatically exempt. They cross trust boundaries in most estates, their responses get logged by the caller, and an internal caller pasting your message into a ticket ends up in the same place. If you do relax the rule internally, say so explicitly and say why. - **Public partner APIs** need the opposite of prose: a documented, versioned code set that partners can actually program against, plus a stable field for the correlation id. This is where investment pays, and where a free-text message is a liability you cannot change later. - **Debug and admin surfaces** may carry more, if they are authenticated, audited and clearly separated from the normal path — never as a flag that changes what the production path emits. ## The pattern to refuse Environment-gated verbosity — full chain in staging, fixed text in production — is the most common proposal and the one to reject. It means the code path that ships is the path nobody exercises, one bad config reinstates the leak, and the fault mode is silent. Keep one path; put the detail in the log. ## Owning the decision Finally, be explicit about who owns it and what would change it. The security reviewer can overrule an exception; the service owner writes the posture; a change in threat model (a public API launch, a compliance scope change, a partner integration) is a reason to revisit it. Write those three sentences down. The reason this is a principal-level question is not that the technique is hard — the helper is fifteen lines — it is that the rule has to survive people who never read the incident that motivated it.

  • How do you answer a team that says stripping detail slows triage?
    By making the correlation id strictly better than what they had: it resolves to the full chain plus request context, not one pasted sentence. Then fix the friction — the id in body, header and log, a one-step search, and retention long enough to cover customer-reported issues. If the lookup is painful, the policy will erode.
  • Should verbose errors be enabled in staging and disabled in production by flag?
    No. It makes the shipping path the untested one and turns a config mistake into a disclosure incident, silently. Run the same path everywhere and send detail to logs, which you can read in staging just as easily.
  • Does the same posture apply to internal service-to-service calls?
    Not automatically, but do not exempt them by default. Internal callers log your responses and paste them into tickets, and most estates treat service boundaries as trust boundaries. If you relax the rule internally, state the exemption explicitly and say what threat model justifies it.
  • What do you publish for an external partner integrating with your Go API?
    A documented, versioned set of machine-readable codes plus a stable correlation-id field — things a partner can branch on and quote in a support request. Free-text messages become a de facto contract you cannot change, so keep them fixed and put the variation in the code.

saying these in an interview costs you the question

  • Says the team will just be careful in code review
  • Builds a blocklist of sensitive substrings to strip
  • Gates verbosity on an environment variable
  • Exempts internal service calls without stating why
  • Cannot say who may grant an exception or where it is recorded
  • Gives no correlation id, so support genuinely loses ground