skip to content

Which parts of a Go service's log output should a platform owner mandate fleet-wide, and which stay a team's call?

level: principalimportance: should knowfreq 30%

answer

  1. mandate only what cannot be fixed later
  2. one stream, one format, no file
  3. buffering trades the crash tail for syscalls
  4. counter-offer: log less, not later
  5. ship the default, don't write the rule

basics

~20 s

Mandate only what the collector depends on: one output stream, one wire format, no log files inside the container, and no buffering of the sink without an exemption. Leave levels, record contents, sampling and handler implementation to teams.

solid answer

~50 s

I mandate the three things the platform cannot repair afterwards: every service writes records to one process stream and never to a file inside the container; every service emits the same wire format with the same shared field names; and the sink is unbuffered unless a team holds an exemption. Those are the collector's contract, and a service that gets them wrong is unreadable during the incident it caused. Everything above that line stays with the team: what they log, at which level, how they sample, how their handler is built. When a team wants to buffer for throughput I ask for the measurement first, because record volume is usually the real cost and cutting it saves more without trading away the tail. If they still need it, I attach flush-on-error, a flush ticker, SIGTERM handling and a test proving the tail survives — and I enforce it by shipping the compliant logger as the default.

go deeper

for a junior

Focus on the compliant behaviour itself: use the shared logger your platform provides, write records to the standard stream, and do not invent your own output destination.

for a middle

Be able to explain why a fleet standardises the stream and the format at all, and what the collector and cross-service queries actually depend on.

for a senior

Show that you would meet a buffering request with a measurement and conditions rather than a flat refusal, and that you would prove the tail survives with a test.

for a principal

Own the line: mandate only what the platform cannot repair after the fact, hand everything else back explicitly, and enforce by shipping the default rather than publishing a rule.

## Draw the line where the platform cannot recover The useful test for a mandate is not "is this a good idea" but "if a team gets this wrong, can the platform fix it afterwards without touching their code?" Almost everything about logging passes that test — a noisy service can be filtered, a badly named field can be aliased in a query, a missing debug record can be added next sprint. Three things do not: 1. **Where the records go.** If a service writes to a file inside its container, the platform never sees the records at all. There is no recovery; the data was destroyed with the container. 2. **What the records look like on the wire.** If every service invents its own format, the collector needs per-service parsing and cross-service correlation becomes manual work. That cost lands entirely on the platform. 3. **Whether the sink is buffered.** Buffered records that never flushed do not exist. The platform cannot retrieve what was never written, and the records lost are always the newest — the ones describing the failure. Those three are the mandate. They are cheap to comply with, mechanically checkable, and each protects a property the platform is accountable for. ## What stays with the team Everything about *content*. Which events are worth a record; what level each one gets; whether per-request logging is sampled at 1% or off entirely; which domain attributes a record carries; whether they build one logger or several; how their handler is implemented. A platform that legislates message wording is spending authority it will need later, and it makes the platform team the bottleneck on every service's debuggability. There is a shared middle: a small set of common field names — service, version, request identifier, level, timestamp — because those are what dashboards and correlation queries join on. Standardise the join keys; leave the rest of the record alone. ## The buffering argument, concretely A team profiles a hot path, sees log writes costing a few percent of CPU, and asks to put a buffer in front of the sink. This is the case the posture exists for, and "no" without reasoning is the wrong answer. What I weigh: - **What buffering actually saves.** It removes system calls, not the formatting and serialisation of the record — and formatting is usually the larger half. The measured saving is often smaller than the profile suggested. - **What it costs.** Every record in the buffer exists only in the process. `os.Exit`, `log.Fatal`, an uncaught SIGTERM, SIGKILL and fatal runtime errors all discard it. The lost records are the newest ones, which is the tail on-call reads first. It also delays visibility, so a stuck service looks silent. - **The better trade.** If logging is genuinely expensive, the cause is volume. Sampling per-request records, or moving them below the enabled level, removes the formatting cost *and* the syscall, and takes nothing away from the crash tail. I offer that first, because it is a bigger win and has no downside at exit. If they still need buffering — a batch job whose output volume is the actual workload, whose final status is recorded elsewhere — I grant it as an exemption with conditions attached: flush on every record at error level and above, a periodic flush that bounds visibility lag, a SIGTERM handler that flushes on the way out, `log.Fatal` banned in that binary, and a subprocess test in their pipeline that runs the real failure path in a child and asserts the last record survived. The exemption is reviewable, and the test is what keeps it honest after the person who requested it moves on. ## Where I can be overruled, and by what Two constraints legitimately beat the default. **Cost**: if log volume dominates the platform bill, something has to give — but the answer is fewer records, not later ones, and I should be able to show that with data rather than asserting it. **A workload whose economics really are different**: a high-throughput pipeline stage where the log stream is the product. In both cases the argument that wins is a measurement, and I should say up front what measurement would change my mind. A posture that cannot be argued with is a posture people route around. ## Enforcement is a default, not a document The part teams get wrong is the part that requires them to remember something. So the posture ships as code: a shared constructor that returns a logger writing unbuffered to the standard stream in the fleet's format, wired into the service template so a new service is compliant before anyone thinks about it. Add one check in the shared pipeline for the obvious violations, and a platform-side alert when a service stops producing records at all — that catches the buffered sink that quietly stopped flushing. A rule in a wiki decays at the rate the team turns over. A default decays at the rate someone deliberately opts out of it, which is far slower and, when it happens, visible. ## The shape of a good answer Name the small mandate and justify it by what the platform cannot repair. Hand back everything else explicitly, so it is clear the mandate is small on purpose. Show that you would meet the buffering request with a measurement and a counter-offer rather than a refusal, and that you would attach conditions rather than trust. Finish with enforcement by default, because a posture nobody can bypass by accident is the only one that holds across dozens of services.

  • A team shows a CPU profile where log writes are several percent and asks to buffer. What do you say?
    Ask what the record volume is first. Formatting and serialising usually costs more than the write that ships the record, so sampling or dropping per-request info records saves more CPU and costs nothing at exit. If they still need buffering, grant it as an exemption with flush-on-error, a flush ticker, SIGTERM handling, log.Fatal banned, and a subprocess test proving the tail survives the failure path.
  • Why standardise the wire format across the fleet rather than let each team choose?
    Because the collector, every dashboard, every alert and every cross-service query parse it. Per-team formats push parsing and correlation cost onto the platform, which cannot fix them without touching each team's code. Standardise the join keys — service, version, request id, level, timestamp — and leave everything else in the record to the team.
  • How do you make the posture hold across dozens of services without policing repositories?
    Ship it as the default: a shared constructor that returns a compliant logger, wired into the service template so a new service starts correct. Add one pipeline check for the obvious violations and an alert when a service stops emitting records at all, which catches a buffered sink that quietly stopped flushing. Documentation alone decays with team turnover.
  • What evidence would make you withdraw the unbuffered-by-default rule for a particular service?
    A measurement showing log syscalls are a material share of that service's cost after record volume has already been cut, plus a credible story for how its failure is reconstructed without the tail — a final status recorded elsewhere, for instance. Saying up front what would change my mind is what keeps the posture arguable instead of routed around.

saying these in an interview costs you the question

  • Mandates everything, down to levels and message wording
  • Bans buffering outright with no measurement and no exemption path
  • Lets each team pick its own wire format and expects the collector to cope
  • Enforces the posture with documentation rather than a default
  • Treats a team's throughput measurement as an argument to be ignored