A service prints summaries to its output stream and full detail to a file inside the container — what does that split cost you?
answer
- two destinations, one collector
- the thin half is the kept half
- detail dies with the instance
- no identifier, no join across the two
- discarding is not reducing
basics
~20 sThe split keeps only the half that was never the useful half. Summaries on the output stream are captured and shipped; the detail is collected by nothing, readable only from inside that running instance, and gone the moment it is replaced.
solid answer
~50 sThe platform captures the output streams and nothing else, so the split decides in advance which half of the evidence survives — and teams reliably put the thin half there. While the instance runs, the detail is readable only by someone who can get a view inside that specific instance; after a rollout, a rescheduling or a crash-and-replace, it is gone with the writable layer. Even when you can read both, they do not join: two writers, two buffers, two stamping regimes, so you cannot reliably line a summary up against its detail. It also feels like a volume control and is not one — the bytes are still produced, just discarded instead of collected. If volume is the real problem, cut it where the line is emitted or downstream in the pipeline, and keep one destination.
code
pseudocode · 9 lineson session_event(e):
emit_to_output_stream("session " + e.id + " state=" + e.state) # captured by the platform
append_to_file("/app/logs/session.log", e.detail) # captured by nothing
on instance_replaced():
# the stream copy was read by the runtime and has already left the node
# /app/logs/session.log was never read by anything, and goes with the instance
kept = ["session " + e.id + " state=" + e.state]
lost = [e.detail]go deeper
Remember which half survives: what goes to the output stream is captured and kept, what goes to a file inside the container is not. Splitting the two means the detail is the part you lose.
Explain the three costs separately — remote reach while running, survival past replacement, and the inability to correlate two independently buffered and stamped writers — rather than just asserting that files are bad.
Demonstrate the incident view: after a replacement you hold only the captured half, so argue the boundary rule up front and show what you would put on the stream so a responder needs no second source.
Own the trade-off you are making for everyone: a single collected destination gives uniform evidence and a predictable ingest bill, and it pushes the volume conversation onto emitting teams. Decide that deliberately rather than letting each service split its own way.
## What the split actually decides A container platform captures what the process writes to its standard output and error streams. That is the collection boundary, and writing to two destinations does not widen it — it **partitions your evidence into a collected half and an uncollected half before you know which half you will need**. The partition is almost always made the wrong way round. Summaries go on the stream because they are cheap and tidy; the per-request or per-session detail goes in the file because it is bulky. But during an incident the summary tells you *that* sessions failed and the detail tells you *why*, so the half you kept answers the question you did not have. ## Three distinct costs 1. **Reach while it runs.** The captured half is queryable by anyone with access to the platform's log view, from anywhere. The file half is readable only by someone who can obtain a view inside that one instance, on that one node, right now. That is a different permission, a different tool and a different person at 02:00. 2. **Survival.** The captured half has already left the node by the time you look. The file half lives in the container's own writable layer and is discarded when the instance is replaced — and replacement is exactly what happens after a crash, a rollout or a rescheduling, which are exactly the events you are investigating. 3. **Correlation.** Two destinations mean two writers with independent buffers and two stamping regimes: the stream copy is stamped when the runtime reads it, the file copy whenever the writing library decided to stamp it. Lining a summary up against its detail then depends on an identifier you remembered to put in both. Most teams discover they did not. | property | line on the output stream | line in a file inside the container | |---|---|---| | collected by the platform | yes, by the default capture path | no, by nothing | | readable remotely while running | yes | no — only from inside that instance | | survives instance replacement | yes, as the shipped copy | no | | stamped by | the runtime, as it reads the line | the writing library, on its own clock | ## Why it is not a volume control The usual defence is cost: "we cannot afford to ship all of that". The split does not reduce what the process produces — it produces exactly the same bytes and then throws half of them into a destination nobody reads. What it reduces is **ingest**, by discarding, without any of the deliberateness of discarding on purpose: no rule about which records go, no record of what went, no way to turn it back on for one workload during an incident. If volume is genuinely the constraint, the honest levers are at the emitting side and in the pipeline that carries the lines away — how much a given call site emits, and what the collection tier does with it. Those are a different subject from this one; what belongs here is the boundary rule: **one destination, and it is the one the platform reads.** There is a second, quieter cost. A file written inside the boundary consumes space the instance is accounted for, and a workload that writes a large uncollected file is spending a resource it gets no value from. What a node does when that space runs short is its own subject; the point here is that the file is pure cost — it buys no reach, no survival and no correlation. ## The shape of the fix - Send everything a responder might need to the output stream, including the detail, and let the platform's capture be the single collection point. - Where a component insists on a file, forward it onto the stream rather than leaving two destinations live — one path in, one path out. - Keep an identifier on every line that ties related records together, so that a single captured stream can be narrowed to one session without needing a second source. - If you must keep an on-instance file for a genuinely local purpose — a scratch artefact a support tool reads live — treat it as a convenience that vanishes, never as the record of what happened. The test to apply before shipping: *if this instance is replaced during the incident, what do I still have?* If the answer is "the summaries", the split is the defect and not the logging level.
- The team says the file is fine because support can read it live when a customer complains. Is that an answer?Only for complaints that arrive while the instance is still alive and someone can get a view inside it. That is a narrow window, it needs a privileged path into a running workload, and it fails for precisely the cases where the instance died. Treat the file as a live convenience, never as the record.
- If shipping everything is too expensive, what do you change?Reduce what the call sites emit, or have the collection tier drop or thin records under an explicit rule — decisions made where they can be reviewed and reversed. Splitting by destination is not a cheaper version of that; it discards silently, keeps no account of what went, and cannot be turned back on for one workload mid-incident.
saying these in an interview costs you the question
- Calls the split a way to cut log volume rather than a way to discard it
- Assumes support can always read the in-container file when it matters
- Expects summary and detail to line up without a shared identifier
- Thinks the file half will be shipped if a collector is installed later
- Believes the platform keeps the container's files after replacement