Many teams make the step that pushes outcomes into a case repository non-fatal, so a failed push cannot fail the job. What does that choice hide, and what has to be added to make it safe?
answer
- the suite owns the verdict
- non-fatal is not the same as unobserved
- partial cycles look complete
- alert on the absence, not the error
- auth errors never self-heal
basics
~20 sNon-fatal reporting is right in principle — a repository outage should not change a suite's verdict — but swallowing the error leaves a half-empty cycle nobody notices. Make the push loud: durable outcomes file, an alert on failure, reconciled counts.
solid answer
~50 sThe instinct is sound. The suite's own verdict is the truth about the code, and a case repository being unreachable says nothing about the code, so letting a reporting error fail the job manufactures a false red. What is wrong is the usual implementation, which suppresses the exit status and nothing else. That produces failures with no consumer: the cycle stays empty or partial, the repository still shows the previous round's outcomes, an expired credential goes unnoticed for weeks, and the run's real result survives only in a console log that ages out. **Non-fatal must never mean unobserved.** Write outcomes to a durable artefact before pushing so the push is replayable rather than the only copy; emit a distinct, visible warning when it fails; alert on a cycle that received nothing from a job that clearly ran; and treat authentication and authorization errors as a genuine failure rather than something to swallow.
go deeper
Understand that the suite decides whether the code is good, and that failing to report that verdict somewhere else should not change it. Know the risk: the outcomes can go missing quietly.
Explain the difference between a partial cycle and an empty one, and why the partial case is more dangerous — it looks finished, so nobody investigates.
Describe the concrete safeguards you would ship: a durable outcomes artefact, a visible warning, an alert on the produced-versus-accepted count gap, and different handling for transient versus credential failures.
Own the framing that non-fatal governs the verdict, not observability, and hold the line that no integration may fail in a way that leaves nothing to replay and nobody informed.
## Why non-fatal is the right default A pipeline asks one question of a suite: does the code behave? The suite answers it from its own execution, and that answer is complete before any push happens. A case repository's availability, credential expiry, or rate limiting has no bearing on that answer. So a job that turns red because a *reporting* call failed is reporting something untrue about the code, and teams learn quickly to distrust a signal that lies. Worse, the fix people reach for under pressure is to disable the reporting step entirely, which is a permanent loss to buy a temporary quiet. The reverse error is just as bad: a push that *succeeds* while the suite was red does not make the run green, and a push that fails does not make a green run suspect. Reporting is downstream of the verdict, never a source of it. ## What the usual implementation hides The common implementation is a suppressed exit status and nothing else. That creates a failure with no consumer, and several distinct problems hide behind it: - **The empty cycle.** The round shows nothing, and nobody notices, because an empty cycle looks the same as a cycle for a round that has not started. - **The partial cycle.** Worse than empty, because it looks complete. A chunked push that failed on chunks three through nine leaves a plausible-looking set of results with no marker saying the rest is missing. - **The stale green.** Where reporting overwrites a round in place, a failed push leaves yesterday's outcomes standing. A reader sees green results dated recently enough to believe. - **The credential that expired weeks ago.** Every push has failed since, every failure was swallowed, and it surfaces during an audit or a release review when someone finally asks why a cycle is thin. - **The evidence that only lived in a log.** The run's real outcomes existed in the console output, which was retained for a while and then was not. ## What has to be added 1. **Persist before you push.** Write the outcomes to a durable artefact as they are produced. This costs nothing, cannot fail for network reasons, and demotes the push from *the only copy of the truth* to *a replayable step over data you still hold*. Every other mitigation depends on this one. 2. **Make the failure visible where humans already look.** A distinct warning annotation on the job, a pipeline summary line, a message on the team's channel — anything with a reader. A line buried a thousand lines into a log has no reader. 3. **Alarm on the absence, not only on the error.** The failure mode that matters is a cycle that received nothing or received too little. Compare the count of outcomes the harness produced with the count the repository accepted, and raise that gap as its own signal; it catches silent row-level rejections that never produced an error at all. 4. **Separate transient from terminal.** Back off and retry on timeouts, rate rejections and gateway errors — these genuinely pass. Do not swallow authentication or authorization failures the same way: they never recover on their own, they mean somebody must act, and treating them as transient noise is exactly how a credential rots for a month. 5. **Replay, do not merely alert.** Because the artefact is durable, a failed push can be re-run later from the stored outcomes. That replay must carry the same idempotency key as the original attempt, or the retry that finally succeeds records everything twice. ## A quick decision table | Push outcome | Should it change the job's verdict? | What must happen | |---|---|---| | Transient network or rate rejection | No | Back off, retry, then warn and leave the artefact for replay | | Credential rejected | No, but it must be actionable | Warn loudly and page an owner; this never self-heals | | Some rows rejected inside a success | No | Raise the count gap explicitly; a status-only check misses it | | Nothing recorded for a job that ran | No | Alert on the empty cycle, which is the failure nobody sees | ## The line to hold Non-fatal is a statement about the **job's verdict**, not about **observability**. A reporting failure that changes nothing and tells nobody is not resilience; it is data loss with a friendly exit code. Which of the *suite's own* failures should stop a pipeline stage is a different question with a different owner. The rule here is narrower: never let the transport for outcomes fail in a way that leaves no artefact, no alert, and no way to replay.
- Which signal actually catches a reporting failure that produced no error at all?The count gap. Compare the number of outcomes the harness produced with the number the repository reports as accepted, and treat any difference as its own alert. That catches row-level rejections hidden inside an overall success, and it catches a cycle that received nothing from a job that clearly ran — neither of which raises an exception anywhere in the push step.
- Why should a rejected credential not be swallowed the way a timeout is?Because a timeout passes and a rejected credential does not. Retrying and backing off is the correct response to something transient; applying the same treatment to an expired or revoked credential simply hides a condition that will persist until a person rotates it. That is how a team discovers, during a release review, that reporting has been silently dead for weeks.
saying these in an interview costs you the question
- Suppresses the exit status and adds nothing else
- Lets a failed push turn a green run red
- Keeps outcomes only in the job log
- Retries an authentication failure forever
- Assumes an empty cycle would be noticed