In Go's os/exec, how do you tell a child's nonzero exit from a failure to launch it at all?
answer
- three outcomes, one error value
- what a nil error actually promises
- one error type means it really ran
- errors.As, then ExitCode
- an underscore turns red into green
basics
~20 sMatch the error with errors.As. An *exec.ExitError means the program ran and exited nonzero, and its ExitCode method gives the status. Any other error, such as an *exec.Error wrapping exec.ErrNotFound, means it never ran. A nil error means it exited zero.
solid answer
~50 sThe error returned by `Run`, `Output` or `CombinedOutput` carries three distinct outcomes. Nil means the process ran and exited with status zero. An `*exec.ExitError`, matched with `errors.As`, means it ran and exited nonzero; `ee.ExitCode()` gives the status, and returns `-1` if the process was killed by a signal rather than exiting normally. Anything else means it never ran — typically an `*exec.Error` that satisfies `errors.Is(err, exec.ErrNotFound)` for a missing binary, or a path error for one that is not executable — and that is a very different incident: a deploy that could not start is a host problem, not a failed step. The failure mode this prevents is the discarded error: `out, _ := cmd.Output()` compiles, returns an empty slice, and reports a red step as green. Always inspect the error, wrap it with `%w`, and include the exit code and the captured output in the message.
code
go · 12 linesout, err := exec.Command(bin, "apply", "-f", file).CombinedOutput()
var ee *exec.ExitError
switch {
case err == nil:
return nil
case errors.As(err, &ee):
// Ran to completion with a nonzero status; -1 means it was signalled.
return fmt.Errorf("apply exited %d: %s", ee.ExitCode(), out)
default:
// Never started: missing binary, bad Cmd.Dir, not executable.
return fmt.Errorf("apply could not start: %w", err)
}go deeper
Recall that these calls return nil only when the program exited zero, and that the error must never be discarded. Know that *exec.ExitError is the type that means the program really ran.
Explain how errors.As separates an exit status from a launch failure, what ExitCode returns for a signalled child, and which call populates the ExitError's captured stderr.
Diagnose a pipeline that reports success on a failed step, and design error handling that names the exit code, includes the child's output, and routes launch failures differently from tool failures.
Own the contract you take on when branching on another tool's exit codes, and set the standard for how orchestrated failures are surfaced so an on-call engineer needs no reproduction.
## Three outcomes, one error value When you run an external program there are three things that can happen, and `os/exec` encodes all of them in the single `error` returned by `Run`, `Output` or `CombinedOutput`: 1. **It ran and succeeded.** The error is `nil`. That is the *only* meaning of nil here — it is a statement about the exit status, not merely about the launch. 2. **It ran and failed.** The error is an `*exec.ExitError`. 3. **It never ran.** The error is something else entirely. Collapsing 2 and 3 into "the command failed" is the mistake that produces the worst incidents, because the operational response differs completely: a nonzero exit is the tool telling you something about your input, while a launch failure means the host is not set up the way you assumed. ## Distinguishing them ```go out, err := exec.Command(bin, "apply", "-f", "service.yaml").CombinedOutput() var ee *exec.ExitError switch { case err == nil: return nil case errors.As(err, &ee): return fmt.Errorf("apply exited %d: %s", ee.ExitCode(), out) default: return fmt.Errorf("apply could not start: %w", err) } ``` `errors.As` rather than a type assertion, because your own wrapping (or the package's) may sit in between. ### `*exec.ExitError` It embeds `*os.ProcessState`, so the state's methods are promoted onto it: - `ExitCode()` — the status the program exited with. It returns `-1` when the process has not exited normally, most commonly because it was terminated by a signal. Treat `-1` as "it did not choose to exit", not as "code minus one". - The state also tells you whether it exited at all, and carries the platform-specific status for the cases where you need the signal. - The `Stderr` field holds captured standard error **only** if the error came from `Output` with a nil `Cmd.Stderr`. After `Run` or `CombinedOutput` it is empty, and reaching for it there is a common source of blank error messages. ### Launch failures - **Program not found.** `exec.Command` calls `exec.LookPath` up front and stores the failure in `Cmd.Err`; running the command returns an `*exec.Error` (which has the attempted `Name` and the underlying `Err`) satisfying `errors.Is(err, exec.ErrNotFound)`. - **Found but not executable, wrong architecture, missing loader.** These surface as operating-system errors from the attempt to start the process, typically a path error naming the file. - **Bad `Cmd.Dir`.** A working directory that does not exist fails at start, not at exit. None of these is an `*exec.ExitError`, and none of them produces a meaningful exit code — which is why code that reaches straight for `ExitCode()` after a type assertion will panic or mislead. ## The failure this question exists for A deployment orchestrator replaces a shell script. The shell script had `set -e`, so any nonzero exit stopped it. The Go rewrite has, at one of forty call sites: ```go out, _ := exec.Command(bin, "apply", "-f", f).Output() log.Printf("applied %s: %s", f, out) ``` The underscore is the bug. The tool exits 1, `out` is empty, the log line prints happily with nothing after the colon, and the pipeline goes green on a deploy that did not happen. Nothing in the language warns you: discarding an error is legal, and the empty output looks like a quiet success. Three things make this class of defect visible: - **Never discard the error from a subprocess.** A vet-style check for ignored errors, or simply a review rule, catches it. Ignoring the *output* is fine; ignoring the *error* never is. - **Put the child's own words in yours.** `%s` on the captured bytes plus the exit code turns `exit status 1` into a message someone on call can act on without reproducing the run. - **Distinguish the two failure classes in the message.** "exited 2" and "could not start" route to different people. ## Exit codes as a contract If you branch on specific codes — treating 2 as "nothing to do" and anything else as fatal, say — you have taken a dependency on the tool's documented exit-code table. That is a legitimate thing to do and often much better than matching on message text, but write down which codes you rely on, because a tool upgrade can add or reassign them. Codes above 128 conventionally indicate a signal in shell reporting, but that is the *shell's* convention: in Go a signalled child gives you `-1` from `ExitCode()` and the signal itself from the process state.
- What does exec.ExitError.ExitCode() returning -1 tell you?That the process did not exit with a status of its own — in practice it was terminated by a signal, or has not exited yet. It is a sentinel, not a real code, so branching on it as if it were one is wrong; consult the process state if you need to know which signal ended it.
- Why is a missing binary a different incident from a nonzero exit?A nonzero exit is the tool's verdict on your input, and the fix is usually in your request or your data. A launch failure means the host does not have the tool you assumed, so the fix is in provisioning or in your own resolution logic. Reporting both as 'step failed' sends the page to the wrong person.
- How do you make the child's own error message reachable when you used CombinedOutput?Attach the returned bytes to your error with `%s`, because `CombinedOutput` leaves `ExitError.Stderr` empty — that field is filled only by `Output`. Without doing so, the caller sees `exit status 1` and nothing about what the tool actually complained about.
- Is it reasonable to branch on specific exit codes from an external tool?Yes, when the tool documents them, and it is far more robust than matching message text. Record which codes you depend on and why, because a version upgrade can reassign them; and keep a default branch that treats unknown nonzero codes as fatal rather than as success.
saying these in an interview costs you the question
- Ignores the error and reads only the output bytes
- Type-asserts to *exec.ExitError without errors.As
- Treats a missing binary as a nonzero exit status
- Reads ExitError.Stderr after CombinedOutput and finds it empty
- Assumes exit code 127 is reported for a program not found
- Interprets an exit code of -1 as a real status