As the owner of a Go orchestrator, how do you decide which external CLIs it may shell out to?
answer
- a dependency graph go.mod cannot see
- four things you depend on, not one
- who can upgrade it under you?
- startup resolution beats mid-run discovery
- one package owns the whole surface
basics
~20 sTreat every os/exec call site as an undeclared dependency on a binary, its flags, its output format and its exit codes, on hosts you may not own. Keep the ones where no usable library exists and the output is machine-readable, resolve them at startup, and agree the versions with whoever maintains the hosts.
solid answer
~60 sEach `exec.Command` call is a dependency your `go.mod` cannot express: a binary that must exist on every host, at a version whose flags, output format and exit codes you have silently baked in. I keep a subprocess when there is no serious Go client for the thing, when the tool's behaviour is genuinely hard to reimplement correctly, and when it offers a machine-readable output mode and documented exit codes. I move to a library when the alternative is scraping human-readable text, or when the call sits on a hot path where fork-and-exec cost matters. Whatever survives goes behind one internal package that owns argument construction, `exec.LookPath` resolution at startup, and version probing, so the whole surface is auditable and swappable in one place. The organisational half is the real work: the platform team owns the host image and can upgrade a CLI under me, so either the pinned version becomes part of that image contract in writing, or I design for skew and detect it loudly at startup rather than mid-deploy.
go deeper
Know that running an external program assumes it is installed on every machine your code runs on, and that this assumption is invisible in the source. Be ready to say how you would check for it.
Explain the concrete dependencies a call site takes on — binary, flags, output format, exit codes — and how startup resolution with exec.LookPath plus a version probe makes them explicit.
Show how you would contain the surface in one package, detect version skew before any state changes, and diagnose a failure that occurs on one host and not another.
Own the policy and the negotiation: which dependencies are acceptable, how versions are pinned or tolerated, who on the platform side agrees to it, and how a call site is migrated to a library without a big-bang cutover.
## What a subprocess actually costs you A Go program that shells out has two dependency graphs. One is in `go.mod`: versioned, checksummed, resolved reproducibly, reviewable in a diff. The other is invisible — the set of external programs your code assumes are installed, at versions nobody wrote down, on machines somebody else provisions. Every `exec.Command` call site adds to the second graph, and it takes a dependency on four things at once: 1. **The binary's presence**, on every host, in the environment your process actually has. 2. **Its command-line surface** — flag names, flag semantics, subcommand structure. 3. **Its output format**, if you parse anything. 4. **Its exit-code contract**, if you branch on status. Any of the four can change in a patch release you did not request. ## The decision, per call site Ask in order: **Is there a real library?** If the tool is a thin wrapper over an API you can call directly from Go, the subprocess is buying flakiness for nothing. If the tool encodes years of hard-won behaviour, reimplementing it is usually the worse trade. **Is its output a machine contract or human text?** A tool with a structured output mode and documented exit codes is a reasonable dependency. A tool whose output you regex is a dependency that will break silently on a cosmetic change, and "we parse its human output" should be recorded as technical debt with a named owner rather than shipped as normal. **How hot is the path?** Each invocation is a fork and exec plus process startup — negligible once per deploy, ruinous in an inner loop. A per-item subprocess is one of the easiest large regressions to introduce and one of the easiest to spot in a CPU profile. **Who upgrades the host?** This is the question that makes it a leadership call rather than a technical one. If a platform team owns the image, they can move the tool under you between one run and the next, and they were not in the room when you took the dependency. Either you get the version into the image contract as a stated requirement, or you accept the skew and detect it. ## Controls, in the order I would add them - **Resolve everything at startup.** Call `exec.LookPath` for each required binary when the program boots and refuse to start with a message naming what is missing. The default behaviour — the lookup error hiding in `Cmd.Err` until the first `Run` — means a missing tool is discovered halfway through a deploy, after other steps have already mutated state. - **Probe and log the version.** Run the tool's version subcommand once at startup, log it with the run, and fail (or warn loudly) below a stated minimum. When a deploy behaves differently on one host, that log line is the whole investigation. - **One package owns the surface.** Every invocation goes through an internal package that builds argument slices from typed inputs, sets `Cmd.Dir` and a curated `Cmd.Env` consistently, uses `CombinedOutput` (or `Output` when the result is parsed), and turns a nonzero exit into an error carrying both the exit code and the tool's own message. Swapping a tool then touches one file, and a reviewer can see the entire external surface in one place. - **Never build a shell string.** Keep everything as an argument list. Beyond correctness, it means every argument is visible to a reviewer as a discrete value rather than buried in interpolated text. - **Pin what you can actually pin.** In descending order of strength: ship the binary with your artifact; pin it in the host image with the platform team's agreement; state a minimum version and enforce it at startup; or nothing, and hope. Know which rung you are on and say so. ## What to say when someone proposes removing all of them "No subprocesses" is as unserious a position as "shell out to everything". The honest framing is that shelling out is a legitimate way to reuse a mature tool you cannot afford to rewrite, and its cost is a dependency you must manage as deliberately as any other. The test I would apply before adding one: can I name the version I need, name the person who controls it on every host, and describe what my program does when that version changes? If I cannot answer all three, the dependency is not ready to take, whatever the code looks like. ## Migration, when you decide to remove one Moving a call site to a library is rarely a big-bang change. Put both behind the same internal interface, run the library path in shadow mode comparing results, and cut over per call site with the subprocess retained as a fallback for one release. Then delete the fallback — a permanently retained fallback is two dependencies instead of one, and the one you no longer exercise is the one that will be broken when you need it.
- A platform team upgrades a CLI on all hosts and one of your flags is renamed. What should have prevented the outage?A startup version probe with a stated supported range would have failed the run loudly, on the first host, before anything changed — instead of a mid-deploy failure at one call site. Longer term, the version belongs in a written contract with that team, so an upgrade is a coordinated change rather than an ambient one.
- When is scraping a tool's human-readable output acceptable?Only as a temporary, recorded compromise: no structured output mode exists, the alternative is reimplementing the tool, and the parse is isolated in one function with tests against captured real output. It should carry an owner and a plan, because it will break on a change the tool's authors consider cosmetic.
- How do you argue for keeping a subprocess when a reviewer wants a pure-Go rewrite?By pricing the rewrite honestly: the tool's edge cases, the ongoing cost of tracking its behaviour, and the risk of being subtly wrong where the original is battle-tested. Then show the containment — one package, startup resolution, version probe, exit codes handled — so the dependency is managed rather than merely present.
- What belongs in your program's startup log about external tools?For each required binary: the resolved absolute path, the version it reported, and the fact that resolution succeeded. That turns 'it behaved differently on host 7' from an investigation into a one-line diff, and it documents the real dependency set for anyone reading a run afterwards.
saying these in an interview costs you the question
- Treats an external binary as free because it is not in go.mod
- Parses human-readable tool output as a permanent design
- Assumes every host has the same version installed
- Discovers a missing binary only when a step runs
- Argues for banning subprocesses outright with no alternative
- Never records which tool versions a run actually used