skip to content

In Rego, what does the walk built-in produce as it traverses a nested document?

level: juniorimportance: should knowfreq 55%

answer

  1. one result per node, not per leaf
  2. you get two things back
  3. location plus subtree
  4. root included, path is []
  5. keys are strings, indexes are integers

basics

~20 s

walk emits one [path, value] pair for every node in the document, starting with the root itself. The path is an array of object keys and array indexes; the value is the whole subtree sitting at that path.

solid answer

~40 s

`walk` is a relation, not a plain function: you write `walk(input, [path, value])` in a rule body with `some path, value` declared, and the body then evaluates once for every node in the document. Each iteration binds `path` to an array of steps from the root — object keys as strings, array indexes as integers — and `value` to the subtree at that path. The root is included, with `path` equal to `[]` and `value` equal to the entire document. That is what makes `walk` the tool for "find this field wherever it is": you never have to enumerate the levels. The cost is that you now see every node, so you filter, typically with `is_object(value)` plus a field reference that simply goes undefined on nodes that lack it.

go deeper

for a junior

Be ready to say what the two bound variables are and to write the some path, value plus walk(input, [path, value]) pairing from memory.

for a middle

Explain that walk is a relation, so the body runs once per node, and that the stream includes the root and every intermediate container, not just leaves.

for a senior

Show the judgment: walk when depth is unknown, a direct reference when the path is fixed, and always use the path half of the pair to make the finding locatable.

for a principal

Own the readability cost. A codebase where every rule walks is one nobody can review; set the expectation that walk is justified by unknown depth, not used as the house style.

## The problem `walk` solves Most Rego rules address a known path: `input.spec.containers[_].image`. That works when you know the shape of what you are handed. It stops working the moment the thing you care about can appear at a depth you did not enumerate — a manifest bundle where one document is a list wrapping other objects, a custom resource that embeds a full object under one of its fields, a template block nested inside another template block. A fixed-path rule silently finds nothing there, and finding nothing in Rego does not look like a failure. `walk` removes the need to know the depth. ## What it actually does `walk(x, [path, value])` is a *relation*: rather than returning a single result, it produces many, and the rule body is evaluated once per result. Written out: ``` some path, value walk(input, [path, value]) ``` Each evaluation binds: - `path` — an array of steps from the root of `x` down to this node. Object keys appear as strings, array indexes as integers, so a path can be mixed: `[0, "spec", "template", "metadata"]`. - `value` — the subtree at that path. It may be an object, an array, or a scalar. The traversal includes the root: the first pair you get is `[]` paired with the entire document. It also includes every intermediate container, not just leaves. So over `{"a": {"b": 1}}` you get three pairs: `[]` with the whole object, `["a"]` with `{"b": 1}`, and `["a", "b"]` with `1`. ## Filtering the stream Because you see everything, a `walk`-based rule is mostly a filter. Two habits matter: 1. **Guard the shape.** `is_object(value)` (or `is_array`, `is_string`) narrows the stream before you look at fields. Strictly this is optional — referencing a field on a scalar is *undefined* in Rego, not a type error, so the iteration just yields nothing — but the guard makes the intent readable. 2. **Let undefined do the work.** `value.apiVersion` on a node with no `apiVersion` is undefined, so that iteration produces no result and evaluation moves on to the next node. You do not need an existence check before the comparison. ## Paths are data, and that is the point The reason `walk` beats a hand-rolled recursive helper for reporting is the `path` half of the pair. A message can name exactly where the offending value lives — `"at [3, \"spec\", \"template\", \"spec\"]"` — which turns a bare "this repo has a problem" into an actionable list a team can work through. If you want a dotted string rather than an array, build one with a comprehension that converts each step with `sprintf("%v", [step])` before `concat`, because `concat` needs strings and array indexes are integers. Do not depend on the order pairs arrive in. Collect results into a set (a partial rule that `contains` messages) and `sort` them if you need stable output for a diff or a report. ## Where `walk` is the wrong reach `walk` visits every node in the document. When the path is known and fixed, a direct reference is shorter, self-documenting, and does far less work — `input.spec.containers[_].image` says what it means in a way `walk` plus three guards never will. Reserve `walk` for the case it exists for: the field can be anywhere, and you would otherwise be writing one rule per level and still missing one. One blind spot to keep in mind: `walk` descends through the *parsed* document only. If a node's value is a string that happens to contain an embedded JSON or YAML document, `walk` sees a string and stops. Decoding that string first is a separate step. ## What an interviewer is checking That you know `walk` gives you both halves — the location and the value — that the root is included, that the stream contains containers as well as leaves, and that you reach for it because the depth is unknown rather than as a default way to read input.

  • Does walk include the root document itself, and why would that matter?
    Yes. The first pair binds `path` to `[]` and `value` to the entire document. It matters because a rule that assumes `value` is a small object will also be handed the whole bundle, so any predicate you apply has to be one that is simply undefined or false for the root rather than accidentally true for it.
  • How do you turn a walk path into a readable location string?
    Convert each step before joining, because array indexes are integers and `concat` only accepts strings: `concat(".", [sprintf("%v", [step]) | some step in path])`. For a first pass it is fine to interpolate the raw array with `sprintf("%v", [path])` — the array form is already unambiguous in a report.
  • When would you not use walk?
    When the path is known. `input.spec.template.spec.containers[_].image` is clearer to the next reader and touches only the nodes it needs, while `walk` visits every node in the document and then filters back down to the same set. Use `walk` when the field can appear at a depth you cannot enumerate.

A fixed path is giving someone a shelf number. walk is handing them a torch and asking them to report back every box in the building along with the aisle it was standing in.

saying these in an interview costs you the question

  • Thinks walk only visits leaf values
  • Thinks walk returns values without their paths
  • Assumes the path comes back as a dotted string
  • Expects walk to skip the root document
  • Reaches for walk even when the path is fixed and known

context