skip to content

You are converting a 400-line shell check suite to policy rules. Which checks should stay scripts, and why?

level: seniorimportance: should knowfreq 52%

answer

  1. split gathering from deciding
  2. is there a verdict, or just a fact?
  3. paging, joining, retrying stay procedural
  4. no print statement, no stack trace
  5. a silent rule looks like a passing rule

basics

~20 s

Keep the procedural parts — paging a listing API, joining two systems, retrying, waiting — because they gather facts rather than decide anything. A pure comparison, such as retention of at least 35 days, is what becomes a rule.

solid answer

~50 s

I split each check into gathering and deciding. The decision — is retention at least 35 days — is a comparison over facts, and that is what belongs in a rule. The gathering — paging a listing API, joining an inventory to another system, retrying a flaky call, waiting for consistency — is procedure; a rule language buys nothing there and costs plenty, so it stays as code that produces the engine's input. I also leave a check as a script when it is genuinely one-off, or when the debuggability tax outweighs the gain: a declarative rule has no print statement and no stack trace, and a rule whose condition never holds produces no result rather than an error, so it goes quiet instead of red. Rules earn their place where the condition is stable, stated often, and has to be readable by someone outside the team.

go deeper

for a junior

Know that a policy rule states a condition over facts it is given, and that fetching those facts — calling an API, paging through results, retrying — is ordinary code that sits outside the rule.

for a middle

Explain the gather-versus-decide split with an example, and be able to say why a rule language handles a comparison well and control flow badly.

for a senior

Demonstrate the judgment call per check and be candid about the debuggability tax, including that an unsatisfied rule produces no result rather than a failure and so goes quiet rather than red.

for a principal

Own the end state: which half of the suite becomes rules, where the seam sits, what the hybrid costs in maintenance, and how you justify not rewriting the scripts that work.

## Split each check into gathering and deciding A 400-line check suite is almost never 400 lines of decision. It is mostly collection: calling something, paging through results, parsing them, joining them to another source, retrying when a call fails, skipping resources that are still being created. The decision hiding inside is usually one line — `retention >= 35`. That split is the whole answer to which checks move. - **The decision moves.** A comparison over facts is exactly what a rule expresses well, and expressing it as a rule is what gives it a name, a per-resource result and a second life at another enforcement point. - **The gathering stays procedural.** Paging, joining, retrying and waiting are sequences of steps with state; a rule language is deliberately not good at them, and the workarounds are worse than the script was. So the honest end state for most teams is not "we replaced the scripts", it is "the scripts now produce an input document, and the conditions live in rules". ## Checks that should stay entirely in a script **Correlation across systems.** "Every production database has a named owner in the service registry and an active retention exception record if it is below the threshold" involves two lookups and a join. The comparison at the end could be a rule, but only if you are willing to assemble both facts first — and if you are not, the script is the honest answer. **Anything with retries, waits or paging.** These are control flow. Squeezing them into a declarative language produces something nobody can read, which destroys the main reason for moving. **One-off and throwaway checks.** A check written for one migration, run at one place, read by its author, does not need a rule set around it. "We rewrote it for consistency" is not a benefit; it is a cost with a nice name. **Checks whose value is the remediation, not the verdict.** If the code's job is to go and fix something, it is not a decision — it is an action, and the rule form has nothing to offer it. ## The debuggability tax, stated plainly This is the part candidates skip and interviewers wait for. Moving a check into a rule language costs you: - **No prints, no stepping.** You cannot scatter output through an evaluation the way you can through a script. - **Silence instead of errors.** A rule body that cannot be satisfied — the field is missing, the type is wrong, the path was renamed — usually produces *no result*, not a false and not an exception. No result reads as "nothing to report". A rule can therefore stop working and look exactly like a rule that is passing. - **Harder bisection.** With a script you run it with different inputs and watch. With a rule, you have to construct the input document that would trip it, which is more work up front. The practical mitigation is to treat "this rule now flags nothing, ever" as suspicious in its own right, and to keep at least one input around that each rule is known to reject. A rule that nothing can trip is a rule you cannot trust. ## How to decide, check by check Four questions, in this order: 1. **Is there a decision worth naming?** If the code's output is a fact rather than a verdict, it is collection. 2. **Is the condition stable?** A threshold that changes with the season and a condition rewritten every quarter both fight the rule form's advantage, which is that the statement stays put. 3. **Does anyone outside the team need to read it?** If yes, the rule form wins even when the script works, because the script's condition is buried in a procedure. 4. **Does it have to hold at more than one enforcement point?** If yes, the rule form wins on reuse alone — this is usually the argument that actually closes the decision. If all four say no, the script stays, and being able to say that without embarrassment is the mark of someone who has run this migration rather than read about it. ## The hybrid's own cost A split suite means two places to look when something is wrong: the collector and the rule set. Keep the seam explicit — the script's job ends when it emits the input document, and every verdict comes out of the engine — and write down which side owns what. A hybrid where some verdicts come from `exit 1` in the collector and others from rule results is worse than either pure option, because nobody can answer "what do we enforce?" from one place.

  • How do you keep the hybrid from becoming two places to look?
    Make the seam explicit: the script's job ends when it emits the input document, and every verdict comes out of the engine. The moment some failures come from an `exit 1` inside the collector and others from rule results, nobody can answer "what do we enforce?" from one place.
  • A rule stopped flagging anything after a field in the input was renamed. How would you catch that?
    Not from the pipeline going green — it will look healthy. An unmatched rule produces no result, not a failure. Treat a rule that has flagged nothing for a long time as a signal to investigate, and keep an input around that each rule is known to reject.
  • Is leaving a 400-line script exactly as it is ever the right call?
    Yes. If it runs at one enforcement point, nobody outside the team reads it, and it works, a rewrite spends real weeks to buy inspectability and reuse that nobody has asked for. Move the checks that fail one of those tests and leave the rest alone.

saying these in an interview costs you the question

  • Insists every check must become a declarative rule
  • Writes retries and API paging into the rule language
  • Assumes a rule that flags nothing is a rule that passes
  • Ignores the debuggability cost of declarative rules
  • Rewrites a working one-off script for consistency alone

context