As security lead, what justifies adopting a policy engine when configurable scanners and shell checks already work?
answer
- the individual check is not the artifact
- what can the current setup not express?
- count the enforcement points
- who reads the rules besides the author?
- toggling shipped checks is a config file
basics
~20 sAdopt when the rule set is the artifact: conditions no shipped catalogue covers, one rule holding at several enforcement points, readers outside the team. If the need is only toggling existing checks, it is a linter with a config file.
solid answer
~50 sI would not adopt an engine to replace checks that already pass. The trigger is a need the current setup cannot meet: conditions specific to us that no shipped catalogue contains, a rule that has to hold at more than one enforcement point, or people outside the team — an auditor, another platform group — who need to read what we enforce without reading code. Against that sits real cost: a language nobody here writes yet, a component to run and upgrade, and a review population of one until others learn it. If what we actually want is to disable one shipped check in a repo and tune a threshold in another, we want a linter with a config file, and an engine buys complexity instead. So I start with the handful of rules the current setup genuinely cannot express and let the case prove itself on those.
go deeper
Know that a configurable scanner ships a catalogue of checks you switch on and off, while a policy engine expects you to write the conditions yourself — and that these solve different problems.
Be able to name what an engine adds over a scanner's config file: conditions specific to your estate, one statement reused at several enforcement points, and a rule set that can be enumerated.
Weigh the real costs — an unfamiliar language whose mistakes are silent, a component to operate, and migration work that reproduces verdicts you already have — against the capability gained.
Own the decision and its reversal criteria: what you would adopt narrowly first, what evidence a year on would show the call was right, and what you would tell the team if it was not.
## The question is not "is policy as code good" A team with a working check suite and a configurable scanner already has automated guardrails. Adopting a general policy engine on top is an organisational decision with a price, and the interviewer is listening for whether you can name the price and the trigger, not for enthusiasm. ## The linter-with-a-config line The cleanest way to sort the situation: **is the thing we need a new rule, or a different setting on an existing one?** If every need on the list is "turn check 17 off in this repository", "raise this threshold to 35 days", "exclude the sandbox account", then the requirement is configuration of checks that already exist. Tools that ship a catalogue plus a config file do that, and doing it in a policy engine means reimplementing a catalogue somebody already maintains. Adopting an engine here is a straight loss: same coverage, more moving parts. The line is crossed when the needs stop being settings: - **Conditions nobody ships.** "A managed database owned by a regulated service keeps at least 35 days of backups, and 7 is acceptable only in the sandbox estate." No general catalogue encodes your estate's structure. - **One condition, several enforcement points.** The moment the same retention requirement must hold on a proposed change and again on the running estate, having it stated once as data stops being elegance and becomes maintenance arithmetic: one statement instead of two copies of `35` that drift. - **Readers who do not read code.** When someone outside the team has to see what is enforced, a list of named rules answers them and a suite of scripts does not. - **Volume.** A dozen custom conditions scattered as script branches is a catalogue that nobody can enumerate. That is the point where the *rule set*, rather than any rule, is the artifact you need. ## The costs, stated to the people who will pay them **A language.** Rule languages are small but genuinely unfamiliar, and their failure mode is quiet: a rule that does not match produces no result rather than an error, so a beginner's mistake looks like a pass. Expect a period where the team writes rules that do nothing. **A component.** An engine is software you run, upgrade and debug, wherever you run it. Whoever adopts it inherits that. **A review bottleneck.** At the start, one person can review rule changes. Until that is two or three people, the engine is a single point of knowledge as well as a single point of failure. **Migration time that buys nothing new.** Rewriting checks that already work produces the same verdicts as before. The benefit arrives only for the checks that needed the new capability. ## How I would actually decide Count three things and be honest about the count: how many conditions the current setup cannot express, how many enforcement points each condition must hold at, and how many people outside the authoring team need to read the rule set. If those numbers are one, one and zero, keep the scripts and the scanner and say so. If they are ten, two and several, the engine pays for itself and the argument writes itself. Then start narrow: put the conditions that fail the current setup into rules, leave the working checks alone, and revisit in a quarter. The scripts are not a moral failing to be cleansed; they are working code, and their replacement costs weeks a security team rarely has spare. ## How you know a year later that you were wrong The rule set is never cited outside the team; every rule still evaluates at exactly one place; nobody but the original author has written a rule; and the loudest incident of the year was the engine itself. That is a runtime you bought rather than a policy capability you built — and noticing it is worth more than defending the original decision.
- Your team writes the rules and nobody outside ever reads them. Does the case still hold?It gets weaker. Inspectability by outsiders was half the argument, so what remains is reuse across enforcement points and a rule set you can enumerate. If the rules also run at one place only, the engine is competing with a script purely on expressiveness, and the script often wins.
- How would you know a year later whether adopting it was the right call?Look for evidence the rule set became an artifact: it is cited outside the team, rules moved between enforcement points without a rewrite, and more than one person authors them. If none of that happened and the engine caused the year's worst incident, you bought a runtime, not a capability.
- A vendor scanner already flags short backup retention. Does that settle it?Only for that one condition, and only where the scanner runs. It does not cover conditions that depend on your estate's own structure, and it does not give you the same statement enforced at a second point. Use the shipped check where it fits and reserve authored rules for what it cannot express.
saying these in an interview costs you the question
- Adopts an engine because the ecosystem is popular
- Ignores that the engine becomes a component to operate
- Confuses tuning shipped checks with authoring policy
- Assumes every team will learn the rule language
- Counts rewriting working scripts as progress