Your rule engine runs customer rules as a tape machine with unbounded scratch space; what termination guarantee should you promise?
answer
- the guarantee is manufactured, not discovered
- meter the machine, not the clock
- space budget does not bound time
- exceeded is not rejected
- the published bound becomes a contract
basics
~20 sPromise a bounded run, not a finishing one. Meter the machine in steps, publish the bound, and report exhaustion as its own outcome. The guarantee is manufactured by enforcement, never discovered by watching a rule run.
solid answer
~50 sThe guarantee you can actually keep is: every rule either halts or is stopped at a published bound. Meter the computation in its own steps rather than elapsed time, so the same rule on the same input gets the same verdict on any host — wall-clock metering makes the outcome depend on load and hardware and is not reproducible. Bound scratch space too, but separately and for a different reason: space caps memory, it does not cap time, since a run confined to a region can circle inside it forever. Make budget exhaustion a distinct, documented outcome — never folded in with a rule error — and capture the configuration at the kill so the author can see where the machine was circling. The real cost of the decision is that the bound becomes part of your public contract.
go deeper
Understand that a runaway computation is stopped from the outside by a limit, not by noticing it has gone wrong. Someone has to choose that limit, and choosing it is a design decision with consequences.
Explain why a step count and elapsed time are different meters, and why a limit on scratch space bounds memory without bounding how long a computation can circle inside it.
Show the operating side: exhaustion reported as its own outcome, the machine's position captured at the kill, and enforcement cheap enough that metering does not dominate the cost of running ordinary rules.
Own the contract. The bound decides which workloads you decline to serve, it is far easier to raise than to lower once authors depend on it, and it sits alongside rate and concurrency limits rather than replacing them.
## What you can and cannot promise The tempting promise is 'every rule we accept will finish'. You cannot keep it by looking at rules: a rule that has been running a while and a rule that will never stop produce the same observation at every finite moment, and no length of run distinguishes them. The guarantee has to be **manufactured by enforcement** rather than discovered by inspection. The promise you can keep is narrower and completely solid: *every execution terminates, either because the rule halted or because the platform stopped it at a bound we publish*. That is checkable, testable and explainable to a customer — and it is a different sentence from the one people expect, which is why it belongs in the contract in writing. ## Choosing the meter | meter | what it buys | what it costs | |---|---|---| | steps of the machine | a verdict that is a property of the computation: same rule, same input, same outcome anywhere | a counter checked on every step, so a throughput tax on every rule, including the well-behaved ones | | scratch space used | a hard cap on memory, and a region small enough to make a loop check conceivable | no bound on time at all: a run confined to a region can circle inside it forever | | elapsed wall-clock time | trivial to implement, and directly matched to the latency you actually owe callers | irreproducible: the same rule passes on an idle host and is killed on a loaded one, and authors cannot test against it | Step metering is the one that makes the guarantee meaningful, because it is intrinsic to the computation. Wall-clock still has a place as an outer safety net for the operator, but it should not be the thing the contract names, or the contract becomes a statement about your hardware. ## The decisions that are actually yours 1. **Which meter the promise names.** Steps for the contract; space as a separate memory cap; wall-clock as an operational backstop. 2. **What number.** Derive it from observed distributions of real rules, not from a round figure, and keep enough headroom that ordinary rules never approach it. 3. **What outcome exhaustion produces.** A distinct result, documented and separately observable — not the same failure a malformed rule returns. Collapsing them destroys the customer's ability to tell 'my rule is wrong' from 'my rule is too big'. 4. **What you record at the kill.** The configuration is what makes the report actionable: the state, the head position, a window of the scratch tape around it, and the counts of which entries fired most. A bare timeout tells the author nothing. 5. **Whether the bound is per rule or per tenant.** A per-rule bound is predictable for authors; a shared budget protects the platform but makes one tenant's behaviour visible in another's outcomes. ## The trade-off a lead owns - **Expressiveness against predictability.** Any bound makes some legitimate long computation impossible. That is the point of a bound, and the honest version of the decision is choosing which workloads you are declining to serve, not pretending none are excluded. - **A bound is a one-way promise.** Once customers build rules that fit it, raising it is easy and lowering it breaks them. Set it knowing you are effectively fixing the ceiling for a long time, and consider publishing it as a floor you will not go below rather than a value you may tune. - **Metering costs throughput.** Checking a counter on every step is a real tax; batching the check coarsens the bound. Decide deliberately how tight the enforcement has to be, because 'within a few thousand steps of the limit' is usually fine and much cheaper. - **Bounded does not mean cheap.** A rule that runs to the limit on every input is within contract and can still ruin capacity planning, so the bound belongs alongside rate limits and concurrency caps rather than instead of them. ## What good looks like An author writing a rule can find the published step budget, run their rule against the same meter locally, and get the same answer the platform will give. When a rule is stopped, they receive an outcome that says so in its own words, with the machine's position at the moment it was stopped. Nobody in the system ever claims to know that a rule would have run forever — only that it did not finish inside the budget, which is the strongest true thing anyone there can say.
- Why not simply reject rules that look like they might loop?Because the inspection cannot be made trustworthy. A syntactic check either rejects legitimate rules that happen to look risky or misses the ones that do not, and either way it tells the author something is wrong with a rule that may be fine. Enforcement at run time is honest about what is being claimed: a bound, not a verdict.
- Does capping the scratch tape make the step budget unnecessary?No. A capped region has finitely many configurations, so a non-halting run must eventually repeat one, but that count is astronomically large in the width of the region. Space caps memory; only a step budget caps time. They are two separate promises and should be two separate numbers.
- How do you set the initial budget when there is no traffic yet?Instrument first and promise later: run early rules with the meter switched on but no enforcement, and publish a bound once you can see where real work sits. If a number must be published on day one, choose it high and describe it as a floor you will not lower, because raising a bound is painless and lowering one breaks working rules.
saying these in an interview costs you the question
- Promises every accepted rule has been checked to terminate
- Meters wall-clock time and calls the outcome reproducible
- Reports a budget kill as an ordinary rule error
- Assumes a capped scratch region also caps the running time
- Sets a bound expecting to lower it once traffic arrives