Your platform lets non-engineers supply regular-expression rules that run inside request handling: how do you keep one pathological pattern from taking the service down?
answer
- a supplied pattern is code
- eliminate, screen, bound, isolate
- structured predicates beat raw patterns
- benchmark at admission, not in production
- isolation survives the shape you missed
basics
~20 sTreat a supplied pattern as untrusted code. Restrict what authors may express, screen patterns at admission for ambiguous shapes, cap input length, run each match under a budget, and isolate execution so a stuck match consumes one bounded, abandonable slot rather than a request worker.
solid answer
~50 sStart from the judgement that the only complete fix is **removing the ability to express ambiguity**; everything else is containment. So the design question is where you put the boundary. At **author time**, offer structured predicates — starts with, contains, one of a list — and accept raw patterns only where they are genuinely required. At **admission time**, parse a submitted pattern and reject the known-dangerous shapes, plus benchmark it against a generated near-match input before it is ever allowed to run. At **request time**, cap field length, and run the match under a step or time budget where the matching machinery exposes one. At the **isolation** layer, run matching in a bounded pool or a separate process so a hung match costs a slot you can abandon, not a request thread. Then observe per-rule timing and keep a kill switch per rule.
go deeper
Take away the principle: a pattern typed into a configuration screen runs inside your service like code, so it deserves the same suspicion as any other supplied input.
Explain the layers and what each bounds: authoring vocabulary, admission screening, input caps and per-match budgets, and isolated execution, and why none of them alone is enough.
Design the request-time behaviour concretely: where the cap is enforced, what happens when an evaluation exceeds its deadline, and how the failure is reported and observed per rule.
Own the trade-off explicitly — how much expressiveness you withdraw from authors to make an outage class impossible, and which layer you rely on for the shapes nobody modelled.
## Treat a supplied pattern as code A pattern supplied at runtime is a program. It is authored outside the review process, it runs inside request handling, and its cost is a function of an input someone else chooses. Once you accept that framing, the familiar controls apply: restrict the language, screen submissions, bound execution, isolate the blast radius, and observe. The framing also settles the false comfort. 'Our authors are internal and well-meaning' is irrelevant, because the dangerous combination is a well-meaning pattern plus a hostile **input**, and the inputs arrive from the public side. ## The four places a guard can sit | Layer | Control | What it actually buys | What it does not do | |---|---|---|---| | Author time | structured predicates instead of raw patterns | removes the defect class entirely for those rules | costs expressiveness; some rules genuinely need patterns | | Admission time | parse and reject ambiguous shapes; benchmark against a generated near-match | catches the shapes before they ever serve traffic | a screen can be evaded by a shape you did not model | | Request time | input length cap; per-match step or time budget | bounds the cost of a single evaluation | caps are weak against doubling-per-symbol growth | | Isolation | bounded worker pool or separate process for matching | keeps a hung match from consuming shared request capacity | the work still runs; abandoning it may require process boundaries | No single row is sufficient, and the rows differ in kind: the first eliminates, the second screens, the third bounds, the fourth contains. A serious design uses several, and says out loud which failures each one is expected to catch. ## What the screen can and cannot do An admission-time screen parses the submitted pattern and looks for the shapes that create many match paths: a quantifier applied to an already-quantified group, alternation branches under a quantifier whose accepted sets overlap, and unanchored patterns beginning with a quantifier. This catches the overwhelming majority of real submissions, because real dangerous patterns are ordinary ones written carelessly, not adversarial constructions. It is a screen, not a proof. Deciding in general whether a pattern has a pathological input is harder than pattern-matching itself, and a determined author can compose shapes your rules do not model. Pair the structural screen with an empirical one: generate a near-match input from the submitted pattern's own alphabet — a long run the repeated part accepts plus one defeating symbol — and time the match at several lengths before admission. A pattern whose time doubles per added symbol is rejected with an explanation of what to change. ## The trade-off you actually own The decision is not technical, it is about who absorbs the cost of safety: 1. **Expressiveness against availability.** Structured predicates make a whole class of outage impossible and will not express some rules people legitimately want. 2. **Author friction against operator risk.** An admission benchmark costs authors seconds and a rejection they must understand. Skipping it moves that cost to whoever is paged. 3. **Isolation cost against blast radius.** Matching in a separate bounded pool or process adds latency and operational surface, and it is the only layer that reliably survives a shape you failed to anticipate — because it does not depend on having anticipated anything. 4. **Uniformity against local judgement.** One shared evaluation path everywhere is auditable; per-service handling is not. Make the safe path the easy one or it will not be used. A further note on matching budgets: ecosystems differ in whether the matching machinery exposes a step or deadline at all, and a design that assumes one exists everywhere will not survive contact with a service where it does not. Build the isolation layer so the platform is safe either way, and treat a budget as a welcome refinement rather than the foundation. ## What I would ship 1. Structured predicates as the default authoring surface, with raw patterns behind an explicit opt-in and a named owner. 2. Admission-time structural screen plus a growth benchmark; rejection tells the author which shape failed. 3. A hard input cap per rule, enforced before evaluation, sized to the domain rather than to the maximum the storage layer allows. 4. Evaluation on a bounded pool separate from request workers, with a deadline per evaluation and a step budget where available. 5. Per-rule timing and failure metrics, a kill switch per rule, and staged rollout so a newly admitted rule meets a slice of traffic before all of it. 6. A written post-incident rule: any outage traced to a rule ends with a screen improvement, not only with that rule disabled.
- Why is an admission-time screen for dangerous shapes not a guarantee?Because it is a syntactic heuristic. Deciding in general whether some input makes a given pattern blow up is harder than the matching itself, so any screen recognises a list of modelled shapes and misses compositions outside it. It is still worth having — most real offenders are careless ordinary patterns — but it must be backed by an execution bound that does not depend on recognising the shape.
- What does isolating matching in a separate bounded pool actually buy, given the work still runs?Bounded blast radius and the ability to abandon. A hung evaluation occupies one slot in a pool sized for it rather than a request worker, so healthy traffic keeps flowing and the symptom is a rule failing rather than a service falling over. It is the only layer that helps against a shape nobody anticipated, which is why it belongs in the design even when a screen exists.
- When would you refuse raw patterns entirely and offer structured predicates instead?When the rules people actually write are covered by a small vocabulary — starts with, contains, one of a list, numeric range — and the surface is exposed to authors outside engineering. You trade a little expressiveness for the elimination of a whole outage class, and you keep an explicit, owned escape hatch for the rare rule that genuinely needs a pattern.
saying these in an interview costs you the question
- Trusts internal authors and skips screening entirely
- Treats an input length cap as a complete defence
- Assumes every matching implementation exposes a time budget
- Relies only on a shape screen with no execution bound
- Runs supplied patterns on the request worker pool
- Closes an incident by disabling one rule and changing nothing else