skip to content

Go's regexp cannot express your rules' lookaheads — how do you decide between rewriting them and adopting a backtracking regex dependency?

level: principalimportance: should knowfreq 26%

answer

  1. count before you argue
  2. most rule sets fail on a small minority
  3. you would be giving up a guarantee, not just adding code
  4. who owns it after the migration
  5. the dialect drifts and the door closes

basics

~20 s

Measure first: compile the whole rule set and count how many rules truly need the missing constructs. Rewriting a handful beats owning an engine with an unbounded worst case, a dependency to govern, and a rule dialect that is hard to walk back.

solid answer

~60 s

Start with a number, not an argument: compile every rule with `regexp.Compile` and count the failures by construct. In most rule sets it is a small minority, and two-stage matching in Go covers most of them, which turns the question into a week of work rather than a dependency decision. Take the dependency only when the unportable share is large or structural — a dialect you do not control, rules authored by another team or a vendor. Then the cost is real and ongoing: an engine whose worst-case matching time is no longer bounded, on input that arrives from untrusted clients; a module that must be maintained, security-reviewed and carried across toolchain upgrades, possibly through cgo; and the fact that once rule authors can write lookaheads, the rule set drifts and the decision stops being reversible. Prefer the containable middle paths — declare the accepted dialect at ingest, keep the stdlib package in the hot path, or isolate the other engine behind one service — and write the choice down with the person who governs dependencies, because that is who can overrule it.

go deeper

for a junior

Understand the two sides in plain terms: rewriting rules costs work now, while a different engine adds a dependency and gives up the guarantee that matching time stays bounded.

for a middle

Be able to argue from the measurement — how many rules actually fail and on which construct — and to show that two-stage matching handles most negative lookaheads without changing engines.

for a senior

Demonstrate that you weigh the ongoing costs: worst-case runtime on untrusted input, maintenance and review of the dependency, build implications if it uses cgo, and the fixtures needed to prove either path preserves behaviour.

for a principal

Own the decision as policy: the measurement, the recommended middle path, the revisit conditions, and the people who must agree — the rule authors whose dialect you are constraining and the dependency owner who can overrule you.

This is a dependency and policy decision wearing a syntax problem's clothes, and it should be made with data and with the right people in the room. ## Establish the real size of the problem Before anyone argues, compile the entire rule set and produce a table: rule id, error, construct. The important outputs are the *share* of rules that fail and the *shape* of the failures. A rule set where four of six hundred rules use a negative lookahead is not a dependency question at all — it is an afternoon of two-stage matching plus tests. A rule set where a quarter of the rules use lookbehind and backreferences because they were generated by tooling you do not own is a genuinely different situation. Teams routinely skip this step and debate the architecture of a problem they have not measured. ## What you take on by adopting a backtracking engine **An unbounded worst case on untrusted input.** The stdlib package's guarantee is that matching time is bounded by input length times pattern size, whatever the pattern. Give that up and the ceiling is now set by the interaction of a pattern someone writes and an input someone else sends. On a proxy filtering traffic from untrusted clients, that is a property you were relying on without noticing. **A dependency with a lifecycle.** Somebody has to own it: watch its maintenance, review its releases, keep it building across Go toolchain upgrades, and answer for it in whatever review governs third-party code. If it binds to a C engine through cgo, it also changes build and deployment properties — cross-compilation, static linking, container images, and a class of crash that no longer produces a Go stack trace. **Irreversibility.** This is the cost most teams underestimate. The day rule authors can write lookaheads, they do. Six months later the rule set is written in the richer dialect, and going back means rewriting all of it, not the four rules you started with. Adoption is a one-way door with a slow closing time. ## What you take on by rewriting **Translation work that must be verified.** Two-stage matching is straightforward, but each rewrite needs fixtures showing the old and new rules agree on real inputs, and the effort scales with the number of rules rather than the number of constructs. **A permanent constraint on rule authors.** Some rules genuinely cannot be expressed and become code. If rules are authored by a different team, you are exporting a restriction to them, which needs their agreement rather than your decision. **Continuing pressure.** The request will come back. Without a stated policy and validation at the point rules are written, the next person meets the same wall and reopens the debate. ## The middle paths, which usually win - **Declare the dialect and validate at ingest.** Reject a pattern when it is submitted, with the parse error attached, instead of at deploy. The constraint becomes visible where it can be acted on cheaply. - **Two-stage matching as the standard technique.** Give the team one documented pattern for negative lookaheads and it stops being a per-rule invention. - **Split the tiers.** Keep the stdlib package on the hot path where every request is evaluated, and route the small unportable set through a separate, contained mechanism — an offline job, a lower-volume path, or a service whose blast radius and resource limits are its own. - **Generate rather than interpret.** If rules come from a spec, compile them into Go once at build time rather than adopting an engine to interpret a foreign dialect at run time. ## How to actually decide Write it down in one page: the measured share of unportable rules, the two options with their ongoing costs, the middle path you recommend, and the conditions under which you would revisit — for example, if the unportable share crosses some threshold, or if rule authorship moves outside the team. Take it to whoever governs dependencies, because a new engine in the binary is their call as much as yours, and a decision that survives being overruled is one you documented rather than one you shipped quietly. And set a review date. The right answer here is a function of who writes the rules and how many need the missing constructs; both of those change, and a policy with no revisit date becomes folklore that nobody can defend or reverse.

  • What single measurement would you take before this discussion starts?
    The share of rules that fail to compile, broken down by construct. Run the whole set through the compile step and tabulate it. If four of six hundred rules fail, this is an afternoon of two-stage matching; if a quarter fail because tooling you do not own generates them, it is a real dependency decision. Almost every unproductive version of this debate skips that number.
  • Why is adopting a backtracking engine hard to reverse later?
    Because the capability changes what people write. Once rule authors can use lookaheads and backreferences, the rule set drifts into the richer dialect, and reverting means rewriting everything rather than the handful of rules that prompted the change. The migration cost grows with time, so the decision should be made as a one-way door, with the review conditions written down.
  • Who should be in the room for this decision, and who can overrule you?
    The team that authors the rules, because a dialect restriction is exported to them; whoever governs third-party dependencies and security review, because a new matching engine in a request path is theirs to approve; and whoever carries the operational risk of the service. Any of the latter two can overrule the engineering preference, which is why the case is written down rather than argued in a thread.

saying these in an interview costs you the question

  • Adopts an outside engine before counting the failing rules
  • Treats the linear-time guarantee as a detail rather than a property being given up
  • Ignores who maintains and security-reviews the new dependency
  • Assumes the decision can be reversed cheaply later
  • Frames it as a personal preference instead of a documented policy
  • Overlooks the cgo build and deployment consequences