Go's regexp panics on a rule set ported from a PCRE engine at proxy start-up — how do you triage which rules are unportable and rewrite them?
answer
- read the parse error, do not guess
- collect every failure in one pass
- group the failures by construct
- one lookahead becomes two patterns
- the rules that compile are the dangerous ones
basics
~20 sStop the crash-loop, then compile every rule in a test so failures surface in CI rather than at start-up. Group the parse errors by construct, rewrite lookarounds as two-stage matches in Go, and add traffic fixtures for rules that still compile.
solid answer
~50 sThe panic text comes from the `regexp/syntax` parser and names the offending construct — "invalid or unsupported Perl syntax" followed by the lookahead fragment, or an invalid escape for a `\1` backreference — so the first move is to collect them rather than fix them one at a time: run the whole rule set through `regexp.Compile` in a table-driven test and print every failure with its rule id. That converts a start-up crash-loop into a build failure and shows the true blast radius. Then triage by construct. Most negative lookaheads become two patterns and a boolean in Go: match the broad rule, then exclude with a second pattern. Backreference rules generally cannot be expressed and need a capture plus a string comparison in code. The subtler half of the job is the rules that *do* compile: modes passed outside the pattern in the old engine are gone, and `$` in Go matches only at end of text, so those rules need fixtures from captured traffic before rollout.
code
go · 7 lines// original rule: ^/api/(?!health)
var apiPath = regexp.MustCompile(`^/api/`)
var healthPath = regexp.MustCompile(`^/api/health(/|$)`)
func filtered(path string) bool {
return apiPath.MatchString(path) && !healthPath.MatchString(path)
}go deeper
Know where to look: the panic text names the exact construct the pattern parser refused, and the fix belongs in the pattern, not in the code that calls it.
Be able to convert the crash into a test that compiles every rule and reports all failures at once, and to rewrite a negative lookahead as two patterns joined by a boolean in Go.
Show that you weight the silent failures higher than the loud one — lost inline flags, anchor differences, changed selection — and that you demand traffic fixtures plus a shadow comparison before cutting a filtering path over.
Own the upstream fix: who declares the accepted pattern dialect, where rules are validated in their lifecycle, and what the rollout contract is for a change that can start blocking legitimate traffic.
A rule set written for a PCRE-style engine and dropped into a Go proxy fails in two distinct ways, and the loud one is the easier one. ## Failure one: it does not compile Go's `regexp` rejects unsupported syntax while parsing, so a rule using a lookahead never becomes a usable `*regexp.Regexp`. If those patterns are compiled at start-up, the process dies before it serves a request, and under an orchestrator that means a crash-loop on rollout. The error is precise and worth reading rather than skimming. `regexp/syntax` reports `error parsing regexp: invalid or unsupported Perl syntax:` followed by the fragment it choked on, so `(?=`, `(?!` and `(?<=` are all named explicitly; a `\1` backreference is reported as an invalid escape sequence. That text is the triage key. **Get the whole list at once.** Instead of restarting and reading one panic at a time, write a test that walks the rule set and calls `regexp.Compile` on every pattern, collecting failures into a report keyed by rule id and error text. Two things fall out immediately: the count of genuinely unportable rules, which is usually far smaller than the panic makes it feel, and a grouping by construct, which tells you how many distinct rewrites you actually have to invent. Wire the same check into CI so a new rule with the wrong dialect fails the build, and keep it as a start-up self-check so a rule set loaded from config at run time is validated before the proxy declares itself ready. **Then rewrite by construct.** - *Negative lookahead* — "match X but not when Y follows" — becomes two patterns and ordinary Go: match the broad pattern, then reject on the second. This covers the large majority of real rules, is more readable, and each half is independently testable. - *Positive lookahead* is often just a group you can consume and ignore, once you accept that the match extent changes; if a caller depends on the reported offsets, do it in two stages as above. - *Lookbehind* usually becomes an anchored prefix check on the surrounding string rather than part of the pattern. - *Backreference* rules — the same token appearing twice — cannot be expressed at all. Capture with a group and compare in Go. - *Atomic groups and possessive quantifiers* exist in the source dialect only to control a backtracker's cost. There is nothing to port: delete the annotations and the pattern behaves the same, because the engine does not backtrack. ## Failure two: it compiles and behaves differently This is the one that reaches production. A pattern can be perfectly valid in both dialects and still select different text. - **Lost flags.** If the old engine took modes as a parameter, they did not travel with the pattern string. In Go a mode must be written inline, and without `(?i)` a rule silently becomes case-sensitive; without `(?s)` a `.` no longer crosses a newline. Nothing errors — the rule just stops firing. - **Anchors.** Go's `$` matches only at the end of the text, not also before a trailing newline as some engines allow. A rule ending in `$` against a header value with a stray newline now behaves differently. - **Selection.** If any part of the rule set came from POSIX tooling rather than PCRE, it may assume longest-match while Go defaults to first-alternative-wins. None of these can be caught by compiling. They need fixtures: capture a sample of real request URLs and header values, record what the old engine decided for each, and assert the Go implementation agrees. That corpus is worth more than any amount of rule reading, and it is also the regression suite for the next rule change. ## Getting it out safely Sequence the rollout so the two failure modes are separated. Ship the compile check first, on its own, so the dialect problem is a build artefact rather than an incident. Then run the new matcher in shadow mode next to the old decision — evaluate both, log disagreements, act on the old one — until the disagreement rate is explained rather than merely small. A rule that fires more often than before is as much a defect as one that stopped firing; a filtering proxy that suddenly blocks legitimate traffic is the worse outcome of the two. Finally, close the loop upstream. Whoever authors the rules is still writing them in the old dialect unless something tells them not to. Declaring the accepted dialect, validating patterns at ingest rather than at deploy, and rejecting a rule at submission time is what stops this from being an annual event. ## What to say in the interview The answer an interviewer is listening for has four beats: read the actual error rather than guessing, convert the crash into a test that reports every failure at once, rewrite by construct with two-stage matching as the default technique, and then spend the larger share of the effort on the rules that compiled — because those are the ones that will otherwise be discovered by a customer.
- Which ported rules worry you more: the ones that panic or the ones that compile?The ones that compile. A panic is loud, complete and fixable before traffic touches it; a rule whose flags did not travel, or whose anchors mean something slightly different, keeps running and quietly stops matching what it used to. Those need fixtures captured from real traffic and a shadow comparison against the old engine, not a compile check.
- How do you stop the same dialect skew from recurring with the next rule?Validate at the point rules are authored or ingested, not at deploy. Compile every submitted pattern and reject it there with the parse error attached, state the accepted dialect in the rule documentation, and keep the compile sweep in CI as a backstop. Otherwise rule authors keep writing the dialect their previous tooling accepted.
- A rule uses an atomic group to stop the old engine from backtracking. What do you port it to?Nothing — drop the annotation and keep the rest of the pattern. Atomic groups and possessive quantifiers exist to bound a backtracking engine's exploration. Go's engine does not backtrack, so those constructs have no behaviour to preserve; the plain pattern matches the same text.
- How would you roll the new matcher out to a live filtering proxy?In shadow first: evaluate both the old and new decision per request, act on the old one, and log every disagreement with the rule id and the input. Ship the compile check separately and earlier, so dialect breakage is a build failure. Cut over only when the disagreements are explained, not merely rare.
saying these in an interview costs you the question
- Fixes the patterns one panic and one restart at a time
- Assumes every rule that compiles behaves as it did before
- Tries to recover from the panic instead of validating rules
- Ports atomic groups and possessive quantifiers literally
- Treats lost inline flags as harmless because nothing errors
- Cuts over without any comparison against the previous behaviour