A Go log-ingest sidecar's CPU profile is dominated by regexp compilation rather than matching. How do you diagnose and fix it?
answer
- parsing frames, not matching frames
- a fixed cost times your line rate
- how often does the pattern change?
- compile at config load, match per line
- one compiled value serves every worker
basics
~20 sPatterns are being compiled inside the per-line path instead of once. Find the call — usually regexp.MatchString in the loop or rules compiled per event — compile every pattern once when the ruleset loads, and reuse the shared *regexp.Regexp values for every line.
solid answer
~50 sA profile weighted toward `regexp/syntax` parsing rather than the matching engine says the same thing every time: the pattern is being built per item, not per program. In an ingest loop the two usual causes are calling `regexp.MatchString(pattern, line)` inside the loop, and compiling config-supplied rules inside the per-line function because they are not source literals. Either way the fix is to move compilation up a lifetime: compile the whole ruleset once when config loads, with `regexp.Compile` so a bad rule is a rejected reload rather than a panic, and hold the resulting `[]*regexp.Regexp`. One compiled value serves every worker — a `*regexp.Regexp`'s matching methods are safe to use concurrently, so you need neither per-goroutine copies nor a lock around matching. For patterns derived from data, keep a bounded cache keyed by pattern text. Then re-profile and benchmark against captured lines to prove it.
code
go · 26 linestype ruleSet struct {
rules []*regexp.Regexp
}
// Called once per config load, never per log line.
func newRuleSet(patterns []string) (*ruleSet, error) {
rs := &ruleSet{rules: make([]*regexp.Regexp, 0, len(patterns))}
for _, p := range patterns {
re, err := regexp.Compile(p)
if err != nil {
return nil, fmt.Errorf("bad rule %q: %w", p, err)
}
rs.rules = append(rs.rules, re)
}
return rs, nil
}
// Shared by every worker; matching methods need no lock.
func (rs *ruleSet) matches(line string) bool {
for _, re := range rs.rules {
if re.MatchString(line) {
return true
}
}
return false
}go deeper
Take away the rule of thumb: compile a pattern once and reuse the compiled value; never put pattern text inside a function that runs for every item.
Explain why the profile points at compilation — parsing is a fixed cost per call, so it scales with throughput, not input size — and where the compile should move to.
Demonstrate the full loop: read the profile, find both call shapes, move compilation to config load with error reporting instead of a panic, share one compiled set across workers, then re-profile and benchmark to prove it.
Own the convention that stops recurrence — pattern text never inside a per-item function — and the reload posture: a bad operator rule rejects the reload and keeps the previous rules serving, rather than taking the ingest process down.
## Reading the profile first The signal is specific: in a CPU profile taken under load (`go tool pprof` against a profile from `runtime/pprof`, `net/http/pprof`, or `go test -cpuprofile`), the heavy frames sit in `regexp/syntax` — the parser — and in the compile path, not in the code that actually walks the input. Matching a short line is cheap; parsing a pattern into a program is not. When parsing outweighs matching, the program is building matchers as often as it uses them. The operational shape fits too. Per-item compilation is a fixed cost multiplied by throughput, so the service looks fine in staging and cliffs at some real-world line rate: latency climbs, the tailer falls behind, and the upstream buffer starts to fill. It is a throughput regression that looks like a mysterious latency regression from the outside, which is why the profile matters more than the dashboard. ## The two causes to look for **1. A convenience call in the hot loop.** `regexp.MatchString(pattern, line)` compiles the pattern, matches, and discards the compiled value — every call. In a loop over log lines that is one compile per line. The tell in review is a pattern literal appearing *inside* a function that runs per item. **2. Runtime rules compiled at the wrong level.** When the patterns come from config rather than source, people often compile inside the function that processes an event, because that is where the rules are read from. Same result: the compile happens per event rather than per config load. A third, subtler variant: a pattern assembled from data (a term, a hostname, a tenant id) and recompiled for every line it is matched against, when it changes only per request. ## The fix: raise the lifetime of the compiled value Compilation belongs at the widest scope where the pattern is known: - **Known in source** → a package-level `var re = regexp.MustCompile(...)`, compiled once during package initialisation, with a malformed literal panicking at start-up where it belongs. - **Known at config load** → compile the whole ruleset in the loader with `regexp.Compile`, wrapping each failure with the offending pattern so the operator can fix it. A bad rule rejects the reload; it does not panic a running ingest process. - **Derived from data** → compile once per distinct pattern and keep a bounded cache keyed by the pattern text, so an unbounded stream of distinct patterns cannot become an unbounded set of retained matchers. Once compiled, a `*regexp.Regexp` is effectively an immutable, reusable value: its matching methods are safe to call from many goroutines at once, so a single compiled ruleset can be shared by every worker with no copying and no mutex on the match path. (Set up any configuration on the value before you publish it to workers, not after.) ## Hot reload without reintroducing the problem The reason people compile per line is often "the rules can change". They can — just not per line. Build the *new* compiled ruleset off the hot path, then publish the whole set as a single value the workers pick up, so each worker is always matching against a coherent old or new set. Compiling inside the match loop to stay current is trading a large constant cost for a freshness guarantee measured in microseconds that nobody asked for. ## Cutting the remaining work After compilation is hoisted, the match cost is what is left, and it is worth a second look in an ingest path: - **Prefilter cheaply.** If a rule can only fire on lines containing a fixed word, a `strings.Contains` check before the regexp skips most lines for far less work. - **Order the rules** so the most selective or most common run first, and stop at the first match if the semantics allow. - **Know the literal prefix.** `(*regexp.Regexp).LiteralPrefix` reports a literal string every match must start with, which can tell you whether a cheap prefix test is available. ## Proving it Do not ship on reasoning alone. Capture a representative sample of real lines, write a benchmark over that sample for the before and after shapes, and run it with `-benchmem` — you should see both the time and the allocations attributable to the discarded compiled values disappear. Then re-take the CPU profile under the same production load and confirm the `regexp/syntax` frames are gone and the remaining time is in matching. The upstream buffer depth and the ingest lag are the operational confirmation. ## Keeping it from coming back This defect is easy to reintroduce, because the wrong version is shorter to write. The durable guard is a convention the team can apply in review without profiling: **pattern text never appears inside a function that runs per item.** Either it is a package-level compiled var, or it is a field on a value built at start-up or config load. If someone needs a matcher in a hot path, they take it from one of those, they do not build one.
- The ruleset is hot-reloaded while workers are matching. How do you refresh it without compiling per line?Build the new compiled ruleset off the hot path, then publish it as a single value the workers read — an atomic pointer swap or a fresh value handed over a channel. Each worker then matches against a coherent old or new set. Compiling in the loop to stay current buys microseconds of freshness for a very large constant cost.
- Patterns are derived per tenant from data, so there is no fixed set. Where does compilation go?Compile once per distinct pattern and keep the result in a bounded cache keyed by the pattern text, so repeated tenants reuse a matcher and a flood of distinct patterns cannot retain an unbounded number. Compile with `regexp.Compile` and treat a failure as bad tenant data, not a crash.
- After hoisting the compiles, matching itself is still the top cost. What next?Reduce how often the engine runs: prefilter with a cheap `strings.Contains` on a required literal, order rules so the most selective run first, and stop at the first match where the semantics allow. `(*regexp.Regexp).LiteralPrefix` tells you whether a required literal prefix exists to test against.
- How do you prove the fix worked before and after deploying it?Benchmark both shapes over a captured sample of real lines with `-benchmem`, then re-take the CPU profile under the same production load and confirm the parsing frames are gone. Operationally, ingest lag and upstream buffer depth should stop growing at the line rate that used to break it.
saying these in an interview costs you the question
- Blaming pattern complexity when the profile shows parsing
- Compiling config rules inside the per-line function
- Giving each worker goroutine its own compiled copy
- Locking a mutex around every match call
- Switching to MustCompile so reload errors disappear