Is re.compile worth it given that the re module already caches compiled patterns?
answer
- The module already avoids recompiling
- There is a cache, and it is bounded
- Eviction is the interesting failure
- Only methods take a window argument
- Named constant compiled at import
basics
~20 sThe re module caches compiled patterns, so re.compile is no large speed win. It still helps: no per-call cache lookup, no eviction when a program uses many distinct patterns, and only its methods take pos and endpos.
solid answer
~50 s`re.search` and friends do not recompile the pattern on every call — `re` keeps an internal cache of compiled pattern objects keyed by pattern text, type and flags, so the hot cost of a module-level call is a dict lookup, not a compile. That means `re.compile` is not the order-of-magnitude win people expect. It still earns its place: you skip the lookup and the argument-hashing on every call, the object cannot be **evicted** — the cache is bounded (512 entries in CPython) and evicts least-recently-used entries, so a program juggling many distinct patterns can silently recompile in a loop; you get a named, testable object that makes the pattern a module-level constant rather than a literal buried in a function; and only the compiled object's methods accept `pos` and `endpos`. `re.purge()` clears the cache, which matters mostly in benchmarks.
code
python · 12 linesimport re
p1 = re.compile(r"\d+p")
p2 = re.compile(r"\d+p")
print(p1 is p2) # True: both came from the cache
re.purge()
p3 = re.compile(r"\d+p")
print(p1 is p3) # False: the cache was cleared
# pos/endpos exist only on the compiled object
print(p1.search("clip_720p_1080p", 9))go deeper
Know that re.compile turns a pattern string into a reusable pattern object with the same match, search and findall methods, and that naming it at module level is the tidy way to write regex-heavy code.
Explain the caching: the module-level functions look the compiled pattern up rather than rebuilding it, so the compile step is not repeated. Be able to say what re.compile still buys beyond speed.
Show the operational judgment: bounded cache plus dynamically generated patterns equals silent recompilation, benchmarks beat folklore, and pos/endpos avoid copying. Diagnosing a regex hot path is the level's real test.
Own the strategy: where regex rules come from configuration, decide how they are compiled, held and invalidated on reload, and make sure your own caching layer is keyed on the pattern text so a rule change cannot serve a stale compiled pattern.
## What actually happens on a module-level call `re.search(r"\d+", s)` does not compile `\d+` afresh each time. Internally the module keeps a cache mapping `(type of pattern, pattern text, flags)` to the compiled `re.Pattern` object. The first call compiles and stores; every later call with the same key finds the ready-made object and dispatches to its method. You can observe the caching directly, because `re.compile` goes through the same cache: ```python import re p1 = re.compile(r"\d+") p2 = re.compile(r"\d+") p1 is p2 # True - the same cached object re.purge() p3 = re.compile(r"\d+") p1 is p3 # False - purge() cleared the cache ``` So the common claim "always `re.compile` your patterns, otherwise Python recompiles them every call" is wrong as stated. What remains after the caching is a *lookup*: hashing the pattern string and the flags, a dict probe, and an extra function call layer. Measurable in a tight loop, invisible almost everywhere else. ## Where re.compile does pay **1. The cache is bounded and it evicts.** CPython's pattern cache holds 512 entries; when it fills, the least-recently-used entry is dropped. A program that generates patterns dynamically — one per rule, one per field, one per file — can push past that and start recompiling on every call, and nothing tells you. Consider a video-metadata extractor that builds a pattern per caption-format rule and runs a 340-case regression pack: hundreds of distinct patterns cycling through a bounded cache is exactly the shape that thrashes. A compiled object held in a variable is immune: it exists because you hold a reference, not because a cache decided to keep it. **2. It removes the lookup from the hot path.** In a loop over millions of lines the saving is real, if modest. Measure it rather than asserting it — `timeit` will tell you in seconds whether it matters for your pattern and input. **3. The `pos` and `endpos` arguments exist only on the methods.** `pattern.search(s, pos, endpos)` restricts the scan to a window **without slicing the string**, so no copy is made. The module-level functions have no such parameters. Tokenisers and incremental scanners need this. **4. It is better engineering.** A compiled pattern named at module scope — `_RESOLUTION = re.compile(r"\d+p")` — is compiled once at import, is visible to a reader looking for the module's regexes, can be tested in isolation, and has an obvious place for the comment explaining it. A pattern literal inlined in a function body is invisible until you grep. **5. It fails fast.** A malformed pattern raises `re.error` at compile time. Compiled at import, a typo is an immediate startup failure; hidden inside a rarely-taken branch, it is a production incident at 3 a.m. ## Flags and the cache key The cache key includes the flags, so the same pattern text with different flags is a different entry — as it must be. It also includes the *type*, which is why a `str` pattern and the equivalent `bytes` pattern are separate compiled objects and never collide. And the `DEBUG` flag deliberately bypasses caching, so its output is printed on every compile. ## The stale-value trap Because a compiled pattern is a cached object keyed by text, the natural instinct to build a "smart" wrapper that caches compiled patterns *itself*, keyed by something else — a rule name, a config version — is where genuine staleness bites. If configuration changes the pattern text but your key does not change with it, the wrapper keeps returning a compiled pattern for the *old* rule, and the extractor silently keeps matching yesterday's format. The module cache does not have this problem precisely because its key **is** the pattern text: a changed pattern is a different key. If you write your own layer, key it on the pattern text too, or clear it when configuration reloads. ## What to say in an interview "Compiling is cached, so `re.compile` is not primarily a performance fix. I use it because it gives me a named constant compiled at import, it cannot be evicted from a bounded cache, and it is the only way to pass `pos` and `endpos`. If someone claims a large speed-up, I ask for the benchmark — and I check whether their program has more distinct patterns than the cache holds, because that is where the real regression usually is." `re.purge()` exists to clear the cache; it is a benchmarking and memory-reclamation tool, not something a normal program calls. If you are timing regex code, purge between runs or your first measurement will be the only honest one.
- How could a program end up recompiling the same pattern on every call?By using more distinct patterns than the cache holds. The cache is bounded — 512 entries in CPython — and evicts least-recently-used entries, so a program that builds patterns dynamically, one per rule or per input, can cycle through them and miss every time. Holding compiled objects in your own structures, keyed by the pattern text, removes the dependence on the cache entirely.
- What does re.purge() do, and when would you call it?It clears the module's internal cache of compiled patterns. Normal code has no reason to call it: the cache is bounded and small. It matters when benchmarking, so a timing run does not measure a warm cache from a previous run, and occasionally when reclaiming memory after a burst of one-off dynamically-built patterns.
- What can a compiled pattern object do that the module-level functions cannot?Its `search`, `match` and `fullmatch` methods accept `pos` and `endpos`, restricting the scan to a window of the string without slicing and therefore without copying — the basis of incremental scanners and tokenisers. It also gives you a first-class object you can name, store, pass around and test, and it surfaces a malformed pattern as an error at import time.
saying these in an interview costs you the question
- Claiming module-level calls recompile the pattern every time
- Promising a large speed-up from re.compile with no benchmark
- Not knowing the pattern cache is bounded and evicts
- Thinking re.purge is something normal code should call
- Believing pos and endpos work on the module-level functions
- Assuming flags do not affect the cache key