How do you safely build a MongoDB $regex filter from user-supplied search text?
answer
- User input is data, never structure
- A string is still a pattern
- Escape before you interpolate
- Anchoring is what a range needs
- Case-insensitivity removes the range
basics
~20 sTreat the input as a literal string: verify it is a string, escape every regular-expression metacharacter, anchor the pattern where possible, and cap its length. Never splice a raw client-supplied object into the filter document.
solid answer
~50 sThree risks stack up. First, **operator injection**: if a client sends `{"$regex": ".*"}` or `{"$ne": null}` where your code expected a string, splicing it straight into the filter turns a lookup into a full match. Check the value is a string before it goes anywhere near the filter. Second, **metacharacter injection**: an unescaped `.`, `*`, `(` or `|` changes the pattern's meaning, and a pathological pattern can burn server CPU, since the regex is evaluated on the server. Escape the input, or build the pattern from a whitelisted shape. Third, **cost**: `$regex` with `$options: "i"`, or any unanchored pattern, cannot be reduced to an index range, so it degrades toward examining every candidate document. A prefix-anchored, case-sensitive pattern such as `/^abc/` is the one shape a normal index can bound. For real full-text search, a search index is the right tool rather than a regex.
code
javascript · 8 lines// unsafe: value may be an object, and is compiled as a pattern
db.users.find({ name: { $regex: req.body.q } })
// safer: assert type, cap length, escape metacharacters, anchor
const q = req.body.q
if (typeof q !== "string" || q.length > 64) throw new Error("bad input")
const esc = q.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")
db.users.find({ tenantId: t, name: { $regex: "^" + esc } }).maxTimeMS(200)go deeper
Know that $regex matches string values, that $options: "i" makes it case-insensitive, and that user input must be escaped rather than pasted into a pattern.
Explain the mechanics: why a leading wildcard or the i option removes the index range, that the regex is evaluated server-side per candidate document, and how $options relates to the native literal form.
Demonstrate the full defence in depth — type assertion at the boundary, metacharacter escaping, length caps, a selective indexed predicate alongside, and maxTimeMS — and explain the operator-injection class of bug concretely.
Own the call on whether regex search belongs in the product at all: set the boundary rules for turning client input into query structure, and decide when a search index replaces regex before the collection makes the choice for you.
## The operator surface `$regex` matches string values against a PCRE-style pattern. Two spellings exist: ``` db.users.find({ name: { $regex: "^kt", $options: "i" } }) db.users.find({ name: /^kt/i }) ``` `$options` accepts `i` (case-insensitive), `m` (multiline anchors), `x` (extended, ignore whitespace and `#` comments) and `s` (dot matches newline). `$options` goes with the `$regex` string form; when you write a native regular-expression literal, put the flags on the literal itself. A regular-expression value may also appear inside `$in`, as in `{ name: { $in: [/^kt/, /^ab/] } }`. A `$regex` predicate only matches string values. A document whose `name` is a number is never matched by any pattern, so a regex filter silently excludes drifted types. ## Risk one: query-operator injection The filter is a document, and in most drivers so is the value you interpolate. If a web handler does the moral equivalent of `{ name: req.body.name }` and the client sends a JSON object rather than a string, that object is interpreted as an operator expression. `{"$ne": null}` turns "find this user" into "find any user". `{"$regex": ""}` matches every string. On authentication paths this has been a real, exploited class of bug. The defence is type discipline at the boundary: assert that the value is a string (or a number, or an ObjectId) before constructing the filter, and reject anything else. Frameworks with schema validation on the request body give you this for free; hand-rolled handlers must do it explicitly. Wrapping the value in an explicit `{ $eq: value }` also removes the ambiguity, because `$eq` compares the operand as a value rather than interpreting it as operators. ## Risk two: metacharacter injection and cost Once you know it is a string, the string is still a *pattern*. A user searching for `a.b` gets matches for `axb`; a user searching for `(a+)+$` submits a pattern whose backtracking behaviour can consume server CPU on long inputs. The regex is evaluated by the server, on the server's cores, once per candidate document — a slow pattern over a large scan is an availability problem, not just a slow query. Mitigations, in order of preference: 1. **Do not accept a pattern.** Escape every metacharacter in the input so it can only match itself, and build the pattern yourself: `new RegExp("^" + escape(input), "i")`. 2. **Bound the input.** Cap length, and reject inputs that fail a whitelist for the field's domain (for example, a SKU is alphanumeric and dashes). 3. **Bound the query.** Apply `maxTimeMS` so a pathological request is killed rather than pinning a core, and make sure the filter also carries a selective, indexed predicate — a tenant id, a status — so the regex evaluates over a small candidate set. ## Risk three: it does not use the index the way you hope An ordinary index stores string values in collation order. A pattern anchored at the start and matched case-sensitively — `/^abc/` — corresponds to a contiguous range of that order, so the planner can bound the scan. Remove either property and it cannot: a leading `.*` or an unanchored substring search has no range, and `$options: "i"` means the ordered values no longer line up with what the pattern accepts. In those cases the predicate is evaluated against each candidate document the plan produces. Two escapes exist. A collection or index built with a case-insensitive collation makes case-insensitive equality index-friendly, though a case-insensitive *regex* still does not gain prefix bounds. And for anything that is genuinely search — substrings, ranking, typo tolerance, multiple fields — a dedicated search index is the right mechanism; regex over a growing collection is a scaling trap that works fine in staging and falls over at production volume. ## A defensible implementation ``` function prefixFilter(field, raw) { if (typeof raw !== "string") throw new Error("bad input"); if (raw.length > 64) throw new Error("too long"); const esc = raw.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"); return { [field]: { $regex: "^" + esc, $options: "i" } }; } ``` Type check, length cap, escape, anchor. Combine the result with a selective indexed predicate and a `maxTimeMS`, and the remaining exposure is bounded. ## What to say in an interview Name all three risks separately — injection of operators, injection of pattern syntax, and query cost — because candidates who name only one usually have not operated such an endpoint. Then state the boundary rule: user input is data, never structure, and never a pattern unless you deliberately decided to expose pattern syntax as a feature.
- Which single $regex shape can a normal index bound, and why only that one?A pattern anchored at the start and matched case-sensitively, such as /^abc/. An index stores string values in collation order, so an anchored prefix corresponds to one contiguous range of that order and the planner can seek to it. Any leading wildcard has no such range, and the i option means the ordered values no longer correspond to what the pattern accepts, so each candidate must be evaluated individually.
- How does query-operator injection actually work if the field is expected to hold a string?The filter is itself a document, so an interpolated value that arrives as an object is read as an operator expression rather than a literal. A client sending {"$ne": null} or {"$regex": ""} converts a point lookup into a match-anything predicate — dangerous on login and lookup-by-token paths. Assert the value's type at the request boundary, and wrap it as { $eq: value } so it can only be compared, never interpreted.
- When should a regex search be replaced by something else entirely?As soon as the requirement is substring matching, several fields, ranking, or typo tolerance over a collection that keeps growing. Regex has no index range for those shapes, so cost scales with the candidate set and the endpoint degrades quietly as data accumulates. A dedicated search index is built for that access pattern; regex should be reserved for anchored, bounded prefix lookups on a already-narrowed set.
saying these in an interview costs you the question
- Interpolates the request body value straight into the filter
- Thinks $regex always uses an index on the field
- Believes escaping is unnecessary because it is not SQL
- Assumes a regex is evaluated on the client
- Adds $options i and expects prefix bounds to survive