skip to content

A client needs to fetch a resource whose identifier is the string "a/b c&d". What goes wrong when that value is placed in a URL path segment versus a query parameter, and how do you handle it correctly?

level: middleimportance: should knowfreq 46%

answer

  1. encoding is per component, not global
  2. path: / → %2F, space → %20
  3. query: & and = must be encoded; + means space
  4. %2F often rejected or pre-decoded by proxies
  5. encodeURIComponent, not encodeURI; decode once

basics

~20 s

Percent-encode per component: in a path segment a slash must become %2F (a raw / splits segments), space becomes %20; in a query, & and = must be encoded or they split parameters, and + may be read as a space. Many servers reject or pre-decode %2F, so identifiers containing slashes are safer in the query or as an opaque encoded id.

solid answer

~60 s

Percent-encoding is **component-specific** — the rules differ between a path segment and a query value. In a **path segment**, `/` is the segment delimiter, so it must be `%2F`; space is `%20`. The trap is that many servers, proxies, and frameworks normalise or reject encoded slashes before your code sees them (Apache's `AllowEncodedSlashes` defaults to off; some stacks decode the path before routing, so `%2F` becomes a real separator and the route no longer matches). So even correctly-encoded input can fail in the path. In a **query value**, `&` and `=` are the delimiters and must be encoded; `?` and `/` are legal raw. Also note `+` means space under form encoding, so a literal plus must be `%2B` — a classic bug with email addresses and base64 values. My handling: use the language's component-aware encoder (`encodeURIComponent` in JS), never string concatenation; prefer opaque, URL-safe ids so the question does not arise; if an id can contain slashes, put it in the query or base64url-encode it; and decode exactly once, server-side.

code

http · 5 lines
http
GET /items/a%2Fb%20c%26d HTTP/1.1
Host: api.example.com

GET /items?id=a%2Fb%20c%26d HTTP/1.1
Host: api.example.com

go deeper

for a junior

Know that special characters must be percent-encoded and that encodeURIComponent is the right tool for a single value.

for a middle

Explain that the reserved set differs by component, name the %2F and plus-sign traps, and show the correct encoding for each position.

for a senior

Add the infrastructure reality — proxies rejecting or normalising encoded slashes, double-decoding, traversal risk — and prefer URL-safe identifiers by design.

for a principal

Make it a platform rule: identifier alphabets constrained to the unreserved set, URL construction only through builders, one canonical decode layer, and gateway behaviour on encoded separators verified explicitly.

## Percent-encoding is per component RFC 3986 defines a set of **reserved** characters — `: / ? # [ ] @ ! $ & ' ( ) * + , ; =` — that carry structural meaning. Which ones are dangerous depends on *where* in the URL you are. This is the point most bugs miss: there is no single "URL encode" function that is correct everywhere. **In a path segment**, the delimiter is `/`. A literal slash inside a value must be `%2F`; otherwise it silently creates an extra segment and your route either misses or matches something wrong. `?` and `#` must also be encoded because they end the path. Space must be `%20`. `&` and `=` are harmless here. **In a query value**, the delimiters are `&` and `=` (and `#` ends the query). Those must be encoded. A raw `/` or `?` inside a query value is legal and common. Space may be `%20` or, under `application/x-www-form-urlencoded` conventions, `+`. **In a fragment**, nothing after `#` is sent to the server at all — worth remembering when someone reports a value "disappearing". ## The plus-sign trap `+` is not reserved as "space" by the URI spec; that meaning comes from HTML form encoding, which most server frameworks apply to query strings. The practical effect: `[email protected]` usually arrives as `jo [email protected]`. Any literal `+` — email plus-addressing, base64 payloads, signed tokens — must be sent as `%2B`. The same string in a path segment is normally *not* form-decoded, so `+` survives, which is why the same value behaves differently in the two positions. ## The encoded-slash trap Even perfectly encoded `%2F` in a path is unreliable in the wild: - Apache httpd rejects encoded slashes in the path by default (`AllowEncodedSlashes Off` → `404`). - Some proxies and gateways normalise the path by decoding percent-escapes *before* routing, turning `%2F` into a real separator. - Some frameworks decode the path once for routing and again in the handler, producing double-decoding bugs. So an identifier that can contain `/` — a file path, a hierarchical key, a fully-qualified name — should not be dropped raw into a path segment. Options: move it to a query parameter (`GET /files?path=a/b%20c`), encode it into a URL-safe alphabet (base64url, which uses `-` and `_` and no `/`), or model the hierarchy as real path segments if it genuinely is hierarchical. A related security concern is that decoding at the wrong layer enables path traversal: `%2e%2e%2f` decoded after your validation check but before file access is the classic bypass. Validate *after* canonical decoding, and decode exactly once. ## Doing it right in code Use component-aware encoders and never hand-build URLs by concatenation: - JavaScript: `encodeURIComponent(value)` for one segment or one query value (`encodeURI` is for whole URLs and deliberately leaves `/` and `&` alone — wrong for values). - Java: `URLEncoder.encode` implements *form* encoding (space → `+`), which is fine for query values but wrong for path segments; use a URI builder that encodes per component. - HTTP clients and frameworks generally offer a builder that takes values unencoded and places them correctly — prefer it. On the server, let the framework decode once, then validate the decoded value. Do not re-decode; do not encode again on the way into a downstream call without re-encoding for that context. ## Design escape hatch The cheapest fix is upstream: choose identifiers that are URL-safe by construction — UUIDs, ULIDs, or a restricted alphabet of `[A-Za-z0-9._~-]` (the unreserved set, which never needs encoding anywhere). Then path-versus-query placement is decided by design meaning rather than by escaping mechanics, and an entire family of intermittent, proxy-dependent bugs never occurs.

  • Why does an email address with a plus sign often arrive mangled in a query parameter but survive in a path segment?
    Query strings are conventionally parsed with application/x-www-form-urlencoded rules, where + decodes to a space; path segments are not form-decoded. The fix is to send a literal plus as %2B whenever it appears in a query value, and to use a component-aware encoder rather than string concatenation.
  • An identifier can legitimately contain slashes. What are your options for addressing it?
    Put it in a query parameter, where a raw slash is legal; encode it into a URL-safe form such as base64url before placing it in a path segment; or, if the value really is hierarchical, express it as multiple real path segments. Relying on %2F in a path is fragile because many servers and proxies reject or pre-decode it.

saying these in an interview costs you the question

  • Believing one encode function is correct for both path segments and query values
  • Assuming %2F in a path always reaches the application intact
  • Treating + in a query value as a literal plus
  • Using encodeURI (or Java's URLEncoder) for path segments
  • Validating a path value before decoding, leaving a traversal bypass

context