Redis strings can be modified in place with APPEND, SETRANGE and read partially with GETRANGE. How do these commands behave — including what SETRANGE does when you write past the end of the value — and what are the limits of treating a Redis string as a growable buffer?
answer
- binary safe: STRLEN counts bytes, not characters
- APPEND returns new length, amortised growth, keeps TTL
- SETRANGE past the end zero-pads the gap
- GETRANGE inclusive, negative indexes, 0 -1 = whole value
- 512 MB hard cap; hundreds of KB is already 'big'
basics
~20 sA Redis string is a mutable byte array. APPEND adds bytes to the end and returns the new length. SETRANGE overwrites bytes at an offset, zero-filling any gap. GETRANGE returns a byte slice and accepts negative indexes. Max length is 512 MB, and rewriting huge values gets expensive.
solid answer
~60 sStrings are binary-safe byte arrays, and three commands treat them as such: - **`APPEND key bytes`** — concatenates to the end (creating the key if absent) and returns the resulting length. Redis over-allocates the underlying buffer, so repeated appends are amortised cheap rather than a copy each time. - **`SETRANGE key offset bytes`** — overwrites starting at `offset`. If the offset is beyond the current length, Redis **zero-pads the gap with null bytes** and grows the value; a single `SETRANGE k 10000000 "x"` therefore allocates ~10 MB instantly. - **`GETRANGE key start end`** — inclusive byte slice, with negative indexes counting from the end (`GETRANGE k 0 -1` returns everything). Out-of-range bounds are clamped, not errors. `STRLEN` gives the length in bytes, so multi-byte UTF-8 characters count as several. The hard ceiling is 512 MB per value. The practical limit is not the API but the cost: values are handled whole on the wire and in memory, so multi-megabyte strings inflate latency, replication traffic and fork-time memory. Fixed-width record slots and bitmap-style offsets are the good use; "append forever to a log string" is not.
code
text · 16 lines> APPEND log "first;"
(integer) 6
> APPEND log "second;"
(integer) 13
> GETRANGE log 0 4
"first"
> GETRANGE log -7 -1
"second;"
> DEL rec
> SETRANGE rec 10 "ABC" # offset beyond end
(integer) 13
> GET rec
"\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00ABC"
> STRLEN rec
(integer) 13go deeper
Know that strings are byte arrays and that APPEND adds to the end while GETRANGE slices, with -1 meaning the last byte.
Explain SETRANGE's zero-padding growth, APPEND's amortised over-allocation, and that these commands keep the key's TTL.
Argue about value size: why hundreds of KB is already large, how big values affect latency and replication, and when to move to lists, hashes or streams.
Set a policy on value-size limits and key design so no single command's cost scales with an unbounded user-controlled value.
## Strings are byte arrays, not text Redis strings are binary safe: they store arbitrary bytes, including NULs and image data, and Redis never interprets an encoding. That is why `STRLEN` counts bytes, not characters — `SET k "héllo"` under UTF-8 has length 6, and slicing with GETRANGE can cut a multi-byte character in half. If you slice user-facing text, slice on boundaries you control, or store fixed-width ASCII fields. ## APPEND `APPEND key value` concatenates to the end of the existing value and returns the new total length. If the key does not exist it behaves like a SET, creating it. If the key holds another type you get `WRONGTYPE`. Crucially, APPEND does **not** clear the key's TTL — it modifies the value rather than recreating the key. The implementation over-allocates the backing buffer (roughly doubling up to a cap, then growing in fixed chunks), which makes a series of appends amortised O(1) in the amount appended rather than O(current length) each time. That still leaves two real costs: memory used is larger than the logical length until the value is rewritten, and every append eventually triggers a reallocation-and-copy of the whole value, which for a 100 MB string is a single very expensive command. ## SETRANGE and the zero-padding trap `SETRANGE key offset value` overwrites `len(value)` bytes starting at `offset`, leaving everything else intact, and returns the new length. The behaviour people miss: **if `offset` is greater than the current length, Redis extends the value and fills the gap with zero bytes (`\x00`)**. So on an empty key, `SETRANGE k 1048575 "A"` immediately produces a 1 MiB value. This is exactly how bitmaps grow when you set a high bit offset, and it is also how a naive "use the user id as the offset" scheme accidentally allocates hundreds of megabytes from a single command. Offsets are limited so the resulting value cannot exceed 512 MB; beyond that the command errors. The useful pattern is fixed-width records: reserve N bytes per field and update field 3 with `SETRANGE k (3*N) <newvalue>` without reading or rewriting the rest — one command, no read-modify-write race. ## GETRANGE `GETRANGE key start end` returns the inclusive slice `[start, end]`. Negative indexes count backwards: `-1` is the last byte, so `GETRANGE k 0 -1` is the whole value and `GETRANGE k -100 -1` is the last 100 bytes. Bounds outside the value are clamped and a fully out-of-range request returns an empty string rather than an error. Cost is O(length of the returned slice), which is why fetching a small window out of a large value is cheap while `GETRANGE k 0 -1` on a big value is not. ## The 512 MB ceiling and the practical ceiling The protocol-level maximum for one string value is 512 MB, and the server also caps a single request/reply bulk length (`proto-max-bulk-len`, 512 MB by default). But the number that matters operationally is much lower. Values are transferred whole in the reply, held whole in memory, copied whole on reallocation, and shipped whole to replicas in the write stream; large values also make memory fragmentation and snapshot-time copy-on-write worse. A few hundred kilobytes is already a "big value" in most production deployments, and multi-megabyte strings are a known source of latency spikes because a single command's work is proportional to the value size. ## When to use these commands, and when not to Good fits: - **Bitmap-style data**, where SETRANGE/GETRANGE are the byte-level counterparts of SETBIT/BITCOUNT. - **Fixed-width binary records** updated field-by-field without a read-modify-write cycle. - **Bounded accumulation**, such as appending small delimiter-separated tokens with a cap and a TTL. Bad fits: - **Unbounded logs or event streams built by APPEND.** There is no trimming primitive, readers must slice by offset they have to track themselves, and the value only ever grows. A list or a stream is the right structure. - **Large documents you mutate frequently.** Every write ships the changed range but every read ships the whole value; a hash with fields, or a smaller key per document part, keeps both sides small. ## Atomicity note Each of these is a single command and therefore atomic: two clients appending concurrently both have their bytes retained, in some serialized order, with no lost update. What you cannot do atomically is "read the length, decide, then write" as two commands — the length may change in between. Use APPEND's returned length (it tells you where your bytes ended up) or a Lua script.
- What does SETRANGE do if the offset is far beyond the current length of the value?It grows the value to offset + len(bytes) and fills everything between the old end and the offset with zero bytes. The allocation is immediate, so a single SETRANGE at offset 100,000,000 creates a ~100 MB value in one command. This is the same mechanism that makes SETBIT at a very high offset expensive, and it is why offsets should be dense, bounded values rather than raw ids.
- Why is building an append-only log out of one Redis string a bad idea?The value only grows — there is no trim primitive — so memory rises without bound and every full read transfers the entire log. Reallocation copies the whole buffer periodically, and replication ships large writes. Lists (with LTRIM) or streams (with XADD MAXLEN) give bounded size, range reads by id, and consumer semantics, which is what an append-only workload actually needs.
saying these in an interview costs you the question
- Thinking SETRANGE past the end errors or leaves a hole, instead of zero-padding and allocating
- Assuming STRLEN and GETRANGE work in characters, so slicing never breaks UTF-8
- Treating a Redis string as a cheap unbounded log because APPEND is amortised O(1)
- Believing GETRANGE with out-of-range indexes raises an error rather than clamping
- Ignoring that multi-megabyte values make every read, replication write and reallocation proportionally expensive