skip to content

In an object store, why is reading part of a large object cheap while changing part of it usually rewrites the whole object?

level: middleimportance: should knowfreq 48%

answer

  1. the two directions have different contracts
  2. an object is one immutable unit
  3. no offset to write into
  4. reads may ask for a span
  5. replace the key, do not patch bytes

basics

~20 s

Reads can be served for any byte span because the store already holds the bytes and can return a slice. Writes create a new immutable object under the key, so there is no place to patch — the store replaces what the key points at.

solid answer

~50 s

The asymmetry comes from the write model. An object under a key is treated as **one immutable unit**: writing does not mutate bytes in place, it stores a new object and re-points the key at it. There is no offset to write to, so changing one span means sending the whole content again. Reading has no such constraint — the bytes already exist, so the store can serve any requested span, which is what a range request does: ask for bytes 1,048,576 through 2,097,151 and get a partial response back rather than the whole file. That is why a transcoder can probe a container's header or pull one segment of a video without downloading it, but appending a line to a log object means rewriting the log. Some stores add a narrow append or compose operation; the general contract is still replace, not edit.

code

http · 9 lines
http
GET /renders/clip-482.mp4 HTTP/1.1
Host: objects.internal.example
Range: bytes=1048576-2097151

HTTP/1.1 206 Partial Content
Content-Range: bytes 1048576-2097151/734003200
Content-Length: 1048576

<one megabyte of the object>

go deeper

for a junior

Hold on to the one-line version: you can ask for part of an object when reading, but writing replaces what the key points at. That alone explains most of what surprises people about object storage.

for a middle

Explain why the write side is built that way — whole-object writes make replication and reader isolation tractable — and name what a range read is good for beyond saving bytes.

for a senior

Demonstrate that you spot the anti-pattern in a design review: a mutable file living under one key, a growing append target, two workers publishing to the same key. Say what you would move and where.

for a principal

The angle is where mutable state is allowed to live at all. Deciding that the object store holds immutable published artefacts and that everything mutable lives in a store built for it removes a whole class of recurring incident.

## Reads slice, writes replace An object store gives you two very different contracts on the two directions of traffic, and it is worth saying them precisely. **On the read side**, the object exists and its bytes are addressable by the store internally. So the API can accept a byte span and return only that span — a **range read**. The response is a partial one, carrying the span you asked for and telling you the span and the total size. Nothing about the object changes, and no other reader is affected. **On the write side**, an object is treated as a single immutable unit. Storing under a key does not modify anything; it creates a new object and makes the key resolve to it. There is no "write at offset" because there is no mutable extent to write into. The consequence is blunt: to change ten bytes in the middle of a large object you send the whole object again. ## Why the write side is built that way This is not an oversight. The write model buys three properties the store depends on: - **Replication is simple.** A whole object is one unit to copy, verify with a checksum and place in more than one failure domain. Partial in-place updates would mean coordinating an update across every copy, which is the problem a filesystem journal exists to solve on one machine and which is far harder across a fleet. - **Readers never see a half-written object.** Because a write lands as a complete new object before the key resolves to it, a concurrent reader gets either the old content or the new — never a mixture. A store that allowed in-place patching would have to answer what a reader sees mid-patch. - **Concurrent writes to one key collapse to last-write-wins.** There is no lock and no merge; the final write is what the key resolves to. That is a real limitation, and it is why two workers must never publish to the same key. | | range read | write | |---|---|---| | what it addresses | a byte span of an existing object | the key, wholesale | | bytes transferred | only the span | the entire content | | effect on other readers | none | they see old or new, never partial | | useful for | headers, indexes, one segment, resuming | publishing a finished artefact | | the wrong use | pretending it is a seekable file handle | patching a few bytes repeatedly | ## What this means for the transcoder Range reads are genuinely useful and under-used: 1. **Probe before you pull.** A container's header sits at a known place, so a worker can read the first chunk, decide the file is the wrong codec, and never transfer the rest. 2. **Fetch one segment.** A player or a worker that needs a ten-second window asks for the byte span that covers it. 3. **Resume a transfer.** A pull that failed at eighty percent restarts from that offset rather than from zero. 4. **Parallelise a read.** Several ranges fetched at once can beat a single sequential stream, because each request gets its own connection. And the write model tells you what not to do: - **Do not keep a mutable file in an object store.** A log, a counter file, a progress marker or an index that changes every minute is rewritten in full every minute. Costs and latency both grow with the size of the thing, not with the size of the change. - **Do not model an object as scratch.** In-progress work belongs on the attached block volume, where positioned writes are cheap; publish once when the render is complete. - **Do write once and version by key** where you need history — a new key per finished artefact, rather than patching one key repeatedly. ## The nuance to state honestly Providers differ here and it is worth acknowledging in an interview rather than claiming a universal law. Some stores offer an append operation on specific object types, some offer a way to compose a new object out of existing ones without re-uploading their bytes, and most offer a way to upload a very large object in parts that the store then assembles. What none of them offer is the filesystem contract — seek to an offset and write a few bytes, with the rest of the object untouched and no new object created. When you are asked this question, the answer that lands is: **reads address a span, writes address the key**, and every surprise about object-store cost and behaviour follows from that one sentence.

  • A worker appends one line per minute to a progress file in the object store. What happens as the file grows?
    Every append rewrites the whole file, so the bytes written per minute grow with the file's size and so does the write latency. The pattern is quadratic in total work over the day. Progress state belongs somewhere that supports small mutations — a database row, or a key per entry that you list rather than one growing object.
  • Two workers write to the same key at the same time. What does a reader get?
    One of the two complete contents — whichever write the store settled on last — never a blend of both. There is no locking or merging on a key, so the other worker's output is silently lost. The fix is to give each worker its own key and choose between them afterwards, not to coordinate the writes.
  • Why can a range read make a transfer faster rather than just smaller?
    Several ranges of the same object can be fetched concurrently, each on its own connection, so you are no longer limited by one stream's throughput. It also lets a failed transfer resume from an offset instead of restarting. The trade is more requests and client-side assembly.

saying these in an interview costs you the question

  • Treats an object as a file handle you can seek and write
  • Appends to an object and expects only the new bytes to be written
  • Thinks a range read locks or alters the object
  • Expects two writers on one key to be merged rather than one lost
  • Claims no object store offers any form of append at all
  • Downloads a whole file to read a header at the front