Which storage shape fits a transcoder's scratch space while rendering, and which fits the finished videos many clients fetch?
answer
- three rentals, three access units
- what does one operation address
- seek and patch, or fetch by key
- one machine attached against many readers
- scratch is block, published is object
basics
~20 sScratch belongs on a block volume: it attaches to one machine and behaves like a local disk, so seeks and in-place writes are cheap. Finished videos belong in an object store, fetched by key over HTTP by any number of readers.
solid answer
~50 sThe three shapes differ in what one operation addresses. A **block volume** addresses a block by offset on a device attached to a single machine; the machine puts a filesystem on it, so the transcoder can seek, append and rewrite a few kilobytes in the middle of a half-finished render. That is exactly what scratch work needs. An **object store** addresses a whole object by key over an HTTP API; there is no mount and no seek-and-write, but any number of clients in any zone can `GET` the same key concurrently, and capacity grows without anyone provisioning it. That is exactly what a published render needs. **File storage** — a shared mount several machines hold at once — is the middle option, and you reach for it only when more than one machine genuinely needs to write into one tree with filesystem semantics.
go deeper
Be able to name the three shapes and the unit each one addresses: a whole object by key, a block by offset, a file by path. Then say which one you would put a temporary working file on and which one a published artefact.
Explain the consequences that follow from the access unit — why partial writes are natural on two shapes and a rewrite on the third, and why only one of the three attaches to a single machine at a time.
Show that you have lived with a bad choice: code that assumed a cheap rename or a cheap directory listing, a mount that became the bottleneck, scratch that was written to the wrong shape. Say what you would measure before moving.
The angle is what the shape costs the organisation later. Every shape sets an interface your code is written against, so the decision is really about how much rewriting a future move costs and which teams inherit that bill.
## Three rentals, three access units Every storage product a platform sells is one of three shapes, and the shape is decided by one thing: **what a single operation addresses**. Everything else — how many machines can write, whether you can change part of a file, what an operation costs in latency — falls out of that. - **Object storage** addresses a **whole object by key** over an HTTP API. There is no device, no mount and no file handle. You store an object under a key such as `renders/clip-482.mp4` and fetch it back by that key. The namespace is flat: what looks like a folder is only a naming convention inside the key. Any number of machines, in any zone, can read the same key at the same time. - **Block storage** addresses a **fixed-size block by offset** on a volume attached to one machine. The machine formats it and mounts it, so it behaves like a local disk — `open`, seek to byte 3,000,000, write four kilobytes, flush. The filesystem, its cache and its journal all live on that one machine. - **File storage** addresses a **path in a shared tree** that several machines mount at once. Filesystem semantics survive, but the directory structure and the locking live on a service the mount talks to across the network. | | object store | block volume | shared mount | |---|---|---|---| | unit of access | whole object, by key | block, by offset | file, by path | | interface | HTTP request | attached device plus a filesystem | a mount over the network | | concurrent writers | many, on different keys; same key is last-write-wins | one machine at a time | many machines into one tree | | change part of a file | rewrite the object | yes, in place | yes, in place | | latency per operation | a network request | local-device latency | a network round trip | | capacity | grows as you store more | a size you provision and grow | provisioned or elastic, depending on the provider | | failure domain | region-wide in most designs | the one zone the volume lives in | usually zone- or region-scoped | ## Which one the transcoder actually wants A worker rendering a video does an enormous number of small, positioned writes: frames land, a container file's index is patched, a temporary file is truncated. Those are *block* operations. Doing them against a key-addressed store would mean rewriting the whole output for every change, and doing them across a network mount would turn each one into a round trip. So the scratch directory goes on a **block volume attached to that worker** — or on the machine's local disk, if losing the work on a restart is acceptable. The finished render is the opposite workload. It is written once, read many times, by clients the worker never meets, possibly for years. Nobody needs to modify byte 900,000 of it. It needs to be reachable from anywhere by name, to survive the worker being replaced, and to cost nothing to keep when no worker is running. That is an **object store**: publish it under a key, and the workers become disposable. The shared mount sits between them, and it is the one people over-reach for. It is the right answer when several machines must see one tree with filesystem semantics — a rendering farm where every worker must see the same font and preset directory, or a legacy tool that only knows how to open a path. It is the wrong answer when what you actually wanted was "somewhere both workers can put a file", because an object store does that with fewer moving parts and no provisioned size. ## How to decide, in order 1. **Does anything need to change part of the file after it is written?** If yes, you need block or file storage; an object store makes you rewrite the object. 2. **Does more than one machine need to write at the same time?** If yes, block storage is out — the volume attaches to one machine at a time. 3. **Is the reader population open-ended, or outside your network?** If yes, an object store is the natural home: it is reachable by key over HTTP and needs no mount. 4. **Is the data hot scratch with a lifetime measured in minutes?** Then local or block storage, and stop paying for durability you will not use. ## Why the choice sticks The shape leaks into your code. Code written against a mount assumes rename is instant, assumes a directory listing is cheap, assumes you can append. Code written against an object store assumes keys are immutable and that a read can be a range request. Moving a workload from one to the other is not a configuration change; it is a rewrite of every place the code touched storage. That is why this is a first-screen interview question: the shape is picked in the first week and lived with for years.
- Why is a shared mount not simply the safe default, given it does everything the other two do?It pays a network round trip for every metadata operation — every open, stat and directory lookup — so workloads with many small files fall off a cliff that a local block volume never hits. It also has a provisioned or metered cost while idle, and it gives you a tree several machines can corrupt each other's assumptions in. Reach for it only when shared filesystem semantics are the actual requirement.
- If the object store is reachable from anywhere, why keep a block volume on the worker at all?Because the render is built by thousands of small positioned writes. Against a key-addressed store each of those becomes a whole-object write and a network request; against an attached volume it is a local write into page cache. The volume is where the work happens, the object store is where the result lands.
- A worker dies mid-render. What differs about recovering scratch on a block volume versus on the machine's local disk?A block volume outlives the machine: detach it and attach it to a replacement in the same zone, and the half-finished work is still there. Local disk on an ephemeral machine usually does not survive the machine, so the render restarts from the source. The trade is durability against latency and cost.
A block volume is a filing drawer bolted to one desk; a shared mount is a filing room several desks can walk into. An object store is a coat check: you can ask for a slice of a bundle you handed in, but you can only ever hand a whole new bundle back in.
saying these in an interview costs you the question
- Calls an object store a filesystem you can mount and seek in
- Thinks several machines can mount one block volume read-write
- Believes folders exist in an object store rather than being key text
- Puts hot scratch in an object store and rewrites it per frame
- Assumes a shared mount costs the same per operation as a local disk
- Picks a shape by storage price alone, ignoring the access unit