A shelve database on a shared volume caches a metrics scraper's results — what does the nightly run risk?
answer
- The cache file is an input, not storage
- The lower layer holds bytes, the upper pickles
- Write access equals shipping code into the job
- No locking, so concurrent writers corrupt it
- Read-only mode does not make loading safe
basics
~20 sshelve stores every value as a pickle, so reading a key unpickles it. Anyone able to write that file — any workload sharing the volume — gets code execution inside the scraper the moment the nightly run reads its cache.
solid answer
~50 sA shelf is two layers: a `dbm`-style store that holds only bytes, and `shelve.Shelf` on top of it turning objects into bytes with `pickle`. So every read of a key is a `pickle.loads` on data from disk, and a pickle stream can import and call whatever it names. On a shared volume that makes write access to the cache file equivalent to write access to the scraper's source: another workload edits a value and owns the process, with its credentials and network access, during an unattended six-hour run. Opening the shelf read-only does not help, and inspecting the value afterwards is too late. Separately, shelves have no locking or transactions, so concurrent writers interleave or corrupt the file, and `writeback=True` widens that window by deferring writes until sync or close. Keep the cache host-local and owner-only, or store JSON or `sqlite3` rows and rebuild objects yourself.
code
python · 20 linesimport dbm
import os
import pickle
import shelve
import tempfile
class Payload:
def __reduce__(self):
return (print, ("executed while reading the cache",))
with tempfile.TemporaryDirectory() as tmp:
path = os.path.join(tmp, "scrape-cache")
with shelve.open(path) as cache:
cache["latest"] = {"samples": 6}
with dbm.open(path, "w") as raw:
raw["latest"] = pickle.dumps(Payload())
with shelve.open(path) as cache:
cache["latest"]go deeper
The takeaway to remember is that a shelf is pickles on disk, so reading one is loading untrusted data if anyone else can write the file. Prefer JSON or a database for anything that leaves your own machine.
Explain the two layers — a byte-oriented store underneath, pickling added on top — and why that makes every read a load. Know that shelves offer no locking and what writeback=True defers until sync or close.
Frame the file as a trust boundary and reason about who can write it in a real deployment: shared volumes, sidecars, backups, restores. Give an ordered remediation and say why read-only mode and post-load validation are not controls.
Set the policy that persisted object formats are code-deployment paths, and decide where the organisation accepts them at all. Own the migration to data-only storage, the interim authentication, and the storage-permission standard that keeps the class from recurring.
## What a shelf actually is `shelve.open` gives you something that behaves like a dictionary with string keys and arbitrary Python objects as values, persisted to a file. The persistence is two layers. The bottom layer is a `dbm`-style key/value store, and it only knows how to hold **bytes**. The top layer is `shelve.Shelf`, and the way it turns your objects into bytes is `pickle`: assigning a value pickles it, reading a key unpickles it. That second sentence is the whole answer. Every read of a shelf key is a `pickle.loads` on bytes that came off disk. A pickle stream is a program — it can import a module, call what it names, and push state into objects through `__setstate__` — so the file is not passive storage. It is an input, and it is executed. ## Who can write those bytes For a scraper whose cache sits on a volume shared with other workloads, the trust boundary is whoever can write that path. Another container mounting the same volume, a compromised sidecar, a job with looser permissions, a backup-restore process, a path-traversal bug elsewhere that lets an attacker place a file: any of those can replace or edit a value, and the next read gives them code execution inside the scraper process. Not "corrupted metrics" — code execution, with the scraper's own identity: its credentials for the systems it scrapes, its outbound network access, whatever secrets are mounted beside it. A long unattended nightly run is a particularly attractive place to land, because nobody is watching it and it has hours to work. The mental model to state in an interview: **write access to a shelf file is equivalent to write access to the scraper's source code.** Anyone who has it can ship code into the process. Permission it accordingly — owner-only, on storage the process does not share — or stop storing pickles there. Two things that look like mitigations and are not. Opening the shelf read-only stops *your* process from writing; it does nothing about who else can, and the values are still unpickled on read. Checking the value after loading is too late by construction, because the stream's calls happen during the load. ## The concurrency half Shelves also give you nothing for concurrent access. There is no locking and no transaction: the documentation is explicit that a shelf is not safe for simultaneous access from multiple processes, and the underlying store may not tolerate it either. Two scraper instances writing the same file on a shared volume can interleave writes, lose updates, or corrupt the database outright, and the failure surfaces hours later as an unreadable cache or a value that never took. `writeback=True` makes that worse rather than better. It keeps an in-memory cache of every entry accessed so that in-place mutation of a stored value works, and flushes only on `sync` or `close`. The write window stretches across the whole run, memory grows with the number of keys touched, and a process killed mid-run silently loses everything it thought it had stored. On a six-hour nightly job that is a long time to be holding unwritten state. The two problems compound. A corrupted or half-written file is not merely a data-integrity issue here, because the bytes that come back out are fed to an unpickler. ## What to do instead Order the fixes by how much trust they remove. 1. **Do not store pickles in a shared location.** Keep the cache local to the host or the container, owner-readable only, and treat it as process-private state. 2. **Store data, not objects.** Put JSON text under each key, or move to a `sqlite3` database with real columns, and rebuild your objects with a constructor you wrote. Now a hostile file yields a parse error rather than a call. `sqlite3` also brings the transactions and locking that a shelf never had, which fixes the concurrency half at the same time. 3. **If the pickled shape genuinely has to stay**, authenticate each value: an HMAC over the pickled bytes with a key from a secret store, verified with `hmac.compare_digest` before the value is loaded. That turns "is this cache entry safe" into "did the writer hold the key", which is at least a decidable question — and it inherits every weakness of your key handling. A restricted unpickler that allowlists the globals your values may name is a reasonable interim narrowing while a migration runs, but it is a narrowing and not a sandbox: allowed classes still run their construction path and `__setstate__`. The generalisation worth voicing is that `shelve` is one of several places where a pickle hides behind an innocent API. Objects handed between processes by `multiprocessing` are pickled and unpickled on the far side; plenty of caching helpers and "save this object" utilities are pickle underneath. The interview skill is noticing the call site, not reciting the warning.
- Does opening the shelf with flag='r' reduce the risk?Only for your own process, and that is not the risk. Read-only mode stops the scraper writing; it says nothing about who else can write the file, and the values are still unpickled on every read. It is a correctness and concurrency measure, not a security control.
- Two scraper instances write the same shelf concurrently. What happens?There is no locking and no transaction: a shelf is not safe for simultaneous access from multiple processes, so writes interleave, updates are lost and the underlying store can be corrupted outright. writeback=True makes it worse by holding mutations in memory until sync or close. If concurrent writers are real, use sqlite3 with proper transactions instead.
- What would you store instead, without giving up the convenience?Store data rather than objects: JSON text under each key, or a sqlite3 table with real columns, and rebuild your objects with a constructor you wrote and can validate. A hostile file then produces a parse error instead of a call, and sqlite3 also supplies the transactions the shelf never had. Where the pickled shape must stay, HMAC each value with a key from a secret store and verify before loading.
A shelf looks like a filing cabinet, but each folder contains instructions the clerk carries out on opening it. Leaving the cabinet in a shared corridor is not a filing decision, it is a decision about who may give your clerk orders.
saying these in an interview costs you the question
- Thinks shelve is just a dict saved to disk
- Assumes only the writing process can affect the file
- Believes opening read-only makes unpickling safe
- Expects shelve to lock across concurrent writers
- Treats a local cache file as inherently trusted input
- Suggests validating the value after it is loaded