A Django weather-forecast site runs on three servers; which built-in cache backend would you pick for its CACHES default, and why?
answer
- who must see the same entry
- network cache server versus in-process
- a table or a directory as fallbacks
- development-only choices
basics
~20 sPick a shared network backend, RedisCache or PyMemcacheCache, so all three servers and every worker read the same forecasts. DatabaseCache works when no cache server is allowed; LocMemCache, FileBasedCache and DummyCache do not share across servers.
solid answer
~40 sThe deciding question is **who has to see the same entry**. Three servers each running several workers need a store outside every process, so `django.core.cache.backends.redis.RedisCache` (Django's native Redis backend, via `redis-py`) or `django.core.cache.backends.memcached.PyMemcacheCache` is the default choice; both take a `LOCATION` pointing at the server, and a refresh job's `cache.set()` becomes visible everywhere at once. If the team cannot run a cache server, `DatabaseCache` shares entries through a table created by `createcachetable`, at the cost of database load. `FileBasedCache` shares only among processes that see one directory, `LocMemCache` only within one process, and `DummyCache` caches nothing, so those suit a single box, development or tests. Memcached rejects keys over 250 characters or containing whitespace, which the other backends only warn about.
code
python · 15 lines# settings.py (production)
CACHES = {
"default": {
"BACKEND": "django.core.cache.backends.redis.RedisCache",
"LOCATION": [
"redis://cache-primary:6379/0",
"redis://cache-replica:6379/0",
],
"TIMEOUT": 900,
"KEY_PREFIX": "forecast",
}
}
# settings.py (local development)
# CACHES = {"default": {"BACKEND": "django.core.cache.backends.dummy.DummyCache"}}go deeper
Recall the built-in backend names and that a multi-server site needs a cache all servers can reach.
Explain each backend's LOCATION and sharing scope, createcachetable for DatabaseCache, and the key-validity difference with Memcached.
Justify a production choice from traffic, operations and failure modes, and keep development settings on a backend with the same semantics.
Weigh running a cache server against leaning on the database, including who operates it and what happens to page latency when it is down.
## Start from the sharing question Every Django cache backend implements the same API, so the choice is not about features you call but about **where the entries live and who can see them**. A weather-forecast site is a good example: forecasts are recomputed every few minutes by a scheduled job, then read thousands of times by page views spread across three servers, each running several worker processes. Whatever the job writes must be visible to every worker on every server immediately. ## The built-in backends compared | Backend (`django.core.cache.backends.…`) | `LOCATION` | Shared by | Typical use | |---|---|---|---| | `redis.RedisCache` | Redis URL, or a list: first is the write server, others are read replicas | Every process that can reach the server | Production default for many teams | | `memcached.PyMemcacheCache` | `host:port` list or `unix:` socket | Every process that can reach the servers | Production, pure-memory cache | | `memcached.PyLibMCCache` | same, using the `pylibmc` binding | same | Production where `pylibmc` is preferred | | `db.DatabaseCache` | table name | Everything connected to that database | No cache server available | | `filebased.FileBasedCache` | absolute directory path | Processes that can see the directory | Single host, large values | | `locmem.LocMemCache` | a name for the store | Threads in one process | Development, tests, per-process memo | | `dummy.DummyCache` | ignored | Nobody; nothing is stored | Turning caching off without code changes | ## Why the network backends win for the forecast site - **One copy of each forecast**: the refresh job writes once and all workers read it, instead of each worker recomputing on its own miss. - **Invalidation works**: a `delete` reaches the only copy there is. - **Scaling is independent**: adding a fourth web server needs no change to the cache. - **Replicas are supported natively by `RedisCache`**: when `LOCATION` is a list, Django writes to the first URL and reads from a randomly chosen replica among the rest. Redis versus Memcached is mostly an operational choice for the team: both are memory-first network caches, and Django's backends expose the same cache API over either. Redis-specific topologies and eviction tuning belong to the Redis and system-design topics, not to Django's configuration. ## When the other backends are honest choices - **`DatabaseCache`**: needs `python manage.py createcachetable` (the test runner calls it for you when creating test databases). It adds reads and writes to the database you were trying to protect, and expired rows are culled only when writes push the table past `MAX_ENTRIES`. Acceptable for low traffic or when no cache server is allowed. - **`FileBasedCache`**: pickles each value to its own file. Use an absolute path outside `MEDIA_ROOT` and `STATIC_ROOT`; the system checks warn about both mistakes (`caches.W003` for a relative path, and a deploy check for a location inside a served directory). Unpickling a file an attacker could write is code execution. - **`LocMemCache`**: per process, so each forecast would be recomputed per worker and invalidations would be partial. - **`DummyCache`**: implements the API and stores nothing; handy for development settings. ## Keeping development honest A frequent mistake is to run production on Redis and every other environment on `LocMemCache` or `DummyCache`. Each hides a different class of bug: - **`DummyCache`** never returns a value, so code paths that depend on a hit (and bugs that only appear with stale data) never run locally. - **`LocMemCache`** accepts long or whitespace-containing keys with only a warning, and it keeps values inside one process, so both key-validity errors and cross-process staleness stay invisible. - **Staging on the production backend type** catches both: point staging at its own Redis or Memcached, with its own `LOCATION` and `KEY_PREFIX`, and let the fast local loop keep a lightweight backend. For the forecast site, a reasonable split is `RedisCache` in production and staging, `LocMemCache` for the test suite so tests are isolated and need no server, and either of the two for local development depending on whether developers run Redis. ## Portability traps between backends 1. **Key validity**: Memcached refuses keys longer than 250 characters or containing whitespace or control characters. The Memcached backends raise `InvalidCacheKey`; the other built-in backends only emit `CacheKeyWarning`, so a key that "works in dev" can fail in production. 2. **Culling options**: `MAX_ENTRIES` (default `300`) and `CULL_FREQUENCY` (default `3`) apply only to the backends that cull themselves, local-memory, file-based and database; Redis and Memcached evict by their own server policies. 3. **Serialisation**: every built-in backend that stores data pickles values (the Redis backend leaves plain integers unpickled so `incr()` stays atomic), so values must be picklable.
- The team is not allowed to run Redis or Memcached. What do you configure, and what must you run first?Use `django.core.cache.backends.db.DatabaseCache` with `LOCATION` set to a table name, then run `python manage.py createcachetable`, which creates any missing cache tables and leaves existing ones alone. The entries are then shared by every process that uses that database, at the cost of extra queries against it.
- How does RedisCache treat a LOCATION that lists three URLs?It treats the first URL as the server that receives writes and picks one of the remaining URLs at random for reads. Keys and the API are unchanged; only the connection used per operation differs, so the replicas must be kept in sync by Redis replication, not by Django.
- Why can a cache key that works with LocMemCache fail after switching to PyMemcacheCache?Memcached forbids keys over 250 characters or with whitespace or control characters. Django's Memcached backends raise `InvalidCacheKey` for such keys, while the other built-in backends only emit a `CacheKeyWarning`. Hashing long keys, or a custom `KEY_FUNCTION`, avoids it.
saying these in an interview costs you the question
- LocMemCache is fine behind a load balancer because all servers use the same settings.
- DatabaseCache works as soon as migrate has run, with no extra command.
- FileBasedCache shares entries across servers with no shared filesystem.
- MAX_ENTRIES caps how many keys Redis or Memcached will hold.
- Only Memcached keys need to be short, so other backends accept any key silently.