Why does a Plotly Dash app that caches a DataFrame in a global break with many users?
answer
- callbacks do not remember each other
- who else is inside that same process?
- now start four of those processes
- the dev server hides it perfectly
- state belongs in the browser or a real store
basics
~20 sDash callbacks are independent stateless requests. A module-level global is shared by every user of a worker process and absent from the other workers, so users see each other's filtered data or inconsistent results depending on which process answered.
solid answer
~50 sA Dash app is a Flask app: each callback invocation is an HTTP request with no memory of the last one. A module-level `df` assigned inside a callback is therefore **process state, not user state**. With one worker, user B sees whatever user A last filtered. Under gunicorn with four workers, consecutive requests from the same user hit different processes, so the value is there sometimes and stale or missing other times — the bug that looks intermittent and "works on my laptop". The fixes, in order of preference: recompute from the source and put a cache in front of it; keep small per-user state in a `dcc.Store`, which round-trips JSON through the browser and is naturally per-session; or keep large per-user state in an external store such as Redis or Flask-Caching keyed by a session id. A read-only global loaded once at startup — a lookup table, a fitted model — is fine, because nothing writes it.
code
python · 8 lines# Broken: process state masquerading as user state
df_cache = None
@callback(Output("tbl", "data"), Input("region", "value"))
def load(region):
global df_cache
df_cache = query(region)
return df_cache.to_dict("records")go deeper
Remember that a Dash callback is just a request handler with no memory of the previous one, so a variable set in one callback is not reliably there in the next.
Explain the two failure modes precisely — sharing within a worker, absence across workers — and name the three places state can legitimately live: the request, dcc.Store, an external store.
Demonstrate the diagnosis: intermittency that tracks worker count, a user seeing another's data, dev-only success. Then state the structural rule and the eviction and TTL story for whatever store you choose.
Own the deployment consequence. Statelessness is what lets these apps scale on plain workers without sticky routing; decide how the platform enforces it and where the shared cache tier lives for every app the team runs.
## The property being violated Dash's server is stateless by design. `app.layout` and the callback registry are built once at import; after that, every user interaction is an isolated POST to the update endpoint. The function runs, returns JSON, and the process forgets everything that was not written down somewhere durable. Assigning to a module-level variable inside a callback writes to the **process**, and a process is shared by every user routed to it and invisible to every other process. ``` df_cache = None # module level @callback(Output("tbl", "data"), Input("region", "value")) def load(region): global df_cache df_cache = query(region) # process state return df_cache.to_dict("records") ``` ## The two failure modes **Cross-user leakage (single worker).** User A selects EMEA; `df_cache` now holds EMEA rows. User B's callback reads `df_cache` and gets EMEA. If the data is permissioned, this is a data-leak incident, not a bug report. **Nondeterminism (multiple workers).** Production runs `gunicorn app:server -w 4`. Each worker is a separate OS process with its own copy of the module. A user's first callback writes the global in worker 2; the next callback lands on worker 3, where the global is still `None`. The app fails roughly (workers − 1)/workers of the time, in a pattern that looks random and never reproduces under a single-process dev server. Threaded workers add the same problem plus a race on the assignment. A third, quieter mode is memory: a global keyed by user id and never evicted grows until the worker is OOM-killed and silently restarted, taking every other user's in-flight state with it. ## Where the state should go **Recompute, and cache the computation, not the session.** The cleanest Dash apps hold no session state at all: each callback derives what it needs from its inputs. Make that affordable by memoising the expensive step on its arguments — Flask-Caching backed by Redis, or `functools.lru_cache` for genuinely process-local, read-only work. The cache key is the query parameters, not the user, so sharing it is correct rather than a leak. **`dcc.Store` for small per-user state.** A `dcc.Store` component in the layout holds JSON in the browser and is passed into callbacks like any other property. `storage_type="memory"` clears on refresh, `"session"` survives reload within the tab, `"local"` persists across sessions. Because it round-trips with every callback that reads it, keep it to kilobytes — filter selections, a computed summary, an id — never a full result set. **An external store for large per-user state.** Write the intermediate frame to Redis or object storage under a session-scoped key, and pass only the key through `dcc.Store`. This is the standard "expensive query shared by three callbacks" pattern: one callback computes and stores, the others read by key. Give the keys a TTL, or the store is a leak with a slower fuse. **Read-only globals are fine.** A reference table, a fitted model, a shapefile loaded once at import and never written is exactly what a global is for. Each worker gets its own copy — budget the memory for that — but there is no correctness problem, because there is nothing to race and nothing user-specific in it. ## Diagnosing it in the wild The tells are characteristic: works in development and fails in production; failure rate that changes when you change the worker count; a user reporting someone else's numbers; a bug that disappears on retry. If you suspect it, set workers to 1 and see whether the intermittency vanishes into a consistent cross-user bug — that pair of symptoms is close to conclusive. The structural fix is a rule, not a patch: **no callback assigns to anything outside its own scope.** State either rides in the request as component properties, sits in the browser in a `dcc.Store`, or lives in a real store with a key and a TTL. Applied consistently, the app scales horizontally with no session affinity, which is Dash's actual operational advantage over session-bound frameworks — and throwing it away with a global is a needless loss.
- When is a module-level global in a Dash app perfectly correct?When it is read-only and loaded once at import — a reference lookup table, a fitted model, a geometry file. Nothing writes it, so there is no race and nothing user-specific to leak. Budget memory for one copy per worker process, and reload it by restarting the app rather than mutating it from a callback.
- What are the limits of dcc.Store for sharing data between callbacks?Its contents are JSON that round-trips to the browser on every callback that reads or writes it, so a multi-megabyte payload turns into per-interaction network cost and slow serialisation. Keep it to selections, ids and small summaries. For large intermediates, write to Redis or object storage under a session key with a TTL and pass only the key through the Store.
- How would you confirm this is the bug rather than a caching or query problem?Drop the app to a single worker. If the intermittency collapses into a consistent cross-user bug — user B reliably seeing user A's data — that pair of symptoms is close to conclusive. Then grep the callbacks for assignments to names outside the function scope, including mutation of a module-level dict or DataFrame, which is the same bug without the global keyword.
It is a shared whiteboard in one room of a four-room office: everyone in that room reads what the last person wrote, and everyone in the other three rooms finds it blank.
saying these in an interview costs you the question
- Believes each user gets their own copy of module globals
- Fixes cross-user leakage by keying the global on a user id
- Assumes it works because the single-process dev server never shows it
- Puts a whole result DataFrame in dcc.Store and calls it solved
- Adds session affinity to preserve a global instead of removing it