In Plotly Dash, how do you stop a 60-second query in a callback from freezing the app?
answer
- the request thread is doing the waiting
- count how many workers you have
- hand the job to something else and poll
- disable the button while it runs
- and ask why it takes a minute at all
basics
~20 sMove the work off the request thread with a Dash background callback, which hands the job to a Celery or Diskcache manager and polls for the result. Add a running spec to disable the button and a cancel input, and show progress rather than a frozen page.
solid answer
~50 sA normal Dash callback runs inside the web request, so a 60-second query occupies a worker for a minute; enough concurrent users and every worker is busy and the whole app stops responding. Dash 2.x offers **background callbacks** — `@callback(..., background=True, manager=...)` — where the manager is a `CeleryManager` (a real broker and worker pool, the production choice) or a `DiskcacheManager` (local development only). The callback body runs outside the request; the browser polls until it finishes. The decorator also takes `running=[(Output("run","disabled"), True, False)]` to disable controls while it works, `cancel=[Input("cancel","n_clicks")]` to abort, and `progress=` to stream updates into a progress bar. Wrap the target component in `dcc.Loading` so the user sees a spinner. And ask the prior question first: a 60-second query is often a missing aggregate or a per-keystroke trigger that should have been a submit button.
code
python · 18 linesfrom dash import Dash, callback, Input, Output, CeleryManager
background_manager = CeleryManager(celery_app)
@callback(
Output("result-table", "data"),
Input("run-btn", "n_clicks"),
background=True,
manager=background_manager,
running=[
(Output("run-btn", "disabled"), True, False),
(Output("cancel-btn", "disabled"), False, True),
],
cancel=[Input("cancel-btn", "n_clicks")],
prevent_initial_call=True,
)
def heavy_query(n_clicks):
return run_long_query().to_dict("records")go deeper
Know that a callback runs inside the web request, so a slow one holds up the app, and that dcc.Loading shows a spinner but does not make anything faster.
Explain the worker-blocking arithmetic and the background-callback mechanism: enqueue via a manager, poll from the browser, and wire running and cancel so the UI stays honest.
Show the operational judgment: Celery versus Diskcache and why, cancellation and progress for a user who will not wait, caching in front, and attacking the query and the trigger pattern before adding a broker.
Own the boundary. Decide what interactive apps are allowed to run against the warehouse at all, which workloads belong in a scheduled pipeline producing aggregates, and what the team pays to operate a broker and worker tier for every app.
## Why a slow callback is an availability problem Every ordinary Dash callback executes synchronously inside the HTTP request that triggered it. The WSGI worker handling that request is blocked for the duration. With `gunicorn -w 4`, four concurrent 60-second callbacks consume every worker; the fifth user cannot even load the page, because serving the initial layout is also a request. A slow callback is therefore not merely a slow chart — it is a denial of service against your own app, and the failure looks like a total outage rather than one sluggish panel. ## Background callbacks Dash 2.x supports running a callback outside the request cycle. The mechanism arrived as `long_callback` and was reworked into a flag on the normal decorator: ``` from dash import Dash, callback, Input, Output, CeleryManager background_manager = CeleryManager(celery_app) @callback( Output("result", "data"), Input("run", "n_clicks"), background=True, manager=background_manager, running=[(Output("run", "disabled"), True, False)], cancel=[Input("cancel", "n_clicks")], prevent_initial_call=True, ) def heavy(n): ... ``` When the callback fires, Dash enqueues the job with the manager and returns immediately; the browser polls for completion and patches the outputs when the result lands. The request worker is free the entire time. **Manager choice matters.** `DiskcacheManager` uses the local `diskcache` library and is documented for development: it depends on a shared local filesystem, which does not hold across containers or hosts. `CeleryManager` submits to a Celery broker (typically Redis or RabbitMQ) with separate worker processes you deploy and scale independently — that is the production configuration, and it is also what makes the job survive a web-process restart. **The supporting arguments.** `running` takes a list of `(Output, value_while_running, value_when_done)` triples — disable the run button, enable the cancel button, swap a label. `cancel` takes inputs that abort the job. `progress` designates an output the function can update through an injected `set_progress` callable, driving a progress bar (dash-bootstrap-components' `dbc.Progress`, say) or a plain text counter. ## The lighter tools, and when they are enough `dcc.Loading` wraps a component and shows a spinner while any callback writing into it is in flight. It is a **UX** fix, not a concurrency fix: the worker is still blocked. Use it always, because an un-spinnered app looks broken during a two-second query — but never as the answer to a 60-second one. `dcc.Interval` fires a callback on a timer and is the right tool for polling a job you launched elsewhere, or for a live-refresh panel. Used carelessly it is a self-inflicted load generator: a one-second interval on a page with fifty viewers is fifty queries a second against the warehouse forever. Caching sits in front of everything. If the 60 seconds is spent recomputing what ten users already asked for, memoise on the callback's arguments with Flask-Caching over Redis and the second caller waits milliseconds. This is often the entire fix. ## Ask the upstream question first Before reaching for background execution, interrogate the 60 seconds: - **Is it firing too often?** A dropdown wired as `Input` recomputes per change. Move the form fields to `State` and trigger on a submit button, and the query count collapses. - **Is the warehouse doing work the model should have done?** A dashboard scanning raw fact tables per interaction wants a pre-aggregated table or a materialised view. Sub-second interaction on a modelled aggregate beats a beautifully engineered progress bar over a bad query. - **Is the payload the problem?** Returning a hundred thousand rows to a table serialises to JSON, crosses the wire, and renders in the browser. Paginate or aggregate server-side. - **Is this an app at all?** Some "slow dashboards" are batch reports wearing a dashboard costume. Materialise the answer on a schedule and serve it. ## What an interviewer is checking They want to hear the availability framing — that a synchronous callback holds a worker and that concurrency, not latency, is what breaks — plus the concrete mechanism (background callback, Celery manager in production, `running`/`cancel`/`progress`), plus the instinct to attack the query and the trigger pattern before adding infrastructure. Naming only `dcc.Loading` is the weak answer, because it makes the frozen app look prettier while it is still frozen.
- Why is DiskcacheManager not the production choice?It stores job state on the local filesystem through the diskcache library, which assumes one machine. Across containers or hosts the web process and the process that ran the job do not share that directory, so results go missing. It is documented for development. CeleryManager routes through a broker with independent workers, which survives restarts and scales separately.
- How does dcc.Loading differ from a background callback?dcc.Loading only renders a spinner over a component while a callback writing to it is in flight — the WSGI worker is still blocked for the full duration, so concurrency still collapses. A background callback actually moves the work off the request. Use dcc.Loading on top of either, because a page with no feedback reads as broken.
- When is dcc.Interval the wrong tool for a live dashboard?When the refresh rate multiplied by the concurrent viewers exceeds what the source can absorb. A one-second interval with fifty open tabs is fifty queries per second, forever, including overnight on forgotten tabs. Lengthen the interval, cache the underlying query so viewers share one result, or push refresh onto a schedule that writes an aggregate the app reads cheaply.
saying these in an interview costs you the question
- Answers only dcc.Loading, leaving the worker blocked
- Thinks adding gunicorn workers fixes a minute-long callback
- Uses DiskcacheManager in a multi-container deployment
- Polls a one-second dcc.Interval against the warehouse for every viewer
- Never questions why the query takes 60 seconds