Adding Keras's TensorBoard callback doubled epoch time — which settings cause that?
answer
- hooks run inside the loop, synchronously
- cadence is the multiplier
- per-step writes hit every batch
- histograms touch every weight tensor
- the profiler is for short runs
basics
~20 sPer-step and per-weight work. update_freq='batch' writes a summary every training step, histogram_freq greater than zero walks every weight tensor each epoch, write_images renders them, and profile_batch runs the profiler. The defaults — update_freq='epoch', histogram_freq=0 — cost almost nothing.
solid answer
~50 s`keras.callbacks.TensorBoard` is cheap at its defaults and expensive when you opt into the diagnostic features. Three arguments dominate. `update_freq` accepts `"epoch"`, `"batch"`, or an integer number of batches; anything below epoch granularity puts a summary write on the per-step critical path, and on a fast model the I/O can rival the step itself. `histogram_freq=N` computes histograms of every weight and bias every N epochs, which means pulling the full parameter set out and binning it — noticeable on a large model, plus a much larger event file. `write_images=True` renders weights as images on the same cadence. `profile_batch` runs the tracing profiler over the given batches and, in Keras 3, only does anything under the TensorFlow backend. Diagnose by removing the callback and re-timing an epoch; fix by returning to `update_freq="epoch"`, `histogram_freq=0`, and profiling in a short dedicated run.
code
python · 16 linesimport keras
cheap = keras.callbacks.TensorBoard(
log_dir="logs/run1",
update_freq="epoch",
histogram_freq=0,
)
expensive = keras.callbacks.TensorBoard(
log_dir="logs/debug",
update_freq="batch",
histogram_freq=1,
write_images=True,
)
csv = keras.callbacks.CSVLogger("logs/history.csv", append=True)go deeper
Know that callbacks run inside the training loop, so logging is not free, and that writing summaries every batch is far more expensive than writing once per epoch.
Explain what each setting actually does — per-step summary writes, per-weight histogram computation, weight images, profiler tracing — and why the defaults are the cheap configuration.
Demonstrate the diagnosis: time epochs with callbacks removed and re-added one at a time, attribute the delta, and apply the same reasoning to hand-written callbacks with expensive batch hooks and unbounded accumulation.
Set the standing policy — what observability every training job gets by default, what is opt-in for an investigation, and how event-file growth and retention trade against the debugging value of per-step and per-weight data.
## Callbacks are on the critical path A callback is not an observer running elsewhere — its hooks execute synchronously inside the training loop, between steps. Anything a hook does is added directly to wall-clock training time. The TensorBoard callback is the usual culprit because it is the one people configure for maximum visibility and then forget about, but the reasoning generalises to any callback you write. The multiplier that matters is *cadence*. An epoch hook runs a few dozen or a few hundred times per run; a batch hook runs once per step, which on a large dataset is tens of thousands of times per epoch. The same body of work is negligible in one and dominant in the other. ## update_freq `update_freq` controls how often losses and metrics are written. `"epoch"` (the default) writes once per epoch — a handful of scalars, effectively free. `"batch"`, or an integer like `10`, writes every step or every N steps. Each write serialises scalars and touches the event file. On a small model where a step is a couple of milliseconds, the write can cost as much as the computation, which is exactly how an epoch doubles. Per-batch curves are genuinely useful when you are debugging a divergence in the first few hundred steps; they are waste for the rest of a run. If you want them, use a small integer rather than every batch, and turn the setting back to `"epoch"` when the investigation ends. ## histogram_freq and write_images `histogram_freq=0`, the default, means no weight histograms. Setting it to 1 makes the callback, at every epoch end, read every weight and bias tensor in the model, bin the values, and write the histograms. The cost scales with parameter count, and there is a second cost that shows up later: event files grow by orders of magnitude, which makes the TensorBoard UI slow to load and the artifacts expensive to keep. In Keras 3, computing histograms also requires materialising weight values on the host, so on an accelerator there is a transfer involved. `write_images=True` additionally renders weight tensors as images on the same schedule — more work, much more file. Both are diagnostic tools for "are my weights exploding or dead", not defaults for routine runs. `write_graph=True` is a one-off cost at the start rather than a per-epoch one, but it can still bloat the event file for a large model. ## profile_batch `profile_batch` runs the tracing profiler over a batch or a range of batches. Profiling is intentionally heavy — that is the point — and it is meant for a short run of a few hundred steps whose only purpose is to produce a trace. Leaving it enabled on a full training job pays the tracing cost and produces trace data nobody reads. In Keras 3 this argument is only meaningful under the TensorFlow backend; with the JAX or PyTorch backend there is no TensorFlow profiler to invoke, which is worth remembering when a snippet copied from a Keras 2 tutorial appears to do nothing. ## Diagnosing it The method is dull and reliable. Time an epoch with the callback list empty, then with the callbacks one at a time. Because callbacks are independent and synchronous, the difference attributes cleanly. If the gap appears only with `update_freq` below epoch level, it is I/O per step; if it appears at epoch boundaries only, it is histograms. The same test catches hand-written callbacks with expensive `on_train_batch_end` bodies — computing a metric over a full batch, appending to a list that grows without bound, writing a row to disk per step, or pulling weights to host memory. The fix is the same shape: accumulate something cheap per batch, and do the expensive part once per epoch. ## What logging should actually cost A sensible default set is `TensorBoard(log_dir=..., update_freq="epoch", histogram_freq=0)` plus `CSVLogger(filename, append=True)`, which writes one row per epoch and is essentially free — and, unlike a TensorBoard event file, is trivially readable when the run is on a machine you cannot point a browser at. Turn on histograms deliberately, for a specific question, and turn them off again. If per-step visibility is a standing requirement rather than a debugging trip, the honest fix is to reduce the *volume* per write and the frequency, not to accept a permanent tax on every epoch.
- Why is heavy work in on_train_batch_end worse than the same work in on_epoch_end?Purely because of how often it runs. The batch hook fires once per training step — tens of thousands of times per epoch on a real dataset — and executes synchronously between steps, so its cost multiplies by the step count and lands directly in wall-clock time. The epoch hook runs a few hundred times per run. Accumulate cheaply per batch, flush once per epoch.
- How do you attribute a slowdown to a specific callback?Time one epoch with an empty callbacks list, then add them back one at a time and re-time. Callbacks execute synchronously and independently, so the deltas attribute cleanly with no profiler needed. If the gap only appears with sub-epoch update frequency it is per-step I/O; if it shows up at epoch boundaries only, suspect histogram or image writing.
- Does profile_batch work under every Keras 3 backend?No. Keras 3 runs on TensorFlow, JAX or PyTorch, but the TensorBoard callback's profiling path is the TensorFlow profiler, so profile_batch is only meaningful under the TensorFlow backend. Under the others it does not produce a trace. This is a good example of why a Keras 2 snippet pasted into a Keras 3 project can appear to run fine while quietly doing nothing.
- What is a reasonable default logging setup that costs almost nothing?TensorBoard with update_freq='epoch' and histogram_freq=0, plus CSVLogger with append=True. That gives one scalar write and one CSV row per epoch — unmeasurable against training cost — and the CSV stays readable on a headless box where nobody can open a TensorBoard UI. Histograms, per-batch updates and profiling get switched on for a specific question and switched off again.
saying these in an interview costs you the question
- Assumes callbacks log asynchronously off the training path
- Leaves update_freq='batch' on a full production run
- Enables histogram_freq=1 by default on every model
- Leaves profiling enabled for an entire training job
- Blames the data pipeline without timing callbacks separately