Should a TF 2.15 codebase set TF_USE_LEGACY_KERAS or migrate to Keras 3?
answer
- two upgrades, not one
- who owns the blocking code?
- a bridge needs a date
- artifacts outlive the code
- stop new code joining the debt
basics
~20 sDecide by blast radius. TF_USE_LEGACY_KERAS=1 with the tf_keras package unblocks the TensorFlow upgrade in a day but pins you to a maintenance-only Keras 2. Treat it as a time-boxed bridge and schedule the Keras 3 migration behind it.
solid answer
~50 sFrame it as two separate upgrades that a naive bump merges into one: the TensorFlow runtime and the Keras API. Setting `TF_USE_LEGACY_KERAS=1` (with `tf_keras` installed) decouples them — you get the newer TensorFlow runtime while Keras 2 code keeps running unchanged. That is the right first move when the change is urgent, the surface is large, or the blocker is a third-party dependency you do not control. It is a bridge, not a destination: `tf_keras` is maintenance-only, so no new Keras capability lands there, and every month on the flag makes the eventual move larger. Migrate directly instead when the Keras surface is small, when the breaking changes are mechanical (saving paths, custom-layer internals, symbolic-tensor usage), or when you want the Keras 3 features. In practice I inventory custom layers, training loops and saved artifacts first, migrate one service behind CI, then delete the flag with a dated owner rather than leaving it as permanent environment config.
go deeper
Know that the legacy flag exists and what it does, and that new projects should start on Keras 3 rather than opting into the older API by default.
Be able to name what breaks in a Keras 3 migration — saving paths and formats, custom layers that touch internals, raw TensorFlow ops on symbolic tensors — and to check which Keras a job actually loaded.
Show that you separate the runtime upgrade from the API upgrade, migrate one service behind CI with an output-equivalence check, and convert saved artifacts with the old path kept as a fallback.
Own the framing and the deadline: size the exposure before choosing, decide when the fleet standardises, stop new code joining the debt, and put an owner and a date on the bridge so an environment variable does not become permanent architecture.
## The decision is about decoupling, not about Keras Upgrading from TensorFlow 2.15 to 2.21 bundles two unrelated changes: a newer TensorFlow runtime (ops, CUDA support, security fixes, performance) and a different Keras API (Keras 3 replacing the bundled Keras 2). Teams get into trouble by treating that as one atomic step and then reverting the whole thing when Keras breaks. `TF_USE_LEGACY_KERAS=1`, with the `tf_keras` package installed, exists precisely to split them. Set it before importing TensorFlow and `tf.keras` routes back to Keras 2 while the runtime is modern. That is a genuinely useful lever, and knowing it exists is half the answer. ## When the flag is the right first move - **The blocker is somebody else's code.** A dependency that subclasses Keras 2 classes cannot be fixed by you on your schedule. - **The upgrade is urgent for a runtime reason** — a CVE, a driver or CUDA requirement, a hardware move — and the Keras migration is not on the critical path. - **The Keras surface is large and unaudited**: many custom layers, custom training loops, home-grown serialization, checkpoints written years ago by people who left. - **You want to de-risk**: ship the runtime upgrade, prove it in production, then change one variable at a time. ## When to migrate straight to Keras 3 - The Keras surface is small — a handful of Sequential/Functional models and stock layers usually port with little more than saving-path changes. - The breakages you find in the inventory are mechanical rather than architectural. - You want what only exists on the Keras 3 side: backend portability, `keras.ops`, the newer saving format, ongoing feature work. - You are starting a new service. New code should never be born on the legacy flag. ## What the inventory should look for Before choosing, spend an afternoon listing the real exposure: 1. **Custom layers and models** that reach into Keras internals rather than using the public `Layer`/`Model` API. Public-API layers usually port; internals-touching ones are the expensive class. 2. **Saving and loading paths.** Keras 3 changed defaults: `model.save()` writes the `.keras` archive, exporting a TensorFlow SavedModel directory is `model.export()`, and `save_weights` expects a `.weights.h5` filename. Anything that hard-codes a save format, or that reloads an old artifact, needs attention — and artifacts outlive code, so this is where migrations actually stall. 3. **Symbolic-tensor usage.** Code that calls raw `tf.*` ops on the output of `keras.Input` fails under Keras 3, because those are backend-agnostic `KerasTensor` placeholders. 4. **Input-pipeline helpers.** Keras 3's base class for Python-side data producers is `keras.utils.PyDataset`; older code written against the Keras 2 sequence base class needs updating. 5. **Downstream consumers.** Serving stacks, evaluation jobs and model registries that load your artifacts have to move roughly in step; a training job that unilaterally changes artifact format breaks them. That list also tells you the cost, which is what turns the decision from opinion into estimate. ## Running the bridge responsibly If you take the flag, treat it as debt with a name on it: - **Set it in one place** — the container/job environment or a single entry-point line — never scattered across notebooks and shell profiles, or you get hosts that disagree and "works on my machine" incidents. - **Log the resolved Keras version at job startup.** One line (`tf.keras.__version__`) turns a mystery into a diagnosis. - **Stop the leak.** New services start on Keras 3; the flag covers existing code only. Otherwise the migration surface grows while you are paying to shrink it. - **Date it.** An owner and a target release, tracked like any other dependency deadline. Environment variables have no natural expiry, which is exactly why they become permanent. ## Sequencing a real migration Pick the smallest service with a representative model. Migrate it behind CI with a numerical equivalence check — same inputs, same weights, compare outputs within tolerance — because silent behavioural drift, not crashes, is what makes people distrust a framework migration. Convert saved artifacts deliberately: load with the legacy stack, re-save in the new format, and keep the old artifact until the new path has run in production. Then repeat, and delete the flag when the last consumer is off it. ## What the interviewer is listening for Not a preference. They want to hear that you separate runtime from API, that you size the work before choosing, that you know the escape hatch is a maintenance-only package rather than a supported long-term configuration, and that you attach an owner and a date to the bridge. "It depends" is the correct answer only if you can say what it depends on.
- What is the risk of leaving TF_USE_LEGACY_KERAS set indefinitely?The package behind it is maintenance-only, so you accumulate a growing gap: no new Keras features, no fixes beyond maintenance, and a migration whose surface keeps expanding as new code is written against Keras 2. It also becomes invisible infrastructure — an environment variable nobody remembers setting, that one host lacks, producing a heisenbug.
- How do you verify a migrated model is behaviourally identical, not just runnable?Pin the weights and compare outputs: load the same trained weights on both stacks, run a fixed evaluation batch, and assert predictions agree within a numerical tolerance, plus the headline metric on a held-out set. Crashes are cheap to find; silent drift in a preprocessing or layer default is what erodes trust, so gate the migration on that comparison in CI.
- Which part of a Keras 2 to Keras 3 migration usually takes longest?Artifacts and their consumers, not model code. Saved-model formats and default paths changed, checkpoints predate the current team, and serving, evaluation and registry jobs all read them. Model code is edited once; artifacts must be converted, revalidated and rolled out while the old path stays alive as a fallback.
- How do you prevent the migration surface from growing while the flag is on?Make new work start on Keras 3: a repository or service policy, enforced in review and ideally in CI, that the legacy flag covers only listed existing services. Publish the inventory and the target date so the bridge reads as a deadline rather than as the default configuration.
saying these in an interview costs you the question
- Treats the legacy flag as a permanent supported configuration
- Bumps TensorFlow and Keras together with no inventory
- Assumes migration is done when the code stops crashing
- Forgets saved artifacts and their downstream consumers
- Lets new services start on the legacy flag