skip to content

What does a TensorFlow SavedModel directory contain, and why does TF Serving need it?

level: middleimportance: must knowfreq 72%

answer

  1. A directory, not a file
  2. Graph plus checkpoint plus assets
  3. No Python in the server
  4. Keras 3 export, not save

basics

~20 s

A SavedModel is a directory: saved_model.pb holds the serialized graph plus signature definitions, variables/ holds the weight checkpoint, assets/ holds files the graph reads, and fingerprint.pb identifies the export. TF Serving needs the graph, not just weights.

solid answer

~50 s

`tf.saved_model.save()` writes a **directory**, not a file. Inside it: `saved_model.pb` (the serialized `MetaGraphDef`s — the computation graph, the tag set such as `serve`, and the `SignatureDef` map that says which callables are exposed and with what input/output tensors), `variables/` (a checkpoint shard plus index holding the trained values), `assets/` (extra files the graph loads at runtime, like a vocabulary), optionally `assets.extra/` (files the graph itself ignores but a serving system may read), and `fingerprint.pb` since TF 2.13. TF Serving is a C++ binary with no Python. It cannot re-run your model-building code, so a weights-only artifact is useless to it — it needs the graph and the signatures. That is why a Keras `.keras` file or an `.h5` weights file is not loadable by TF Serving: in Keras 3 you call `model.export(path)` to get a SavedModel, while `model.save("m.keras")` produces a Python-side artifact.

go deeper

for a junior

Know that a SavedModel is a folder containing saved_model.pb plus a variables/ directory, and that it is what deployment tools consume rather than a weights file.

for a middle

Be ready to name each entry and say what it is for, and to explain why the server needs the graph and signatures rather than just the checkpoint. Know that Keras 3 uses export() for this, not save().

for a senior

Expect to be asked what breaks in production: preprocessing left in Python, assets lost during copy, a non-atomic publish, or a model placed outside a numeric version directory. Show that you verify an export before shipping it.

for a principal

Own the artifact contract between training and serving: what a build produces, how it is versioned and fingerprinted, what is guaranteed to be inside the graph versus done by callers, and how you stop a team from shipping an export that silently skips its own preprocessing.

## What a SavedModel actually is A SavedModel is TensorFlow's **language-neutral serialization of a program**, not of a network's weights. It is always a directory, and every consumer of it — TF Serving, the `tf.lite` converter, TensorFlow.js converters, other TF runtimes — reads that directory rather than executing any of your Python. That distinction is the whole reason the format exists. A training script defines a model in Python; the moment you want that model to run in a process that has no Python, something has to carry the *computation* across, not just the numbers. ## What is on disk A freshly written export looks like: - **`saved_model.pb`** — the serialized protocol buffer holding one or more `MetaGraphDef`s. Each `MetaGraphDef` is tagged (the serving one carries the tag `serve`, available as the constant `tf.saved_model.SERVING`) and contains the graph itself plus a `signature_def` map. The signatures are the callable surface: names, input tensor specs, output tensor specs. - **`variables/`** — `variables.index` and one or more `variables.data-*-of-*` shards. This is a standard TF checkpoint holding the values of every tracked `tf.Variable`. - **`assets/`** — files the graph reads at load time. The classic case is a vocabulary file backing a lookup table; the graph holds a filename, and the SavedModel keeps the file next to the graph so the pair travels together. - **`assets.extra/`** — an optional directory that TensorFlow itself never reads. Higher layers use it; TF Serving's model-warmup feature, for example, reads a `tf_serving_warmup_requests` file from here. - **`fingerprint.pb`** — added in TF 2.13, a set of hashes identifying the export so tools can tell whether two SavedModels are the same content. ## Why serving needs the graph TF Serving is written in C++ and embeds the TensorFlow runtime, not the Python interpreter. When it loads a model it maps the `serve` tag set, restores the variables from `variables/`, and exposes each `SignatureDef` as a callable. Nothing in that path can call your `build_model()` function, import your custom layer module, or evaluate a Python lambda. So the practical rule is: **whatever must run at inference has to be inside the graph**. Preprocessing written as a Python function that you call before `predict()` in your notebook does not exist in the SavedModel unless you put it inside the exported `tf.function`. This is the single most common production surprise — the served model returns garbage because the normalization that lived in the training notebook never got exported. ## Keras 3: save versus export TensorFlow 2.21 ships Keras 3, and Keras 3 draws a hard line between its own format and the TF deployment format: - `model.save("m.keras")` writes the Keras v3 archive. It is for Python-side reload — it round-trips architecture, weights and config through Keras. - `model.export("m/1")` writes a **SavedModel** directory with a `serving_default` signature derived from the model's call, which is what TF Serving can load. Handing TF Serving a `.keras` or `.h5` file produces a load failure, and the fix is always to re-export rather than to rename anything. ## Verifying what you produced Two cheap checks before shipping. First, list the directory and confirm `saved_model.pb` and a non-empty `variables/` are present. Second, run `saved_model_cli show --dir m/1 --all`, which prints every tag set, every signature, and the dtype/shape of every input and output. If the signature list is empty, TF Serving will load the model and then reject every request; if the shapes are fully static, requests with a different batch size fail. ## Failure modes worth naming - **Copying a half-written directory.** The export is not atomic across files. Write elsewhere and rename into place. - **Assets lost in transit.** Zipping or copying with a tool that skips subdirectories loses `assets/`, and lookup tables then fail at load. - **Variables directory absent.** Some exports of pure-function graphs legitimately have no variables, but for a trained model an empty `variables/` means you exported an untrained or unrestored object. - **Version directory missing.** TF Serving expects the SavedModel to sit under a numeric subdirectory of the model base path, e.g. `models/my_model/1/saved_model.pb`. Pointing it at the SavedModel directory itself finds no servable versions.

  • What ends up in assets/ and why does it have to travel with the graph?
    Any file the graph opens at load time — most often a vocabulary or label file backing a `tf.lookup` table. The graph stores a filename, so the file must be shipped alongside it; the SavedModel writer copies such files into `assets/` and rewrites the graph's filename to the copied location. Lose `assets/` and the model fails at load with a missing-file error rather than at request time.
  • You have only a checkpoint from a training run. Can TF Serving serve it?
    No. A checkpoint holds variable values keyed by name, with no graph and no signatures, so a C++ server has nothing to execute. You need the original Python model-building code: rebuild the model, restore the checkpoint into it, then `tf.saved_model.save()` (or `model.export()` in Keras 3) to produce a servable directory.
  • Does a SavedModel ever contain more than one MetaGraphDef?
    Yes — the format supports several, distinguished by tag sets, historically used to separate serving from training or TPU graphs. In practice modern exports from `tf.saved_model.save()` write a single `serve`-tagged graph, and TF Serving loads exactly that tag set. When inspecting an unfamiliar model, `saved_model_cli show --dir m/1` lists the available tag sets first.

saying these in an interview costs you the question

  • Calling a SavedModel a single .pb file
  • Thinking weights alone are enough to serve
  • Assuming TF Serving can import Python custom layers
  • Believing model.save('m.keras') produces a servable artifact
  • Expecting Python preprocessing to run inside the served model

context