How does TF Serving pick which SavedModel version to serve, and how do you roll one out?
answer
- Numeric directories under a base path
- Filesystem polling, not an API call
- Load new, then swap, then unload
- Rename atomically into place
- Labels address, they do not split
basics
~20 sVersions are numeric subdirectories under the model base path. TF Serving polls that path and by default serves the highest number, loading the new version fully before switching traffic and then unloading the old one. A model config file changes the policy and adds labels.
solid answer
~40 sLayout is the mechanism: `/models/my_model/1/saved_model.pb`, `/models/my_model/2/...`. The server polls the base path (`--file_system_poll_wait_seconds`) and, under the default policy, serves the **highest numeric** directory. A rollout is therefore a file copy. The server loads the new version in the background, and only when it is ready does it swap the servable so requests start hitting it; the old version is then unloaded. Both are briefly resident, so plan for roughly double the model's memory during the overlap. Copy into a staging path and **rename** into place — rename is atomic, so the poller never sees a half-written directory. For more control, run with `--model_config_file`. It lets you set `model_version_policy` (`latest { num_versions: 2 }`, `specific { versions: 42 }`, `all {}`) and `version_labels` such as `stable` and `canary`, which clients address as `/v1/models/my_model/labels/canary:predict`.
go deeper
Know that TF Serving looks for numbered folders under a model base path and serves the highest number, so deploying a new model is largely a matter of adding a folder.
Explain the polling interval, the load-then-swap-then-unload sequence, and why the copy must be made atomic with a rename. Know that a config file can pin or keep several versions.
Show you plan the rollout: memory for two resident versions, warmup to avoid the post-swap p99 spike, labels rather than hard-coded numbers in clients, and a rollback that repoints a label instead of deleting artifacts.
Own the release process end to end — who may publish a version, how canary traffic is shaped in front of the server, what the rollback SLA is, whether the previous version stays hot, and how model versions are correlated with the data and code that produced them.
## The directory contract TF Serving does not track models in a database. The model base path is a directory whose **children are integers**, and each integer directory is a complete SavedModel: ``` /models/my_model/ 1/saved_model.pb variables/ 2/saved_model.pb variables/ ``` Non-numeric children are ignored. Pointing `--model_base_path` at the SavedModel directory itself, rather than its parent, is the classic first-run error: the server starts, finds no numeric children, and reports no servable versions. ## Polling and the swap The server periodically re-lists the base path — `--file_system_poll_wait_seconds` controls the interval, and setting it to 0 disables polling so the set of versions is fixed at startup. When a new numeric directory appears and the policy selects it, the server: 1. loads the new version (reads the graph, restores variables, allocates buffers) while the old one keeps serving, 2. flips the servable handle so new requests route to it, 3. unloads the previous version once it is no longer selected and its in-flight requests drain. The consequence people forget is step 1 and 2 overlapping: **two versions are in memory at once**. On a large model that is the difference between a comfortable node and an OOM kill during every deploy. The other consequence is that the new version's first requests are cold — buffers not warm, no cached shapes — which shows up as a p99 spike right after each rollout. TF Serving addresses this with model warmup: a file named `tf_serving_warmup_requests` under the version's `assets.extra/` directory, holding serialized `PredictionLog` records that the server replays at load time before marking the version available. ## Publishing atomically An export writes many files. If the poller lists the directory mid-copy, it may try to load an incomplete SavedModel. The standard discipline is to write the version somewhere else on the **same filesystem** and rename it into the base path — rename is atomic, so the directory either is not there or is complete. On object storage, prefer writing to a temporary prefix and then a single copy/rename that the server observes as one step. ## The version policy Without a config file, the policy is "serve the latest one version". A config file makes the choice explicit: - `latest { num_versions: 2 }` — keep the two highest versions loaded. This is how you keep the previous version hot for instant rollback, at the cost of holding both in memory permanently. - `specific { versions: 42 }` — pin exactly. Nothing new is picked up automatically; deploys become deliberate config changes rather than file drops. - `all {}` — serve every version present. Rarely wise for large models. The file is re-read on an interval set by `--model_config_file_poll_wait_seconds`, so a policy change can be applied without a restart. The same config file is how one server hosts several models, each with its own base path and policy. ## Labels and canaries Hard-coding version numbers into clients is brittle. `version_labels` maps a name to a number: ``` version_labels { key: "stable" value: 1 } version_labels { key: "canary" value: 2 } ``` Clients then call `/v1/models/my_model/labels/canary:predict`, or set the label on the gRPC `model_spec`. Promoting a canary is repointing the `stable` label at the new number — clients change nothing. By default a label may only point at a version that is already loaded, which prevents pointing traffic at something that does not exist; `--allow_version_labels_for_unavailable_models` relaxes that for the assign-then-load ordering. Note what labels do and do not do: they let a caller *address* a version. They do not split traffic by percentage — that shaping belongs to whatever routes requests in front of the server, which sends some share to the canary label and the rest to stable. ## Rollback Because selection is a function of the directory contents and the config, rollback has two shapes. With `latest { num_versions: N }`, deleting or moving the bad version out of the base path makes the server fall back to the next highest — fast, but destructive to the artifact. With pinned versions or labels, rollback is repointing `stable` (or the `specific` list) at the old number and letting the config poll pick it up — non-destructive and auditable. Teams running real traffic generally choose the second, and keep the previous version loaded so the switch costs no load time. ## What to watch Memory during overlap, load duration for the new version, p99 immediately after a swap (the warmup signal), and the model-status endpoint reporting the expected version. A deploy is not done when the file lands; it is done when `GET /v1/models/my_model` shows the intended version `AVAILABLE` and a real predict against it succeeds.
- Why does p99 latency spike right after a new version is swapped in, and what fixes it?The freshly loaded version has cold buffers and has never executed its graph, so the first requests pay one-off costs. TF Serving's model warmup fixes it: place a `tf_serving_warmup_requests` file of serialized PredictionLog records under the version's `assets.extra/` directory, and the server replays them at load time before marking the version available, so real traffic arrives at an already-exercised model.
- What happens to memory during a rollout of a large model?Both versions are resident between the moment the new one starts loading and the moment the old one is unloaded, so peak usage is roughly the sum of the two. Nodes sized to hold exactly one copy get OOM-killed on every deploy. Either size for two copies, or keep the fleet rolling so a given instance loads while others serve.
- A client asks for /versions/3:predict but only 1 and 2 are loaded. What does the server do?It returns an error rather than falling back to a nearby version — pinned requests are honoured exactly. This is why labels are preferable to numbers in clients: repointing `stable` from 2 to 3 changes the served version without any client edit, and a mistyped label fails loudly instead of silently serving the wrong model.
- How do you make deploys deliberate rather than triggered by a file appearing?Run with a `--model_config_file` using a `specific { versions: N }` policy, or disable filesystem polling by setting `--file_system_poll_wait_seconds` to 0. New exports can then be staged into the base path harmlessly, and traffic moves only when the config changes and is re-read on the config poll interval.
saying these in an interview costs you the question
- Thinking versions are registered through an API call
- Pointing model_base_path at the SavedModel itself
- Copying a version directly into the base path mid-write
- Assuming version labels split traffic by percentage
- Sizing a node for exactly one copy of the model