skip to content

SavedModel and Serving

SavedModel is the deployment format: signatures define the callable surface, TF Serving exposes it over REST and gRPC with versioning, and TFLite conversion targets the edge. Interviewers ask how a signature is chosen and what versioning buys you at rollout.

on this pageshow

questions

6

How do you send a prediction request to TensorFlow Serving's REST API?

level: juniorimportance: must knowfreq 58%

answer

  1. Port 8501 for HTTP
  2. Colon verb in the path
  3. Row list versus named columns
  4. predictions versus outputs
  5. Base64 under a b64 key

basics

~10 s

POST JSON to http://host:8501/v1/models/MODEL:predict with an "instances" list of examples; the response is a JSON object with a "predictions" list in the same order. Port 8501 is REST; 8500 is gRPC.

solid answer

~40 s

TF Serving exposes each loaded model at `/v1/models/{name}`. To infer, POST to `http://host:8501/v1/models/my_model:predict` with a JSON body. There are two body shapes. **Row format** — `{"instances": [ ... ]}` — is a list with one entry per example; the response comes back as `{"predictions": [...]}`, same length, same order. **Columnar format** — `{"inputs": {"features": [[...]]}}` — maps input tensor names to already-batched values, and the response key is `"outputs"`. Row format is what most clients use; columnar is handy when inputs have different leading dimensions. The request uses `serving_default` unless the body includes `"signature_name"`. You can pin a version with `/v1/models/my_model/versions/2:predict`. `GET /v1/models/my_model` returns model status, and `/metadata` returns the signature definitions.

go deeper

for a junior

Be able to write the curl by hand: POST to /v1/models/NAME:predict on 8501 with an instances list, and read the predictions list back in the same order.

for a middle

Explain the row versus columnar formats and their matching response keys, the b64 convention for binary inputs, and how to pin a version or label in the path.

for a senior

Show that a real predict call is your readiness probe, that you know the JSON encoding cost of large tensors, and when you would move a hot path to gRPC on 8500.

for a principal

Own the client contract: which format the organization standardizes on, how input names and versions are pinned by callers, and where the boundary sits between an HTTP gateway doing business logic and the server doing tensor math.

## The endpoint shape TF Serving's HTTP surface is small and regular. With the server started as, for example, `tensorflow_model_server --rest_api_port=8501 --model_name=my_model --model_base_path=/models/my_model`, you get: - `GET /v1/models/my_model` — is the model loaded, and in what state. - `GET /v1/models/my_model/metadata` — the signature definitions, i.e. what inputs and outputs exist. - `POST /v1/models/my_model:predict` — run the `Predict` API. - `POST /v1/models/my_model:classify` and `:regress` — the two legacy typed APIs, used with the classification and regression signature types. A version or a label may be inserted: `/v1/models/my_model/versions/2:predict`, `/v1/models/my_model/labels/canary:predict`. Omit both and you get whatever version the server currently considers servable. ## instances versus inputs This is the part candidates get wrong, and the error message is not always obvious. **Row format** treats the request as a list of examples: `{"instances": [{"height": 1.0, "width": 2.0}, {"height": 3.0, "width": 4.0}]}` or, for a single unnamed input, just a list of values: `{"instances": [[1.0, 2.0], [3.0, 4.0]]}`. Every named input must appear in every row, and every input must share the same leading (batch) dimension — that is what makes rows meaningful. The response is `{"predictions": [...]}`, one entry per row. **Columnar format** skips the row abstraction: `{"inputs": {"features": [[1.0, 2.0], [3.0, 4.0]]}}` Here the value is the whole batched tensor for each named input. The response key is `"outputs"`, not `"predictions"`. Use this when inputs legitimately do not share a batch dimension, or when it is simply easier to build the batched arrays directly. Mixing the two — sending `instances` and reading `outputs`, or vice versa — is a common integration bug. ## Binary data JSON has no byte type. For a `DT_STRING` input holding, say, a JPEG, base64-encode it and wrap it in an object with the single key `b64`: `{"instances": [{"image_bytes": {"b64": "/9j/4AAQ..."}}]}` The server decodes it to raw bytes before feeding the graph. The same convention applies to binary outputs coming back. ## Reading the response A successful `:predict` returns HTTP 200 with `predictions` (row format) or `outputs` (columnar). Errors return a non-200 with `{"error": "..."}` — shape mismatches, unknown signature names, and dtype conversion failures all surface here. Because the model can load successfully and still reject every request, treat a 200 from `GET /v1/models/my_model` as necessary but not sufficient; a real `:predict` call is the actual readiness check. ## When to move to gRPC REST is the right default for getting started, for debugging with `curl`, and for clients that live outside the TensorFlow ecosystem. It is also verbose: every float becomes a decimal string, so a large image or a wide feature vector costs meaningful serialization time and bandwidth. gRPC on port 8500 sends `TensorProto`s over a persistent HTTP/2 connection, which is markedly cheaper for large payloads. ## Practical checklist - Confirm the port: 8501 REST, 8500 gRPC. Sending JSON to 8500 fails in a confusing way. - Confirm the verb form: the colon is part of the path (`:predict`), not a query parameter. - Match the batch dimension the signature declares; a signature pinned to a fixed batch rejects a one-row request. - Read `/metadata` when unsure of input names — it prints exactly the keys `instances` rows must carry. - Send one real request in your deployment smoke test, not just a status check.

  • How do you send an image to a served model over the REST API?
    If the signature takes a `DT_STRING` of encoded bytes, base64-encode the file and wrap it as `{"b64": "..."}` inside the instance. If the signature takes a float tensor, you must decode and resize client-side and send the numeric array, which is large in JSON. Pushing the decode into the exported graph — so the endpoint takes bytes — keeps requests small and the preprocessing consistent.
  • You get HTTP 200 from GET /v1/models/my_model but every :predict returns an error. What is going on?
    The status endpoint only reports that a version loaded and is available; it says nothing about whether requests match the signature. The usual causes are a shape mismatch against a pinned batch dimension, wrong input names in the rows, a dtype the JSON values cannot convert to, or a signature name that does not exist. `GET /metadata` shows the true expected inputs.
  • When does the response key come back as outputs instead of predictions?
    When the request used the columnar `"inputs"` body rather than the row-oriented `"instances"` body. The two formats are symmetric: `instances` in, `predictions` out; `inputs` in, `outputs` out. A client that builds one format and parses the other appears to work in tests against a stub and fails against the real server.

saying these in an interview costs you the question

  • Sending JSON to the gRPC port 8500
  • Using /predict with a slash instead of :predict
  • Mixing instances in with outputs out
  • Putting raw bytes in JSON without base64
  • Treating a model-status 200 as proof inference works

context

open as a page

What does a TensorFlow SavedModel directory contain, and why does TF Serving need it?

level: middleimportance: must knowfreq 72%

basics

~20 s

A SavedModel is a directory: saved_model.pb holds the serialized graph plus signature definitions, variables/ holds the weight checkpoint, assets/ holds files the graph reads, and fingerprint.pb identifies the export. TF Serving needs the graph, not just weights.

open as a page

How do you set and inspect the serving_default signature of a SavedModel?

level: middleimportance: must knowfreq 64%

basics

~20 s

Pass signatures= to tf.saved_model.save() mapping a key such as serving_default to a concrete function, or let Keras 3's model.export() create it. Inspect with saved_model_cli show --dir m/1 --all, or load the model and read .signatures.

open as a page

How does TF Serving pick which SavedModel version to serve, and how do you roll one out?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Versions are numeric subdirectories under the model base path. TF Serving polls that path and by default serves the highest number, loading the new version fully before switching traffic and then unloading the old one. A model config file changes the policy and adds labels.

open as a page

Why does a SavedModel exported without an input_signature reject a different batch size?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The signature written into a SavedModel is a concrete function, and a concrete function fixes the exact dtype and shape it was traced with. Trace with tf.TensorSpec([None, ...]) so the exported signature accepts any batch size.

open as a page

When would you serve a SavedModel with TF Serving instead of a custom Python service?

level: principalimportance: should knowfreq 38%

basics

~20 s

TF Serving is worth it when the endpoint is mostly tensor math and you want versioning, request batching, warmup and a C++ runtime for free. A custom service wins when heavy Python-side logic surrounds the model and cannot move into the graph.

open as a page