skip to content

How do you send a prediction request to TensorFlow Serving's REST API?

level: juniorimportance: must knowfreq 58%

answer

  1. Port 8501 for HTTP
  2. Colon verb in the path
  3. Row list versus named columns
  4. predictions versus outputs
  5. Base64 under a b64 key

basics

~10 s

POST JSON to http://host:8501/v1/models/MODEL:predict with an "instances" list of examples; the response is a JSON object with a "predictions" list in the same order. Port 8501 is REST; 8500 is gRPC.

solid answer

~40 s

TF Serving exposes each loaded model at `/v1/models/{name}`. To infer, POST to `http://host:8501/v1/models/my_model:predict` with a JSON body. There are two body shapes. **Row format** — `{"instances": [ ... ]}` — is a list with one entry per example; the response comes back as `{"predictions": [...]}`, same length, same order. **Columnar format** — `{"inputs": {"features": [[...]]}}` — maps input tensor names to already-batched values, and the response key is `"outputs"`. Row format is what most clients use; columnar is handy when inputs have different leading dimensions. The request uses `serving_default` unless the body includes `"signature_name"`. You can pin a version with `/v1/models/my_model/versions/2:predict`. `GET /v1/models/my_model` returns model status, and `/metadata` returns the signature definitions.

go deeper

for a junior

Be able to write the curl by hand: POST to /v1/models/NAME:predict on 8501 with an instances list, and read the predictions list back in the same order.

for a middle

Explain the row versus columnar formats and their matching response keys, the b64 convention for binary inputs, and how to pin a version or label in the path.

for a senior

Show that a real predict call is your readiness probe, that you know the JSON encoding cost of large tensors, and when you would move a hot path to gRPC on 8500.

for a principal

Own the client contract: which format the organization standardizes on, how input names and versions are pinned by callers, and where the boundary sits between an HTTP gateway doing business logic and the server doing tensor math.

## The endpoint shape TF Serving's HTTP surface is small and regular. With the server started as, for example, `tensorflow_model_server --rest_api_port=8501 --model_name=my_model --model_base_path=/models/my_model`, you get: - `GET /v1/models/my_model` — is the model loaded, and in what state. - `GET /v1/models/my_model/metadata` — the signature definitions, i.e. what inputs and outputs exist. - `POST /v1/models/my_model:predict` — run the `Predict` API. - `POST /v1/models/my_model:classify` and `:regress` — the two legacy typed APIs, used with the classification and regression signature types. A version or a label may be inserted: `/v1/models/my_model/versions/2:predict`, `/v1/models/my_model/labels/canary:predict`. Omit both and you get whatever version the server currently considers servable. ## instances versus inputs This is the part candidates get wrong, and the error message is not always obvious. **Row format** treats the request as a list of examples: `{"instances": [{"height": 1.0, "width": 2.0}, {"height": 3.0, "width": 4.0}]}` or, for a single unnamed input, just a list of values: `{"instances": [[1.0, 2.0], [3.0, 4.0]]}`. Every named input must appear in every row, and every input must share the same leading (batch) dimension — that is what makes rows meaningful. The response is `{"predictions": [...]}`, one entry per row. **Columnar format** skips the row abstraction: `{"inputs": {"features": [[1.0, 2.0], [3.0, 4.0]]}}` Here the value is the whole batched tensor for each named input. The response key is `"outputs"`, not `"predictions"`. Use this when inputs legitimately do not share a batch dimension, or when it is simply easier to build the batched arrays directly. Mixing the two — sending `instances` and reading `outputs`, or vice versa — is a common integration bug. ## Binary data JSON has no byte type. For a `DT_STRING` input holding, say, a JPEG, base64-encode it and wrap it in an object with the single key `b64`: `{"instances": [{"image_bytes": {"b64": "/9j/4AAQ..."}}]}` The server decodes it to raw bytes before feeding the graph. The same convention applies to binary outputs coming back. ## Reading the response A successful `:predict` returns HTTP 200 with `predictions` (row format) or `outputs` (columnar). Errors return a non-200 with `{"error": "..."}` — shape mismatches, unknown signature names, and dtype conversion failures all surface here. Because the model can load successfully and still reject every request, treat a 200 from `GET /v1/models/my_model` as necessary but not sufficient; a real `:predict` call is the actual readiness check. ## When to move to gRPC REST is the right default for getting started, for debugging with `curl`, and for clients that live outside the TensorFlow ecosystem. It is also verbose: every float becomes a decimal string, so a large image or a wide feature vector costs meaningful serialization time and bandwidth. gRPC on port 8500 sends `TensorProto`s over a persistent HTTP/2 connection, which is markedly cheaper for large payloads. ## Practical checklist - Confirm the port: 8501 REST, 8500 gRPC. Sending JSON to 8500 fails in a confusing way. - Confirm the verb form: the colon is part of the path (`:predict`), not a query parameter. - Match the batch dimension the signature declares; a signature pinned to a fixed batch rejects a one-row request. - Read `/metadata` when unsure of input names — it prints exactly the keys `instances` rows must carry. - Send one real request in your deployment smoke test, not just a status check.

  • How do you send an image to a served model over the REST API?
    If the signature takes a `DT_STRING` of encoded bytes, base64-encode the file and wrap it as `{"b64": "..."}` inside the instance. If the signature takes a float tensor, you must decode and resize client-side and send the numeric array, which is large in JSON. Pushing the decode into the exported graph — so the endpoint takes bytes — keeps requests small and the preprocessing consistent.
  • You get HTTP 200 from GET /v1/models/my_model but every :predict returns an error. What is going on?
    The status endpoint only reports that a version loaded and is available; it says nothing about whether requests match the signature. The usual causes are a shape mismatch against a pinned batch dimension, wrong input names in the rows, a dtype the JSON values cannot convert to, or a signature name that does not exist. `GET /metadata` shows the true expected inputs.
  • When does the response key come back as outputs instead of predictions?
    When the request used the columnar `"inputs"` body rather than the row-oriented `"instances"` body. The two formats are symmetric: `instances` in, `predictions` out; `inputs` in, `outputs` out. A client that builds one format and parses the other appears to work in tests against a stub and fails against the real server.

saying these in an interview costs you the question

  • Sending JSON to the gRPC port 8500
  • Using /predict with a slash instead of :predict
  • Mixing instances in with outputs out
  • Putting raw bytes in JSON without base64
  • Treating a model-status 200 as proof inference works

context