skip to content

Why does a SavedModel exported without an input_signature reject a different batch size?

level: seniorimportance: should knowfreq 46%

answer

  1. Concrete functions are single specializations
  2. Specs are copied from the trace
  3. None marks a free dimension
  4. dtype gets pinned too
  5. Read the -1 in saved_model_cli output

basics

~20 s

The signature written into a SavedModel is a concrete function, and a concrete function fixes the exact dtype and shape it was traced with. Trace with tf.TensorSpec([None, ...]) so the exported signature accepts any batch size.

solid answer

~40 s

A `SignatureDef` records fixed `TensorSpec`s. If the function you exported was traced by calling it with a real batch of 32 rows, the spec written to disk is `[32, 10]`, and the server rejects a request carrying 1 or 64 rows with a shape-mismatch error — at request time, from a model that loaded fine. The fix is to control the trace instead of letting a sample input control it: declare `@tf.function(input_signature=[tf.TensorSpec([None, 10], tf.float32)])`, or call `get_concrete_function(tf.TensorSpec([None, 10], tf.float32))` and register that. `None` marks a dimension as unknown, so the exported signature accepts any batch size. The same trap applies to dtype: a signature traced on `float64` rejects `float32` input. Verify before shipping with `saved_model_cli show --dir m/1 --all`, which prints the shape as `(-1, 10)` when the batch dimension is genuinely free.

go deeper

for a junior

Know that the shapes a served model accepts are decided when it is exported, and that a None batch dimension is what lets it take any number of rows.

for a middle

Explain that a signature stores a concrete function's fixed TensorSpecs, and show the two spellings that free the batch dimension: input_signature on the tf.function, or get_concrete_function with a TensorSpec.

for a senior

Demonstrate that you catch this before deploy — assert on the exported shape in the build, run the artifact locally at an unusual batch size, and connect a pinned batch dimension to server-side batching being unusable.

for a principal

Own the export contract as a build-time gate: which dimensions are declared free, what dtypes clients may send, and an automated check that fails the pipeline rather than relying on anyone reading saved_model_cli output.

## Where the pinning comes from A SavedModel signature is not a loosely typed function. It is a `SignatureDef` holding, for each input, a name, a dtype, and a shape — and those values are copied from the concrete function you registered. A concrete function is a single specialization: one dtype and one shape per argument. So the question of what your served model accepts is decided entirely at export time, by whatever specs the concrete function carries. If you produced that concrete function by calling your `tf.function` on a real sample — `fn(sample_batch)` — the specs come from that sample, batch dimension included. ## What the server sees TF Serving loads the model, publishes the signature, and validates every incoming request against those specs. A request whose batch dimension differs produces an error response mentioning the expected and actual shape. Two things make this painful in practice: - **It loads cleanly.** Nothing warns you at deploy time; the failure is per-request. - **It often passes your smoke test.** If your smoke test sends the same batch size you happened to trace with, it succeeds, and the first real traffic with a different batch fails. A fixed batch size also destroys server-side batching: a batching-enabled server assembles requests into batches of varying size, and a model that only accepts one exact size cannot participate. ## Fixing it with TensorSpec The explicit spelling is the fix: `@tf.function(input_signature=[tf.TensorSpec(shape=[None, 10], dtype=tf.float32, name="features")])` Now the `tf.function` has exactly one form, it can be handed to `tf.saved_model.save(..., signatures={"serving_default": fn})` directly, and the exported spec has an unknown leading dimension. Equivalent: `fn.get_concrete_function(tf.TensorSpec([None, 10], tf.float32))`. Use `None` for every dimension that genuinely varies at inference — batch always, and sequence length or image height/width when your graph supports them. Do **not** blanket-`None` everything: a fully unknown shape prevents shape inference inside the graph, can force ops to be less efficient, and hides real bugs where a dimension should have been fixed. The `name=` on the spec is worth setting; it becomes the input key clients use in a columnar request, and an unnamed spec gets a generated name. ## dtype is the other half The spec pins dtype as well. Tracing on a NumPy array of Python floats gives `float64`, while most clients and most Keras models send `float32`; the resulting signature then rejects perfectly reasonable input. JSON requests to the REST endpoint are parsed to the signature's dtype, so a signature expecting `int64` will reject `1.0` and one expecting `float32` will accept `1`. Being deliberate about dtype in the `TensorSpec` avoids a whole class of confusing client-side errors. ## Keras models A Keras 3 model exported with `model.export(path)` derives its serving signature from the model's own input spec, which normally already has an unknown batch dimension — Keras builds layers with `batch_size=None` unless you told it otherwise. You get the fixed-shape problem back when you explicitly built the model with a fixed batch, or when you export a custom `tf.function` yourself with `keras.export.ExportArchive` and pass a concrete sample instead of a spec to `add_endpoint(..., input_signature=[...])`. ## The pre-ship check One command answers the question definitively: `saved_model_cli show --dir m/1 --all` Read the `inputs:` block. A free batch dimension prints as `(-1, 10)`; a pinned one prints as `(32, 10)`. Make this check part of the export job rather than a thing someone remembers to do, and assert on it: if the first dimension is not `-1`, fail the build. It costs nothing and catches the defect before an endpoint ever sees traffic. A second, stronger check is `saved_model_cli run` with an input of a *different* batch size than the one used in training — if the artifact is right, it succeeds locally, and you have proven the contract without deploying anything.

  • Should you just use None for every dimension to be safe?
    No. Unknown dimensions block shape inference inside the graph, which can prevent optimizations and pushes shape errors from export time to run time. Mark free only what actually varies — batch nearly always, sequence length or image size when the graph truly handles them — and keep genuinely fixed dimensions fixed so a wrong-width input fails loudly at the boundary.
  • How does a pinned batch dimension interact with TF Serving's server-side batching?
    Badly. With `--enable_batching`, the server merges concurrent requests into batches whose size varies with load and the configured timeout, then splits the results. A signature that accepts only one exact batch size cannot serve those merged batches, so batching either fails or is effectively unusable. A free leading dimension is a prerequisite for turning batching on.
  • Your signature expects float32 but a client sends JSON integers and gets an error. What is happening?
    REST request values are parsed into the dtype the signature declares, and the conversion is not permissive in every direction. The durable fixes are to declare the dtype clients naturally produce, or to add a cast at the very start of the exported function so the graph normalizes input itself. Silently widening the spec to float64 usually just moves the problem.

saying these in an interview costs you the question

  • Believing a SavedModel accepts any shape at inference
  • Tracing on one real batch and shipping it
  • Assuming the server reshapes input for you
  • Marking every dimension None to be safe
  • Only smoke-testing with the batch size used at export

context