skip to content

What does a minimal TGI docker run command need to serve a Hub model?

level: juniorimportance: must knowfreq 70%

answer

  1. Docker flags before the image, launcher flags after
  2. the container listens on 80
  3. mount something at /data
  4. shared memory default is too small
  5. health and info tell you what loaded

basics

~20 s

Give the container --gpus all so it sees the GPU, publish its port 80, mount a host volume at /data so downloaded weights survive restarts, and pass --model-id with the Hub repo id. Sharded runs also need --shm-size 1g.

solid answer

~40 s

Text Generation Inference ships as a container whose entrypoint is `text-generation-launcher`, so the command splits in two: everything **before** the image name is Docker's, everything **after** it is a launcher flag. A working single-GPU launch is `docker run --gpus all --shm-size 1g -p 8080:80 -v $PWD/data:/data ghcr.io/huggingface/text-generation-inference:3.3.5 --model-id HuggingFaceH4/zephyr-7b-beta`. `--gpus all` exposes the device through the NVIDIA container runtime — without it the server starts and then fails looking for CUDA. The server listens on port 80 inside the container. The `/data` mount is the Hub cache: skip it and every restart re-downloads tens of gigabytes. `--shm-size 1g` matters once you shard, because the shard processes talk over shared memory and Docker's 64 MB default is far too small. Then poll `/health` until it returns OK and `/info` to confirm which model actually loaded.

code

bash · 7 lines
bash
model=HuggingFaceH4/zephyr-7b-beta
volume=$PWD/tgi-data

docker run --gpus all --shm-size 1g -p 8080:80 \
  -v "$volume:/data" \
  ghcr.io/huggingface/text-generation-inference:3.3.5 \
  --model-id "$model" --revision main

go deeper

for a junior

Be able to type the launch command from memory: --gpus all, a published port mapped to 80, a volume at /data, then --model-id after the image name. Know that everything after the image name is a server flag.

for a middle

Explain why each piece is there — the container runtime for GPU access, the cache mount for restart cost, the raised shared memory for sharded runs — and how to confirm readiness with /health and /info rather than by sending a generation.

for a senior

Show the production shape: pinned image tag and pinned --revision, weights on a persistent or node-local volume, readiness probe on /health with a startup delay that survives weight download and warmup, and no Hub dependency in the restart path.

for a principal

Own the deployment contract around the image: which tag is approved, how model revisions are promoted, where weights are staged so a node failure does not cost a multi-gigabyte download per pod, and whether the hardware variant you standardise on (CUDA, ROCm, Gaudi) constrains which models the fleet can serve.

## What you are actually starting TGI is not a Python library you import; it is a server distributed as a container image. Inside that image the entrypoint is the `text-generation-launcher` binary, which supervises three things: a **downloader** that fetches the model weights from the Hugging Face Hub into a cache directory, one or more **shard** processes that hold the model on GPU and run the forward passes, and a **router** that owns the HTTP surface, validates requests, and batches them onto the shards. This matters for the command line. Because the entrypoint is the launcher, `docker run` arguments split at the image name: `--gpus`, `-p`, `-v`, `-e`, `--shm-size` are Docker's and go before it; `--model-id`, `--num-shard`, `--quantize`, `--max-total-tokens` are the launcher's and go after it. Putting `--model-id` before the image name is the single most common first-run mistake, and Docker's error message about an unknown flag is not obvious. ## The minimum viable launch ``` docker run --gpus all --shm-size 1g -p 8080:80 \ -v $PWD/data:/data \ ghcr.io/huggingface/text-generation-inference:3.3.5 \ --model-id HuggingFaceH4/zephyr-7b-beta ``` **`--gpus all`** hands the container GPU devices through the NVIDIA container runtime. Without it the container starts, the launcher runs, and the shard dies when it cannot find a CUDA device. On a multi-GPU box you can narrow this to `--gpus '"device=0,1"'`, or set `CUDA_VISIBLE_DEVICES` inside the container, when you want to leave other cards for other workloads. **`-p 8080:80`** publishes the port. The container serves HTTP on **80**; the host port is your choice. Every curl example that hits `http://localhost:8080/generate` assumes this mapping, and forgetting that the container side is 80 (not 8080) is a frequent source of connection-refused confusion. **`-v $PWD/data:/data`** is the weight cache. On a cold container the launcher downloads the repo — for a 7B model that is roughly 14 GB at bf16, for a 70B far more. Mounting a host directory at `/data` means the second start is seconds instead of minutes, and it means a crash-looping container is not re-pulling gigabytes each time. On Kubernetes this is the same argument for a persistent volume or a node-local scratch disk rather than the pod's ephemeral filesystem. **`--shm-size 1g`** raises the container's `/dev/shm`. Docker defaults it to 64 MB. A single-shard run usually survives that, but as soon as you shard across GPUs the shard processes coordinate through shared memory and the run hangs or dies with an opaque error. Setting it unconditionally is the cheap habit. **`--model-id`** takes a Hub repo id such as `meta-llama/Llama-3.1-8B-Instruct`. It can also take a **local path** to a directory of weights inside the container — useful when your weights come from an artifact store rather than the Hub, and mandatory in air-gapped environments. `--revision` pins a branch, tag, or commit sha so a redeploy cannot silently pick up a changed repo; leaving it unpinned means "whatever main is today", which is not what you want in production. ## Confirming it came up Startup is not instant: download, then load and shard the weights, then a warmup pass that probes how much memory the server can actually use. The launcher logs the stages, ending with a line reporting the shards ready and the router connected. Two routes make this scriptable — `/health` for readiness (this is what a Kubernetes readiness probe should hit, not the generation route) and `/info`, which reports the model id, revision, dtype and the effective token limits the server settled on. Reading `/info` after a deploy is the fastest way to catch "we shipped the wrong revision" or "the limits are not what the manifest said". A first smoke test is a POST to `/generate` with `{"inputs": "...", "parameters": {"max_new_tokens": 20}}`; a 200 with a `generated_text` field means the whole chain works. ## Image variants and pinning The published image has hardware variants — the default CUDA image, plus `-rocm` for AMD GPUs and `-gaudi` for Intel Gaudi, and a `latest-trtllm` variant. Pin an explicit tag (`:3.3.5`) rather than `latest`: an inference server upgrade can change default token limits, kernel selection and supported quantization values, and you want that to be a deliberate change with a test behind it rather than something that arrives on the next node restart. ## What this launch does not yet handle A gated repo needs a token in the environment. A model too large for one card needs sharding. A model whose context you want to bound needs explicit token limits. Those are separate flags layered onto exactly this command — the shape above does not change.

  • Why put the model cache on a mounted volume rather than leaving it in the container filesystem?
    Because the weights are the expensive part of a cold start. A 7B repo is on the order of 14 GB and a large model far more; without a mount, every restart, rescheduling or crash loop re-downloads it from the Hub. A host volume (or a node-local persistent volume in Kubernetes) turns a multi-minute cold start into a load-from-disk, and removes the Hub from your restart path entirely.
  • How would you pin exactly which model bytes a deployment serves?
    Pass `--revision` with a commit sha alongside `--model-id`, so the launcher resolves a fixed commit rather than whatever `main` points at, and pin the image tag rather than `latest`. Then verify after rollout by reading `/info`, which reports the model id and revision the server actually loaded. Without both pins, a redeploy can silently change the weights or the server behaviour.
  • What should a Kubernetes readiness probe hit on a TGI pod?
    `/health`. It reports whether the shards are loaded and the router is ready, which is the thing the probe needs to gate traffic on. Probing a generation route instead burns GPU time on every probe interval and can itself queue behind real work. Give the probe a generous initial delay too, since weight download and warmup can take minutes.

Think of the image as a sealed appliance with one dial on the front: Docker's flags decide what the appliance is plugged into, and the flags after the image name are the dial settings.

saying these in an interview costs you the question

  • Putting --model-id before the image name
  • Forgetting --gpus all and blaming the image
  • Assuming the container serves on 8080 internally
  • Leaving the weight cache inside the container
  • Running :latest in production without a pinned tag

context