skip to content

Text Generation Inference (TGI)

You will learn Hugging Face's TGI server: a container you point at a Hub model id to get a sharded, continuously batched endpoint. Interviewers meet it wherever a team already lives on the Hugging Face stack, and it is the usual point of comparison against vLLM.

on this pageshow

questions

12

What does a minimal TGI docker run command need to serve a Hub model?

level: juniorimportance: must knowfreq 70%

answer

  1. Docker flags before the image, launcher flags after
  2. the container listens on 80
  3. mount something at /data
  4. shared memory default is too small
  5. health and info tell you what loaded

basics

~20 s

Give the container --gpus all so it sees the GPU, publish its port 80, mount a host volume at /data so downloaded weights survive restarts, and pass --model-id with the Hub repo id. Sharded runs also need --shm-size 1g.

solid answer

~40 s

Text Generation Inference ships as a container whose entrypoint is `text-generation-launcher`, so the command splits in two: everything **before** the image name is Docker's, everything **after** it is a launcher flag. A working single-GPU launch is `docker run --gpus all --shm-size 1g -p 8080:80 -v $PWD/data:/data ghcr.io/huggingface/text-generation-inference:3.3.5 --model-id HuggingFaceH4/zephyr-7b-beta`. `--gpus all` exposes the device through the NVIDIA container runtime — without it the server starts and then fails looking for CUDA. The server listens on port 80 inside the container. The `/data` mount is the Hub cache: skip it and every restart re-downloads tens of gigabytes. `--shm-size 1g` matters once you shard, because the shard processes talk over shared memory and Docker's 64 MB default is far too small. Then poll `/health` until it returns OK and `/info` to confirm which model actually loaded.

code

bash · 7 lines
bash
model=HuggingFaceH4/zephyr-7b-beta
volume=$PWD/tgi-data

docker run --gpus all --shm-size 1g -p 8080:80 \
  -v "$volume:/data" \
  ghcr.io/huggingface/text-generation-inference:3.3.5 \
  --model-id "$model" --revision main

go deeper

for a junior

Be able to type the launch command from memory: --gpus all, a published port mapped to 80, a volume at /data, then --model-id after the image name. Know that everything after the image name is a server flag.

for a middle

Explain why each piece is there — the container runtime for GPU access, the cache mount for restart cost, the raised shared memory for sharded runs — and how to confirm readiness with /health and /info rather than by sending a generation.

for a senior

Show the production shape: pinned image tag and pinned --revision, weights on a persistent or node-local volume, readiness probe on /health with a startup delay that survives weight download and warmup, and no Hub dependency in the restart path.

for a principal

Own the deployment contract around the image: which tag is approved, how model revisions are promoted, where weights are staged so a node failure does not cost a multi-gigabyte download per pod, and whether the hardware variant you standardise on (CUDA, ROCm, Gaudi) constrains which models the fleet can serve.

## What you are actually starting TGI is not a Python library you import; it is a server distributed as a container image. Inside that image the entrypoint is the `text-generation-launcher` binary, which supervises three things: a **downloader** that fetches the model weights from the Hugging Face Hub into a cache directory, one or more **shard** processes that hold the model on GPU and run the forward passes, and a **router** that owns the HTTP surface, validates requests, and batches them onto the shards. This matters for the command line. Because the entrypoint is the launcher, `docker run` arguments split at the image name: `--gpus`, `-p`, `-v`, `-e`, `--shm-size` are Docker's and go before it; `--model-id`, `--num-shard`, `--quantize`, `--max-total-tokens` are the launcher's and go after it. Putting `--model-id` before the image name is the single most common first-run mistake, and Docker's error message about an unknown flag is not obvious. ## The minimum viable launch ``` docker run --gpus all --shm-size 1g -p 8080:80 \ -v $PWD/data:/data \ ghcr.io/huggingface/text-generation-inference:3.3.5 \ --model-id HuggingFaceH4/zephyr-7b-beta ``` **`--gpus all`** hands the container GPU devices through the NVIDIA container runtime. Without it the container starts, the launcher runs, and the shard dies when it cannot find a CUDA device. On a multi-GPU box you can narrow this to `--gpus '"device=0,1"'`, or set `CUDA_VISIBLE_DEVICES` inside the container, when you want to leave other cards for other workloads. **`-p 8080:80`** publishes the port. The container serves HTTP on **80**; the host port is your choice. Every curl example that hits `http://localhost:8080/generate` assumes this mapping, and forgetting that the container side is 80 (not 8080) is a frequent source of connection-refused confusion. **`-v $PWD/data:/data`** is the weight cache. On a cold container the launcher downloads the repo — for a 7B model that is roughly 14 GB at bf16, for a 70B far more. Mounting a host directory at `/data` means the second start is seconds instead of minutes, and it means a crash-looping container is not re-pulling gigabytes each time. On Kubernetes this is the same argument for a persistent volume or a node-local scratch disk rather than the pod's ephemeral filesystem. **`--shm-size 1g`** raises the container's `/dev/shm`. Docker defaults it to 64 MB. A single-shard run usually survives that, but as soon as you shard across GPUs the shard processes coordinate through shared memory and the run hangs or dies with an opaque error. Setting it unconditionally is the cheap habit. **`--model-id`** takes a Hub repo id such as `meta-llama/Llama-3.1-8B-Instruct`. It can also take a **local path** to a directory of weights inside the container — useful when your weights come from an artifact store rather than the Hub, and mandatory in air-gapped environments. `--revision` pins a branch, tag, or commit sha so a redeploy cannot silently pick up a changed repo; leaving it unpinned means "whatever main is today", which is not what you want in production. ## Confirming it came up Startup is not instant: download, then load and shard the weights, then a warmup pass that probes how much memory the server can actually use. The launcher logs the stages, ending with a line reporting the shards ready and the router connected. Two routes make this scriptable — `/health` for readiness (this is what a Kubernetes readiness probe should hit, not the generation route) and `/info`, which reports the model id, revision, dtype and the effective token limits the server settled on. Reading `/info` after a deploy is the fastest way to catch "we shipped the wrong revision" or "the limits are not what the manifest said". A first smoke test is a POST to `/generate` with `{"inputs": "...", "parameters": {"max_new_tokens": 20}}`; a 200 with a `generated_text` field means the whole chain works. ## Image variants and pinning The published image has hardware variants — the default CUDA image, plus `-rocm` for AMD GPUs and `-gaudi` for Intel Gaudi, and a `latest-trtllm` variant. Pin an explicit tag (`:3.3.5`) rather than `latest`: an inference server upgrade can change default token limits, kernel selection and supported quantization values, and you want that to be a deliberate change with a test behind it rather than something that arrives on the next node restart. ## What this launch does not yet handle A gated repo needs a token in the environment. A model too large for one card needs sharding. A model whose context you want to bound needs explicit token limits. Those are separate flags layered onto exactly this command — the shape above does not change.

  • Why put the model cache on a mounted volume rather than leaving it in the container filesystem?
    Because the weights are the expensive part of a cold start. A 7B repo is on the order of 14 GB and a large model far more; without a mount, every restart, rescheduling or crash loop re-downloads it from the Hub. A host volume (or a node-local persistent volume in Kubernetes) turns a multi-minute cold start into a load-from-disk, and removes the Hub from your restart path entirely.
  • How would you pin exactly which model bytes a deployment serves?
    Pass `--revision` with a commit sha alongside `--model-id`, so the launcher resolves a fixed commit rather than whatever `main` points at, and pin the image tag rather than `latest`. Then verify after rollout by reading `/info`, which reports the model id and revision the server actually loaded. Without both pins, a redeploy can silently change the weights or the server behaviour.
  • What should a Kubernetes readiness probe hit on a TGI pod?
    `/health`. It reports whether the shards are loaded and the router is ready, which is the thing the probe needs to gate traffic on. Probing a generation route instead burns GPU time on every probe interval and can itself queue behind real work. Give the probe a generous initial delay too, since weight download and warmup can take minutes.

Think of the image as a sealed appliance with one dial on the front: Docker's flags decide what the appliance is plugged into, and the flags after the image name are the dial settings.

saying these in an interview costs you the question

  • Putting --model-id before the image name
  • Forgetting --gpus all and blaming the image
  • Assuming the container serves on 8080 internally
  • Leaving the weight cache inside the container
  • Running :latest in production without a pinned tag

context

open as a page

In TGI, what does --max-batch-prefill-tokens control, and what breaks if it is too high?

level: middleimportance: must knowfreq 58%

basics

~20 s

TGI's --max-batch-prefill-tokens caps how many prompt tokens the router packs into one prefill step. Raising it improves prompt-processing throughput but raises peak activation memory, so setting it too high produces CUDA out-of-memory crashes under concurrent long prompts.

open as a page

TGI fails to download a gated Llama repo — how do you fix the launch?

level: middleimportance: must knowfreq 55%

basics

~20 s

Accept the model's licence with the Hugging Face account that owns the token, then inject the token into the container at runtime as the HF_TOKEN environment variable. The download happens once at startup, so the failure appears in the launcher log before the server ever binds.

open as a page

How do TGI's /generate and /v1/chat/completions routes differ?

level: middleimportance: must knowfreq 58%

basics

~20 s

/generate is TGI's native route: you send raw prompt text under "inputs" with a "parameters" object and get back "generated_text". /v1/chat/completions is the Messages API: you send a "messages" array, the server applies the model's chat template, and the response is OpenAI-shaped.

open as a page

How does TGI's grammar parameter force a response to match a JSON schema?

level: middleimportance: should knowfreq 45%

basics

~20 s

TGI accepts a grammar object in a /generate request's parameters, either type json with a JSON Schema or type regex with a pattern. At each decoding step the server masks out any token that could not continue a valid match, so the output is guaranteed to parse.

open as a page

In TGI, what do --max-input-tokens and --max-total-tokens each cap?

level: middleimportance: should knowfreq 52%

basics

~20 s

--max-input-tokens caps the prompt length TGI will accept for one request; --max-total-tokens caps prompt plus generated tokens together. The router validates both before queueing and rejects an oversized request with a 422 validation error rather than truncating it.

open as a page

Which TGI /metrics series separate queueing time from GPU inference time?

level: seniorimportance: should knowfreq 40%

basics

~20 s

TGI exposes Prometheus metrics on /metrics, where tgi_request_queue_duration records time spent waiting before the GPU touched a request and tgi_request_inference_duration records the model work itself. Splitting slow end-to-end latency across those two is the first triage step.

open as a page

What does TGI's --speculate flag enable, and when does it stop paying off?

level: seniorimportance: should knowfreq 36%

basics

~20 s

TGI's --speculate sets how many extra tokens are proposed per step and verified in one forward pass - from Medusa heads when the checkpoint has them, otherwise by n-gram lookup in the existing context. It helps a lightly loaded server and fades as concurrency fills the GPU.

open as a page

How does TGI's --waiting-served-ratio decide when to pause decoding for queued requests?

level: seniorimportance: should knowfreq 34%

basics

~20 s

TGI's --waiting-served-ratio is the ratio of waiting requests to running requests at which the router will interrupt decoding, prefill the queued requests, and fold them into the running batch. A lower value admits newcomers sooner and protects their time-to-first-token.

open as a page

In TGI, what does --num-shard do and what must the container provide?

level: seniorimportance: should knowfreq 45%

basics

~20 s

--num-shard tells the TGI launcher how many shard processes to start, splitting the model tensor-wise across that many GPUs on one node. The container must expose those GPUs and raise shared memory (--shm-size 1g), and the shard count has to divide the model's attention head count.

open as a page

In TGI, when is --quantize bitsandbytes the wrong way to shrink a model?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Whenever throughput matters. bitsandbytes quantizes weights on the fly at load, so it needs no special checkpoint and buys VRAM headroom, but its kernels are markedly slower than serving a pre-quantized AWQ, GPTQ, Marlin or FP8 checkpoint with the matching --quantize value.

open as a page

When would you choose TGI over vLLM for a production LLM endpoint?

level: principalimportance: should knowfreq 44%

basics

~20 s

Choose on mechanics and fit, not popularity: TGI is the low-friction option for teams already on the Hugging Face stack, with built-in grammar guidance and a Rust router. vLLM moves faster on new architectures and quantization formats. Both do continuous batching and paged KV cache.

open as a page