In TGI, what does --num-shard do and what must the container provide?
answer
- N processes, one model, one node
- shared memory default will bite you
- the count must divide the heads
- every layer ends in a collective
- fit first, throughput is not linear
basics
~20 s--num-shard tells the TGI launcher how many shard processes to start, splitting the model tensor-wise across that many GPUs on one node. The container must expose those GPUs and raise shared memory (--shm-size 1g), and the shard count has to divide the model's attention head count.
solid answer
~60 sThe launcher supervises N shard processes, each holding a slice of every weight matrix on its own GPU; `--num-shard` sets N and `--sharded true` is the boolean form that lets TGI use all visible devices. Which physical cards those are is Docker's decision plus `CUDA_VISIBLE_DEVICES` inside the container — the launcher shards across what it can see. Three container-level requirements bite in practice. The GPUs must actually be exposed (`--gpus all` or an explicit device list). `--shm-size` must be raised, because the shards coordinate through shared memory and Docker's 64 MB default is not enough — this is the classic "it hangs during load" failure. And the shard count must divide the model's attention head count evenly, so 3 shards on a model with 32 heads fails at startup. Sharding is primarily how a model **fits**. It also splits per-layer work, but each layer now ends in a collective across shards, so throughput does not scale linearly with N — and it is single-node: all shards live in one container on one host.
code
bash · 6 linesdocker run --gpus all --shm-size 1g -p 8080:80 \
-v "$PWD/tgi-data:/data" \
-e HF_TOKEN \
ghcr.io/huggingface/text-generation-inference:3.3.5 \
--model-id meta-llama/Llama-3.1-70B-Instruct \
--num-shard 4go deeper
Know that --num-shard splits one model across several GPUs on the same machine, and that the container needs GPU access plus a raised --shm-size for it to start.
Explain that the launcher runs one shard process per GPU with a slice of each weight matrix, that the count must divide the attention head count, and that the shards coordinate through shared memory during load.
Show the fit-versus-scale judgment: shard because the weights do not fit on one card or because per-request latency matters, replicate when the model fits and you want capacity, and check the interconnect before promising any scaling number.
Own the fleet-shape decision — which GPU SKU and interconnect you standardise on, how many shards the chosen model architectures permit, and whether large-model fit should be bought with sharding, a quantized checkpoint, or a different model entirely.
## What the flag starts When you pass `--num-shard 4`, the TGI launcher does not start four servers. It starts one router and **four shard processes**, each pinned to one GPU, each holding a slice of the model's weights. A single logical model spans all four devices; a request is served by all four cooperating on every token. `--sharded true` is the related boolean — it turns sharding on and lets TGI use the devices it can see, while `--num-shard` states the count explicitly. Being explicit is the better habit in a deployment manifest, because it fails loudly if the node presents a different number of GPUs than you expected instead of silently changing the topology. The launcher log makes the topology visible: it reports each shard becoming ready, and only then connects the router. A run that logs "shard 0 ready" and then stalls is almost always the shared-memory problem below. ## The container-level requirements **GPU visibility.** `--gpus all` on the `docker run` line, or an explicit device list like `--gpus '"device=0,1"'`. Inside the container, `CUDA_VISIBLE_DEVICES` narrows further. If you want two TGI containers on a 4-GPU box, give each two devices and set `--num-shard 2` in both; if the containers both see all four, they will fight over memory. **Shared memory.** `--shm-size 1g`. The shard processes exchange data through `/dev/shm`, and Docker's default of 64 MB is too small. The failure is not a clean error message — it is typically a hang or an obscure crash partway through model load, which is why every TGI launch example carries this flag whether it shards or not. In Kubernetes the equivalent is mounting an `emptyDir` with `medium: Memory` at `/dev/shm`, since pods inherit the same small default. **Divisibility.** Tensor parallelism splits attention heads across shards. The shard count must divide the head count evenly, so on a model with 32 attention heads, 1, 2, 4, 8, 16 and 32 shards are valid and 3, 5 or 6 are not. This is checked at startup and it is a hard failure, not a warning. It is also why you cannot casually "just add one more GPU" to a sharded deployment. **One node.** All shards are processes of one launcher inside one container. There is no multi-node form of this flag: a model too large for the GPUs on a single host is a different architecture problem, not a bigger `--num-shard`. ## Why you shard The primary reason is **fit**. A model whose weights exceed one card's memory, once you have also left room for the cache and activations, cannot be served on that card at any batch size. Sharding cuts the per-GPU weight footprint by roughly N, which is what makes a large model servable at all. As a side effect, each shard also holds only its slice of the cache, so the freed memory per device raises how much concurrent context fits. The secondary reason is latency: the per-layer matrix multiplications are split, so with a fast enough interconnect a single request can decode faster than on one card. This is real, and it is the reason a latency-sensitive deployment may shard a model that would technically fit. ## Why it does not double your throughput The expectation that two GPUs give twice the tokens per second is the most common wrong answer here. Each transformer layer ends with a collective communication step to recombine the shards' partial results, and that step happens on **every layer, every token**. On a host with fast GPU-to-GPU links this overhead is modest; on a host where the cards talk over slower PCIe links it can dominate, and a sharded run can even be slower per token than a single card that could have held the model. The operational consequence: check the interconnect before you promise scaling numbers, and measure. Two GPUs running two independent single-GPU replicas of a model that fits will almost always beat one 2-way sharded copy on aggregate throughput, because the replicas communicate not at all. Shard when you must fit, or when per-request latency is the goal; replicate when you want capacity and the model already fits. ## Failure modes worth naming - **Hang during load** — shared memory too small. - **Startup rejection about head divisibility** — invalid shard count for that architecture. - **Out of memory on one shard while others look fine** — the split is not perfectly even across all tensors, and the device that also hosts other work loses first. Check that nothing else (another container, a stray process) holds memory on those cards. - **Every request slow and GPUs at low utilisation** — communication-bound; look at the interconnect and consider replicas instead. ## Sizing the decision The honest procedure is arithmetic before flags: estimate weight bytes from parameter count and dtype, add the cache you intend to support at your configured token limits, add activation and framework overhead, and compare against the memory of one card. If it does not fit, the smallest N that both fits and divides the head count is the answer — not the largest N available, because every additional shard adds collective overhead for a model that was already going to fit.
- You raised --num-shard from 2 to 4 and throughput barely moved. What do you check?The interconnect first. Each layer ends in a collective across shards on every token, so on hosts where the GPUs communicate over slower links that overhead grows with shard count and can swamp the compute you split. Check GPU utilisation — if the devices are busy waiting rather than computing, you are communication-bound. At that point two independent replicas of a model that already fits will usually beat one 4-way sharded copy on aggregate throughput.
- How do you run two separate TGI containers on one 4-GPU node?Give each container an explicit device subset with `--gpus '"device=0,1"'` and `'"device=2,3"'`, and set `--num-shard 2` in both. Each launcher then shards only across the devices it can see. If both containers see all four GPUs, both will try to claim memory on the same cards and one will fail at load or later at warmup — the memory reservation is per process and nothing arbitrates between containers.
- Why does the shard count have to divide the attention head count?Because tensor parallelism splits the attention heads across shards, and an uneven split has no valid mapping — some shard would own a fractional head. So on a 32-head model, 1, 2, 4, 8, 16 and 32 are the legal counts and 3 or 6 are not. The launcher checks this at startup and refuses to run, which is a good thing: the alternative would be a silently wrong model.
- Can --num-shard spread a model across two machines?No. The shards are child processes of one launcher inside one container on one host, so the flag is single-node by construction. A model too large for all the GPUs on your biggest available node is an architecture problem — a different sharding strategy, a smaller or quantized checkpoint, or larger nodes — not a higher shard count.
saying these in an interview costs you the question
- Expecting throughput to scale linearly with shard count
- Leaving Docker's default shared memory on a sharded run
- Picking a shard count that does not divide the head count
- Thinking --num-shard spreads the model across nodes
- Sharding a model that already fits, instead of adding replicas