TGI fails to download a gated Llama repo — how do you fix the launch?
answer
- it dies before the port opens
- two parties: the account and the process
- licence acceptance is per account
- the env var name changed spelling
- a warm cache needs no credential at all
basics
~20 sAccept the model's licence with the Hugging Face account that owns the token, then inject the token into the container at runtime as the HF_TOKEN environment variable. The download happens once at startup, so the failure appears in the launcher log before the server ever binds.
solid answer
~50 sTwo things have to line up. First, the account behind the token must have been granted access to the gated repo — a valid token for an account that never accepted the terms still gets refused. Second, the container has to see that token: add `-e HF_TOKEN=$HF_TOKEN` to the `docker run` line (older TGI docs use the legacy `HUGGING_FACE_HUB_TOKEN` spelling; `HF_TOKEN` is the current one). The symptom is diagnostic: the launcher's **download** stage fails and the process exits, so you see an authorization error in the logs and no listening port at all — quite different from a serving error, which would come back over HTTP. Treat the token as a secret: never bake it into an image layer or a committed compose file; inject it from a Kubernetes secret or the platform's secret store at runtime. If you already have the weights on disk from a prior download, mounting that cache at `/data` means the run needs no token at all.
code
bash · 7 linesexport HF_TOKEN=hf_xxx
docker run --gpus all --shm-size 1g -p 8080:80 \
-e HF_TOKEN \
-v "$PWD/tgi-data:/data" \
ghcr.io/huggingface/text-generation-inference:3.3.5 \
--model-id meta-llama/Llama-3.1-8B-Instructgo deeper
Know that gated repos need a Hugging Face token passed into the container as the HF_TOKEN environment variable, and that the licence has to be accepted by that token's account first.
Explain that the token is consumed once, in the launcher's download stage, so the failure shows up in logs with no port listening — and that you inject it at runtime rather than building it into the image.
Show how you keep the Hub off the startup path: a pre-populated cache mounted at /data or weights staged locally and passed to --model-id as a path, with the token in a secret store, scoped read-only and rotatable by restart.
Own the policy side: who accepts model licences on the organisation's behalf, how those terms are tracked per deployed model, and whether serving gated weights across regions or tenants stays within the licence you agreed to.
## Where the failure actually happens A TGI container starts in stages: resolve and **download** the model repo, load and shard the weights onto GPU, warm up, then bind the HTTP port. Authorization for a gated repo is checked in the very first stage. That gives the failure a distinctive shape — the launcher logs an authorization error against the Hub and the container exits or crash-loops, and nothing ever answers on the published port. Candidates who go looking for a 401 in the HTTP response are looking in the wrong place; there is no HTTP response, because the router never started. This also means the check is **per container start**, not per request. TGI holds one model, downloaded once at boot. There is no per-caller Hub identity at inference time — whatever authentication your endpoint requires from clients is a separate concern from the token that fetched the weights. ## The two conditions **Access must have been granted to the account.** A gated repo requires a human to accept its terms under a specific Hugging Face account. A token belonging to a different account, or to an organisation that never accepted, authenticates fine and is still refused the files. This is the failure that survives three rounds of "but the token is valid" — the token is valid, the account is not authorized. **The token must reach the process.** The launcher reads it from the environment as `HF_TOKEN`. In Docker that means `-e HF_TOKEN=$HF_TOKEN` before the image name; note that the bare `-e HF_TOKEN` form (no value) passes the variable through from your shell, which keeps the secret off the command line and out of shell history. Older TGI documentation uses `HUGGING_FACE_HUB_TOKEN`; that spelling is the legacy name and `HF_TOKEN` is what to write today. ## Handling the token as a secret A read token for a Hub account is a credential, and it belongs where credentials belong: - **Never in the image.** An `ENV HF_TOKEN=...` or a token pasted into a Dockerfile is baked into a layer, survives in every registry copy, and is trivially recoverable from the image history. - **Never in a committed manifest.** A token in a compose file or a plain Kubernetes manifest is a token in git. - **Inject at runtime.** In Kubernetes, a `Secret` projected as an env var; on a managed platform, its secret store. Scope the token to read-only, and prefer a fine-grained token limited to the repos you actually serve. - **Rotate it.** Because the token is only consumed at container start, rotation is unusually cheap: change the secret, restart the pods. Nothing long-lived is holding it. ## The better production answer: don't need the token at boot Depending on the Hub at every pod start is a liveness dependency you may not want. Two ways out: **Pre-populate the cache.** Download the repo once and mount that directory at `/data`. Weights already present are not re-fetched, so the run needs no credentials and no network to the Hub. This is also the fix for scale-out cost: ten pods pulling 140 GB each is ten times the egress and ten cold starts. **Serve from a local path.** `--model-id` accepts a filesystem path inside the container, not only a Hub repo id. Stage the weights into an internal artifact store or a shared volume, point the launcher at that path, and the Hub is out of your critical path entirely. In an air-gapped environment this is the only option. Both approaches turn a runtime authorization problem into a build- or provisioning-time one, which is where a licence acceptance really belongs. ## Pinning alongside the token While you are fixing the download, pin `--revision` to a commit sha. A gated repo is still a mutable repo; the account that gained access today can be served different bytes next month. Pinning the revision, pinning the image tag, and verifying against `/info` after rollout together give you a deployment whose model contents are reproducible. ## Distinguishing the neighbouring failures Three startup failures look similar in a dashboard and are not the same thing: - **Authorization** — the account lacks access or the token is missing; the download stage errors out. - **Not found** — a typo in the repo id, or a `--revision` that does not exist; also a download-stage failure, but the log names a missing repo or revision. - **Out of memory** — download succeeded and the shard died trying to place weights; the log shows a CUDA allocation failure after the download finished. Reading which stage the log stopped at tells you which of the three you have in seconds, and that habit is what an interviewer is actually probing.
- The token is valid and exported, but the download still fails. What is left?Almost always the account behind the token has not been granted access to that gated repo — authentication succeeded, authorization did not. Check that the terms were accepted by the exact account (or organisation) the token belongs to. The other candidates are a typo in the repo id and a `--revision` that does not exist; the launcher log names a missing repo or revision in those cases, rather than an authorization error.
- How would you run TGI on a gated model in an air-gapped cluster?Stage the weights outside the cluster once, publish them to an internal artifact store or a shared volume, mount them into the container, and pass that filesystem path to `--model-id` instead of a Hub repo id. No token and no Hub reachability are needed at run time. Licence acceptance becomes a provisioning-time step performed once by a human, which is where it belongs anyway.
- Does TGI re-check the token for each inference request?No. The token is used only during the download stage at container startup, to fetch the weights for the single model this server holds. After that the process has the files and never consults the Hub again. Any authentication of your own callers is a separate layer — a gateway, proxy or ingress in front of the server — and is unrelated to the Hub credential.
saying these in an interview costs you the question
- Baking the token into the image with ENV
- Assuming a valid token implies repo access
- Looking for an HTTP 401 when the server never started
- Re-downloading gated weights on every pod start
- Committing the token in a compose or Kubernetes manifest