How do TGI's /generate and /v1/chat/completions routes differ?
answer
- who builds the prompt string
- one takes text, one takes roles
- the template lives with the tokenizer
- change a base URL, keep the client
- the model field is a label here
basics
~20 s/generate is TGI's native route: you send raw prompt text under "inputs" with a "parameters" object and get back "generated_text". /v1/chat/completions is the Messages API: you send a "messages" array, the server applies the model's chat template, and the response is OpenAI-shaped.
solid answer
~50 sThey differ in who builds the prompt. On `/generate` you post `{"inputs": "...", "parameters": {...}}` and TGI runs exactly the string you sent — **no chat template is applied**, so if the model expects role markers you must write them yourself. The reply is `{"generated_text": "..."}`, with per-token detail when you ask for it. On `/v1/chat/completions` you post a `messages` array of roles and content; the server renders it through the tokenizer's chat template and returns the OpenAI response shape with `choices`. That is what lets an existing OpenAI client work by changing only its base URL to `http://host:8080/v1` — the `model` field is effectively a label, since one TGI process serves one model. `/generate_stream` is the native streaming counterpart, delivering incremental token events, while the Messages API streams with `"stream": true` in the OpenAI delta shape. Pick native when you want raw prompt control, pick the Messages API when you want existing OpenAI tooling to work unchanged.
code
bash · 6 linescurl http://localhost:8080/generate \
-H 'Content-Type: application/json' \
-d '{
"inputs": "Write a haiku about GPUs.",
"parameters": {"max_new_tokens": 40, "temperature": 0.7}
}'go deeper
Know that one route takes a raw prompt string under inputs and returns generated_text, while the other takes a messages array and returns the OpenAI-shaped response.
Explain that the Messages API applies the tokenizer's chat template server-side while the native route does not, and that this is why an OpenAI client can switch to TGI by changing only its base URL.
Show the consequences: prompt-format bugs that produce plausible-but-worse output on the native route, one model per process regardless of the model field, and auth plus rate limiting living in a proxy in front of the server.
Own the contract choice — standardising on the OpenAI shape keeps hosted-versus-self-hosted a deployment decision rather than a rewrite, at the cost of a compatibility surface you must test rather than assume.
## Two front doors onto one engine A TGI process holds exactly one model and exposes more than one HTTP surface onto it. The routes are not different engines or different capabilities — they are different request contracts, and the real distinction is **who is responsible for turning a conversation into a prompt string**. ## The native route `POST /generate` takes a body of the form `{"inputs": "<prompt text>", "parameters": { ... }}`. The `inputs` string is used as-is. TGI applies no chat template, inserts no system preamble and adds no role markers. Whatever you send is what the model continues. That is a feature when you want it. Instruction-tuned models are trained against a specific formatting — special tokens marking the system, user and assistant turns — and a client that knows this formatting can control it exactly: constructing few-shot prefixes, prefilling the start of the assistant's reply, or serving a base model where the very idea of a chat turn does not apply. It is a footgun when you do not know it: posting bare user text to an instruction-tuned model produces plausible-looking, subtly worse output, and nothing in the response tells you the template was skipped. The response is `{"generated_text": "..."}`. Passing `"details": true` returns additional per-generation detail such as the finish reason and token-level information, which is what you want when you are debugging why a generation stopped where it did. `POST /generate_stream` is the same contract with incremental delivery: instead of one JSON body at the end, the connection stays open and the server emits events carrying successive pieces of the generation, ending with the completed text. Same request shape, different delivery. ## The Messages API `POST /v1/chat/completions` takes `{"model": "tgi", "messages": [{"role": "user", "content": "..."}], ...}`. Here the server does the prompt construction: it renders your message array through the chat template that ships with the model's tokenizer, so the role markers are correct by construction for whatever checkpoint is loaded. Swap the model id in the launch command and the template changes with it; your client code does not. The response comes back in the OpenAI shape — a `choices` array with a message and a finish reason, plus usage information — and `"stream": true` produces the OpenAI incremental delta format rather than TGI's native event shape. The `model` field deserves a note: TGI serves one model per process, chosen at launch by `--model-id`. The field exists because the OpenAI contract requires it, and examples commonly pass the literal `"tgi"`. It is a label, not a selector — you cannot ask a running TGI for a different model. Multi-model routing is a layer above the server. ## Why the compatibility layer matters The practical value of `/v1/chat/completions` is migration. A client library already pointed at a hosted OpenAI-compatible endpoint usually needs one change — its `base_url` set to `http://your-host:8080/v1` — to talk to your self-hosted model instead. No prompt-format research, no rewrite of the streaming handling, and the same code path can be switched back by config. That is why teams standardise on it even when they own both ends: it keeps the option of moving between a hosted provider and self-hosted weights a deployment decision rather than a code change. The cost is that the compatibility is a shape, not a promise of identical behaviour. Provider-specific fields, and behaviours around things the server does not implement, are where the analogy leaks. Test the specific fields your client sends rather than assuming a hosted provider's full surface is reproduced. ## Choosing between them - **Use the Messages API** for ordinary chat traffic, and especially when existing OpenAI-shaped tooling, SDKs or gateways sit in front of the server. Let the server own the template; that is one fewer thing to get wrong per model swap. - **Use `/generate`** when you need the prompt string to be exactly what you wrote: base models, custom few-shot layouts, prefilled assistant openings, or evaluation harnesses that must reproduce a documented prompt byte for byte. A useful rule: if you find yourself hand-writing the model's special role tokens into a `messages` content field, you wanted `/generate`. If you find yourself reimplementing a chat template in client code, you wanted the Messages API. ## Operational notes Both routes hit the same router, so the same validation limits apply to both, and an over-long request is rejected identically whichever door it came through. Neither route authenticates callers — TGI is a model server, not a gateway; if the endpoint needs auth, rate limits per tenant, or request logging, that belongs in a proxy or ingress in front of it. And because both routes share one queue and one set of shards, a heavy batch job on `/generate` and interactive chat on `/v1/chat/completions` are competing for the same GPU; separating those workloads means separate deployments, not separate routes.
- You post plain user text to /generate on an instruction-tuned model and the answers are noticeably worse. Why?Because `/generate` applies no chat template. The model was tuned to see specific role markers around each turn, and without them it is being asked to continue arbitrary text rather than to answer as an assistant. The output stays fluent, which is what makes it a nasty bug — nothing errors. Either format the prompt yourself with the model's expected markers, or move to `/v1/chat/completions` and let the server render the template.
- Can you ask a running TGI server for a different model through the model field?No. A TGI process loads exactly one model, fixed at launch by `--model-id`; the `model` field exists to satisfy the OpenAI request shape and examples commonly pass the literal `"tgi"`. Serving several models means several servers, with routing done by a layer in front of them. That is a real difference from a hosted provider, where one endpoint fronts a catalogue.
- What has to change in an existing OpenAI-compatible client to point it at TGI?Usually just the base URL — set it to `http://your-host:8080/v1` — plus whatever the client requires in the API-key slot, since TGI itself does not authenticate callers. The chat-completions request and response shapes, including streaming deltas, are the ones the client already speaks. Verify the specific fields your client sends rather than assuming a provider's entire surface is reproduced.
- Where should authentication and rate limiting for a TGI endpoint live?In front of the server. TGI is a model server: both routes serve any caller that can reach the port, with no notion of tenants or keys. Put an ingress, gateway or reverse proxy in the path to authenticate, apply per-tenant quotas and log requests, and keep the TGI port off any network the public can reach. Treating the server itself as the security boundary is the common deployment mistake.
saying these in an interview costs you the question
- Assuming /generate applies the model's chat template
- Thinking the model field selects among several loaded models
- Believing OpenAI compatibility means every provider field works
- Expecting TGI to authenticate callers on the /v1 routes
- Treating the two routes as separate engines with different capabilities