How do you convert a Hugging Face checkpoint into a TensorRT-LLM engine?
answer
- HF weights are not built directly
- an intermediate TRT-LLM checkpoint exists
- tp_size shards at conversion, not at launch
- quantization pass replaces plain conversion
- tokenizer is not inside the engine
basics
~20 sTwo steps. A per-model convert_checkpoint.py script rewrites the Hugging Face weights into a TensorRT-LLM checkpoint directory — a config.json plus one weight file per rank, sharded by the --tp_size you choose. Then trtllm-build compiles that directory into an engine.
solid answer
~40 sTensorRT-LLM does not compile a Hugging Face repo directly; there is an intermediate **TensorRT-LLM checkpoint**. The per-model `convert_checkpoint.py` example script reads `--model_dir`, applies `--dtype`, and writes `--output_dir` containing a `config.json` (architecture, dtype, tensor- and pipeline-parallel degrees) plus per-rank weight files. Sharding happens here: `--tp_size 2` produces weights already split for two ranks, which is why the parallel degree is a conversion-time decision rather than a serving flag. If you want a quantized engine, you substitute a calibration/quantization pass for the plain conversion, which emits a quantized checkpoint in the same layout. `trtllm-build --checkpoint_dir ... --output_dir ...` then compiles that into engines, one per rank. At serve time the runtime's world size must match the parallel degree the checkpoint was converted for; mismatches fail at load, not gracefully.
code
bash · 14 lines# 1. Hugging Face weights -> TensorRT-LLM checkpoint (sharded for 2 ranks)
python convert_checkpoint.py \
--model_dir ./Meta-Llama-3.1-8B-Instruct \
--output_dir ./ckpt-llama3-8b-tp2 \
--dtype bfloat16 \
--tp_size 2
# 2. TensorRT-LLM checkpoint -> engines
trtllm-build \
--checkpoint_dir ./ckpt-llama3-8b-tp2 \
--output_dir ./engines/llama3-8b-tp2 \
--gemm_plugin auto \
--max_batch_size 32 \
--max_seq_len 8192go deeper
Remember the order: download the checkpoint, convert it to TensorRT-LLM format, then build the engine. Name the two commands and say the tokenizer is handled separately.
Explain the intermediate checkpoint as its own artifact, and state clearly that --tp_size shards weights at conversion so the parallel degree cannot be changed later without redoing the pipeline.
Talk about running the whole pipeline inside one pinned container, recording an artifact key that includes model revision, precision, parallel degree and library version, and the failure modes you have actually hit — world-size mismatch and tokenizer drift.
Frame conversion coverage as a platform risk: every model family your product may want must have a supported conversion path, and that dependency decides how quickly you can adopt a new architecture at all.
## The three artifacts It helps to name the things, because interviews get muddled when "model" means three different files: 1. **The Hugging Face checkpoint** — what you download: `config.json`, `*.safetensors`, tokenizer files, in the layout the Transformers library expects. 2. **The TensorRT-LLM checkpoint** — an intermediate, framework-specific layout: a `config.json` describing architecture, dtype, quantization and the tensor/pipeline parallel degrees, plus weight files split per rank. 3. **The engine directory** — the compiled output of `trtllm-build`: one serialized engine per rank plus the build config. The conversion step exists because TensorRT-LLM needs weights laid out the way its own layer definitions expect — fused QKV projections, per-rank shards, transposed matrices where a plugin wants them that way — and because quantization scales, when present, have to live alongside the weights in a form the builder understands. ## Step one: conversion The conversion scripts ship per model family in the repository's examples, because each architecture needs its own weight-name mapping. The common arguments are stable across them: - `--model_dir` — the downloaded Hugging Face checkpoint. - `--output_dir` — where the TensorRT-LLM checkpoint is written. - `--dtype` — the compute dtype for an unquantized conversion, typically `float16` or `bfloat16`. - `--tp_size` / `--pp_size` — tensor- and pipeline-parallel degrees. That last pair is the one candidates miss. Sharding is baked into the converted weights: with `--tp_size 4`, the attention heads and MLP dimensions are split four ways and written as four files. You cannot decide at deploy time to run the same artifact on two GPUs instead of four. Changing the parallel degree means going back to conversion and rebuilding. ## Step one, quantized variant If the target precision is FP8, INT8 or INT4-AWQ, you replace the plain conversion with a quantization pass. That pass runs the model over a small calibration set, computes the scaling factors the chosen format needs, and writes a checkpoint in the same layout but flagged with its quantization scheme. `trtllm-build` reads that flag and compiles the matching kernels. The two paths converge: whatever produced the checkpoint, the build command that follows looks the same. ## Step two: build `trtllm-build --checkpoint_dir <converted> --output_dir <engines>` compiles. This is where you declare the runtime ceilings — `--max_batch_size`, `--max_input_len`, `--max_seq_len`, `--max_num_tokens` — and plugin choices such as `--gemm_plugin`. With multiple ranks, `--workers` lets the ranks build in parallel, at the cost of holding several builders in memory at once. ## What travels with the engine, and what does not The engine holds weights and kernels. It does **not** hold the tokenizer. Whatever serves the engine — the Triton TensorRT-LLM backend's preprocessing and postprocessing models, or a higher-level server — must be pointed at the original Hugging Face directory for tokenizer files. A very common first deployment failure is an engine that loads fine while text in and text out is broken, because the tokenizer path was never wired up or points at a different revision of the model than the weights came from. ## Version and reproducibility discipline The conversion scripts live in the same repository as the builder and evolve with it: converting with one release and building with another is asking for a mismatch. Pin both to the same TensorRT-LLM version, and prefer running the whole pipeline inside the matching container image so the CUDA, TensorRT and library versions agree. Because the output is a binary artifact rather than a reproducible-on-demand file, record what produced it. A useful artifact key is: model repo and revision, dtype or quantization format, `tp_size`/`pp_size`, the build ceilings, the TensorRT-LLM version and the GPU architecture. Six months later, that key is the only way to answer "is the engine running in production the one we evaluated?" ## Common failure modes - **World-size mismatch.** The engine was converted with `tp_size 2` but launched on one GPU, or with a Triton `gpu_device_ids` list that does not supply the expected ranks. This fails at load. - **Unsupported architecture.** There is no conversion script for the model family, or the checkpoint uses a variant the mapper does not know. This is a hard stop, and it is the main reason a brand-new model appears in a Python-level engine days before it appears here. - **Silent tokenizer drift.** Engine built from revision A, tokenizer served from revision B. - **Disk.** Converted checkpoints and engines are each roughly model-sized, so a naive pipeline needs several times the model's footprint in scratch space. ## The interview shape A strong answer walks the pipeline in order, names the intermediate checkpoint as a distinct artifact, points out that the parallel degree is chosen at conversion, and mentions that the tokenizer is not part of the engine. A weak answer describes it as "point the server at the Hub model", which is the mental model of a Python-level engine, not a compiled one.
- You want to move a tp_size 4 deployment onto two GPUs. What has to change?You go back to conversion. The parallel degree is baked into the converted weight shards, so you re-run the conversion with `--tp_size 2` and rebuild the engines from that checkpoint. There is no runtime flag that re-shards an existing engine, and the serving world size must match what the checkpoint was converted for.
- The engine loads and serves, but the output text is garbage. Where do you look first?The tokenizer wiring. The engine contains weights and kernels, not tokenizer files, so the serving layer must point at the original model directory — and at the same revision the weights came from. Mismatched or missing tokenizer configuration produces fluent-looking token ids decoded into nonsense, or immediate special-token breakage.
- Why does a brand-new model architecture often reach a Python-level engine before TensorRT-LLM?Because it needs a weight-mapping conversion script and layer definitions the builder can compile, not just a Transformers config the runtime can interpret. Until someone adds the mapping for that family, there is no path from the Hub checkpoint to a TensorRT-LLM checkpoint, so the model simply cannot be built.
saying these in an interview costs you the question
- Says you point trtllm-build straight at the Hub repository
- Thinks tensor-parallel degree is a serving-time flag
- Assumes the tokenizer ships inside the engine
- Mixes conversion from one release with a build from another
- Forgets the pipeline needs several model-sized scratch volumes