Meta Llama
Meta's open-weight family — the models you can download, quantize, fine-tune and serve yourself. Interviews use Llama as the vehicle for every self-hosting question: sizing, serving throughput, and the license fine print.
on this pageshowhide
explore
- Model Family and Sizes5 questions
- Prompt and Chat Templates5 questions
- Quantization6 questions
- Fine-Tuning6 questions
- Self-Hosted Inference6 questions
- Licensing and Compliance6 questions
questions
page 2 of 2At a fixed VRAM budget, do you run a bigger Llama at 4-bit or a smaller one at FP16?
basics
~20 sDown to about 4 bits, the larger model usually wins: parameter count buys more capability than precision does, and big models absorb quantization loss better than small ones. Below roughly 3 bits the ordering flips, and throughput and kernel support can override both.
Llama 70B won't fit one GPU — how do you choose between tensor parallelism, quantization, and a smaller model?
basics
~20 sDecide from the quality floor and the latency SLO, not from what fits. Tensor parallelism keeps full precision but needs multiple GPUs with fast interconnect; quantization fits fewer GPUs at some accuracy cost; a smaller model is cheapest and often good enough once measured on your own evaluations.
In Llama 3.1's chat format, what do <|eom_id|> and the ipython role mean?
basics
~20 s<|eom_id|> ends an assistant message that is waiting on a tool result instead of handing the turn back to the user. The tool's output is fed back as a message with the ipython role, and the model then continues.
How would you assess Llama licence risk before committing a product roadmap to it?
basics
~20 sAssess it per release, not once: check the very-large-operator threshold against group MAU at that version's release date, any regional carve-outs, the attribution and naming obligations your product must carry, the termination clauses, and how expensive it would be to swap models if rights ended.
showing 31–34 of 34