skip to content

llama-server and Endpoints

0 questions

llama-server loads a GGUF with -m or -hf, sets -ngl offload, -c context and -np slots, and serves OpenAI and Anthropic routes with GBNF or JSON-schema output. Interviewers probe which flag fixes what.

questions

no questions here yet

this part of the tree is still being written

>
0 questions in this topic or below it