llama-server and Endpoints
0 questions
llama-server loads a GGUF with -m or -hf, sets -ngl offload, -c context and -np slots, and serves OpenAI and Anthropic routes with GBNF or JSON-schema output. Interviewers probe which flag fixes what.
questions
no questions here yet
this part of the tree is still being written
>
0 questions in this topic or below it