GPU Offload and Memory
0 questions
Ollama decides how many layers fit in VRAM, while num_ctx, OLLAMA_NUM_PARALLEL and the KV-cache type decide what is left. A 48%/52% CPU/GPU split in ollama ps is the classic slow-model question.
questions
no questions here yet
this part of the tree is still being written
>
0 questions in this topic or below it