skip to content

GPU Offload and Memory

0 questions

Ollama decides how many layers fit in VRAM, while num_ctx, OLLAMA_NUM_PARALLEL and the KV-cache type decide what is left. A 48%/52% CPU/GPU split in ollama ps is the classic slow-model question.

questions

no questions here yet

this part of the tree is still being written

>
0 questions in this topic or below it