When should you serve Llama with Ollama versus vLLM?
answer
- one user versus many users
- laptop CPU versus datacentre GPU
- GGUF quantized weights versus VRAM-resident weights
- continuous batching is the dividing feature
- both expose an OpenAI-shaped route
basics
~20 sOllama wraps llama.cpp for local, single-user use: one command pulls a quantized model and runs it on a laptop, CPU included. vLLM is a GPU server built for concurrent production traffic, using continuous batching and a paged KV cache.
solid answer
~50 sThey solve different problems. **Ollama** is a model manager and daemon on top of llama.cpp: `ollama run llama3.1` pulls a quantized GGUF build and serves it, and llama.cpp can run entirely on CPU or split layers between CPU and GPU with `-ngl`. That makes it the right tool for laptops, demos, air-gapped boxes and single-developer use — but it is tuned for a small number of simultaneous requests, and per-request throughput collapses as concurrency rises. **vLLM** is a datacentre inference server. It loads the full-precision or GPU-quantized weights onto CUDA/ROCm GPUs, keeps a paged KV cache, and schedules requests with continuous batching so dozens or hundreds of streams share one GPU efficiently. It also shards a model across GPUs with `--tensor-parallel-size`. The cost is that the model must fit in VRAM and you need real GPUs. Hugging Face TGI occupies roughly the same slot as vLLM. Rule of thumb: one user and no GPU budget → Ollama/llama.cpp; a service with an SLO → vLLM or TGI.
code
bash · 6 linesollama pull llama3.1:8b
ollama run llama3.1:8b "Summarise the CAP theorem in two sentences."
curl http://localhost:11434/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"hi"}]}'go deeper
Be able to say plainly that Ollama is the easy local way to run Llama and vLLM is the GPU server for real traffic, and name one command for each.
Explain the mechanism behind the split: continuous batching and a paged KV cache in vLLM versus llama.cpp's slot model, plus GGUF-on-CPU versus fp16-in-VRAM.
Show you have operated both — cold-load pauses from model unloading, per-slot context division, VRAM headroom, and the metrics you watch before declaring a stack production-ready.
Own the fleet decision: standardising on one OpenAI-compatible contract so laptops, staging and production differ only by base URL, and knowing when the cost of GPU inference beats a hosted API at all.
## The four stacks, briefly **llama.cpp** is a C/C++ inference engine. It loads models in the GGUF format, runs on CPU with SIMD kernels, and can offload some or all transformer layers to a GPU with `--n-gpu-layers` (`-ngl`). Its `llama-server` binary exposes an HTTP API. **Ollama** is a model manager and background daemon built on the same ggml/llama.cpp core. It adds a registry (`ollama pull`), a local model store, a `Modelfile` for baking in parameters and system prompts, and automatic load/unload of models. Most people who say "I ran Llama locally" ran Ollama. **vLLM** is a Python/CUDA inference server. Its distinguishing features are PagedAttention (a block-based KV cache) and continuous batching (iteration-level scheduling), plus multi-GPU sharding via `--tensor-parallel-size`. You start it with `vllm serve <model>`. **Text Generation Inference (TGI)** is Hugging Face's Rust+Python server in the same class: continuous batching, paged KV cache, sharding via `--num-shard`, shipped as a Docker image. ## The dividing line is concurrency, not model quality All four run the same Llama weights and, at temperature 0 with the same quantization, produce broadly similar text. What differs is what happens when more than one person calls at once. Ollama and llama.cpp were designed around an interactive loop: one prompt, one stream of tokens. llama.cpp has parallel slots (`--parallel N`), and Ollama exposes `OLLAMA_NUM_PARALLEL`, but the slot model statically partitions the context window — with `-np 4 -c 8192`, each slot gets 2048 tokens — and the scheduler is far simpler than vLLM's. Throughput per GPU under load is not the design goal. vLLM and TGI were designed around a queue. Requests join and leave the running batch at every decode step, the KV cache is allocated in small blocks so memory is not reserved for the longest possible output, and GPU utilisation stays high because the batch never sits half-empty waiting for one slow generation to finish. ## Hardware and memory format Ollama/llama.cpp read GGUF, a format built for aggressive weight quantization and CPU execution. A 4-bit 8B model is roughly 4.5–5 GB and runs acceptably on a modern laptop, GPU optional. This is the only one of the four you can realistically hand to a colleague with a MacBook. vLLM and TGI expect the model in GPU memory. At fp16 an 8B Llama is about 16 GB of weights before any KV cache; a 70B is about 140 GB and needs several GPUs or GPU-side quantization (AWQ/GPTQ/FP8). There is no meaningful CPU fallback — if it does not fit, it does not serve. ## Operational surface vLLM and TGI give you a Prometheus `/metrics` endpoint, health checks, per-request scheduling knobs (`--max-num-seqs`, `--max-model-len`, `--gpu-memory-utilization`) and Kubernetes-shaped deployment. Ollama gives you a daemon and environment variables (`OLLAMA_HOST`, `OLLAMA_KEEP_ALIVE`, `OLLAMA_MAX_LOADED_MODELS`) and will happily unload a model that has been idle — great on a laptop, surprising in production, where the next request then pays a cold-load penalty. One thing they share: all of them expose an OpenAI-shaped `/v1/chat/completions` route, so application code can move between them with a base-URL change. ## Choosing Ask two questions. First, how many concurrent requests must one machine sustain? If the answer is "one or two, interactively", Ollama is less operational work and runs on hardware you already own. If it is "tens to hundreds with a latency target", you want continuous batching, which means vLLM or TGI. Second, do you have GPUs with enough VRAM for the model you want? If not, the question answers itself: llama.cpp's CPU and partial-offload paths are the only ones that will run at all, and you accept single-digit tokens per second. A common production pattern is both: Ollama on developer laptops for iteration, vLLM behind the staging and production endpoints, with the application talking OpenAI-compatible HTTP to whichever is configured.
- Ollama and llama.cpp run the same engine — when would you drop to llama.cpp directly?When you need control Ollama abstracts away: exact `llama-server` flags such as `--parallel`, `--ctx-size`, KV-cache type, or flash attention; a GGUF file you built yourself rather than one from Ollama's registry; or an embedded/edge build where you want the binary and nothing else. Ollama is the convenience layer; llama.cpp is the engine, and anything Ollama does you can do directly with more effort.
- Your Ollama box is fine for one user but falls over at ten concurrent chats. What is actually happening?Requests are being serialised or squeezed into a few static slots, so each waits behind the others and the shared context budget shrinks per slot. GPU utilisation looks low while latency climbs — the classic signature of no continuous batching. The fix is not more Ollama tuning; it is moving to a server that schedules at the iteration level, such as vLLM or TGI.
- Does switching from Ollama to vLLM require application code changes?Usually only the base URL, the API-key placeholder and the `model` string, because both expose OpenAI-compatible `/v1/chat/completions`. What does change is behaviour at the edges: vLLM wants the exact served model name, some parameters are honoured by one and ignored by the other, and features like tool calling need explicit server flags in vLLM. Re-test streaming and tool paths after the swap.
saying these in an interview costs you the question
- Claiming vLLM runs fine on a CPU-only laptop
- Thinking Ollama and vLLM produce different model quality
- Assuming Ollama scales to production concurrency with more RAM
- Believing you must rewrite client code to move between the two
- Saying llama.cpp cannot use a GPU at all