Ollama vs vLLM vs llama.cpp comes down to one question: how many people will hit the model at once? For one developer on a laptop, pick Ollama or llama.cpp. For a GPU server with many concurrent users, pick vLLM. Below, I explain why. I also benchmarked Ollama and llama.cpp on the same model file, and the defaults changed the result more than the engine did.
If you haven’t installed any of these yet, start with our LM Studio and Ollama setup guide.
Ollama vs vLLM vs llama.cpp at a Glance
| Ollama | llama.cpp | vLLM | |
|---|---|---|---|
| What it is | Model manager and local server | C/C++ inference engine, CLI and server | High-throughput serving engine |
| Best at | Easiest local setup | Running anywhere, fine control | Many concurrent requests on GPUs |
| Model format | Its own library, or import GGUF / Safetensors | GGUF | Hugging Face weights; GGUF “highly experimental” via a plugin |
| Hardware | macOS, Windows, Linux; CPU, NVIDIA, AMD, Apple Silicon | CPU, Metal, CUDA, HIP, Vulkan, SYCL and more | NVIDIA, AMD and Intel GPUs, TPUs, CPUs; experimental on Apple Silicon |
| API | Native API plus OpenAI-compatible | OpenAI-compatible (llama serve) |
OpenAI-compatible, Anthropic Messages API, gRPC |
| Concurrency default | 1 request per model | Parallel slots, auto-sized | Continuous batching |
| Setup effort | Lowest | Low to medium | Highest |
What Each Engine Actually Is
llama.cpp is the engine underneath much of the local-LLM world. It’s a plain C/C++ implementation with no dependencies. Per its README, it targets “a wide range of hardware.” It runs GGUF files and supports 1.5-bit to 8-bit quantization. As of v0.5.0 it ships a single llama command, so the old llama-server workflow is now llama serve:
# Chat with a model straight from Hugging Face
llama cli -hf ggml-org/Qwen3-4B-GGUF
# OpenAI-compatible server on port 8080
llama serve -m ./Qwen3-4B-Q4_K_M.gguf --port 8080
Ollama wraps local inference in a model manager: ollama pull, ollama run, a model library, and a background service. It exposes both its own API and an OpenAI-compatible one on port 11434. It trades some control for convenience, which is why most people start here. Our Ollama API guide covers the endpoints.
vLLM is a serving engine built for throughput. Its headline features are PagedAttention for KV-cache memory, continuous batching, chunked prefill and prefix caching. It also splits models across GPUs with tensor, pipeline and expert parallelism. So it’s what you run behind an internal API that dozens of users or agents call at once.
uv pip install vllm
vllm serve Qwen/Qwen3-4B
Ollama vs llama.cpp: A Real Benchmark on the Same Model File
To compare Ollama and llama.cpp fairly, I loaded the exact same GGUF file into both: Qwen3-4B at Q4_K_M, 2.5 GB. The machine was an Apple M1 with 16 GB of RAM. I ran Ollama 0.34.4 and llama.cpp 0.5.0 with default settings. Each engine generated 256 tokens for the same prompt. First I sent one request at a time, then four at once.
| Setup | Single request | 4 concurrent (aggregate) | Wall time for 4 requests |
|---|---|---|---|
Ollama, default (OLLAMA_NUM_PARALLEL=1) |
18.7 tok/s | 18.9 tok/s | 54 s |
llama.cpp llama serve, default (4 slots) |
13.7 tok/s | 28.9 tok/s | 35 s |
Ollama with OLLAMA_NUM_PARALLEL=4 |
17.8 tok/s | 35.0 tok/s | 29 s |
Three things stand out.
For one user, Ollama was faster. It generated about 18.7 tokens per second against 13.7 for llama.cpp on the same file. The startup logs show different defaults: Ollama picked a 4,096-token context from available VRAM, while llama serve auto-configured 4 slots with a 40,960-token context each. Same weights, different setup.
Under load, Ollama’s default queues. With four requests at once, default Ollama served them one after another. Aggregate throughput stayed at about 19 tokens per second, and the last user waited nearly a minute. llama.cpp’s parallel slots split the work, so each request ran slower (8 tokens per second) but all four finished in 35 seconds.
One setting flips the result. With OLLAMA_NUM_PARALLEL=4, Ollama handled the same four requests in 29 seconds at 35 tokens per second aggregate, the best result in the test.
I also threw away an earlier run. It showed llama.cpp far ahead, but the Mac was busy (load average around 9), and the numbers didn’t hold up on a rerun. Benchmark on an idle machine, and run each test more than once.
⚠️ Note: These are small-model numbers on one laptop. They show the shape of the difference, not a universal ranking. Run the same test on your own hardware before choosing.
Ollama vs vLLM vs llama.cpp Under Concurrent Load
For a single user, the gap between engines comes mostly from configuration: context size, slot count and quantization. Model size and memory bandwidth set the ceiling, and Ollama and llama.cpp share much of the same low-level code.
Concurrency is where they split:
- Ollama processes one request per model at a time by default. Its FAQ lists
OLLAMA_NUM_PARALLELwith a default of 1, and notes that memory use scales withOLLAMA_NUM_PARALLEL × OLLAMA_CONTEXT_LENGTH. Extra requests wait in a queue, and pastOLLAMA_MAX_QUEUEthe server returns a 503. - llama.cpp runs several parallel “slots” in
llama serve. The-np/--parallelflag defaults to auto. - vLLM batches requests continuously, adding new ones to the running batch as others finish. Combined with PagedAttention, this is what lets one GPU serve many users without the queue growing.
If Ollama is your backend and several people or agents share it, raise the parallel setting before blaming the model:
OLLAMA_NUM_PARALLEL=4 ollama serve
Each parallel slot needs its own context memory, so watch RAM or VRAM when you raise it.
The vLLM Memory Default That Surprises People
vLLM claims most of the GPU as soon as it starts. The --gpu-memory-utilization flag defaults to 0.92, meaning vLLM takes 92 percent of GPU memory for the model and its KV cache, per the engine arguments reference. That’s intentional: the spare memory becomes KV cache, which is what makes high concurrency possible.
It also means a second vLLM instance, or anything else on that GPU, fails to allocate. Run two instances on one GPU and you have to split it yourself:
vllm serve Qwen/Qwen3-4B --gpu-memory-utilization 0.45 --port 8000
vllm serve Qwen/Qwen3-1.7B --gpu-memory-utilization 0.45 --port 8001
Can vLLM Run GGUF Models or Run on a Mac?
Mostly no, for now. vLLM’s GGUF support is labelled “highly experimental and under-optimized” in its own docs, and it now needs a separate vllm-gguf-plugin package. If your models are GGUF files, llama.cpp or Ollama is the natural fit.
On Macs, vLLM’s installation docs describe experimental support for Apple Silicon that must be built from source and runs on the CPU. In practice, Apple Silicon users should pick Ollama or llama.cpp, both of which use Metal on the GPU.
Ollama vs vLLM vs llama.cpp: Which One Should You Use?
- One developer, local machine, just want it working: Ollama.
ollama runis the fastest path from nothing to a working model, and the hardware requirements guide tells you what fits. - Maximum control, unusual hardware, or embedding inference in your own tooling: llama.cpp. You pick the exact quantization, the context size, the slot count and the backend.
- A shared endpoint for a team or an application with real traffic, on NVIDIA or AMD GPUs: vLLM. Its batching pays off as soon as requests overlap.
- Small team sharing one Ollama box: Ollama can work, but set
OLLAMA_NUM_PARALLELabove 1 and test under load first.
A common path is to prototype with Ollama, then move the same model to vLLM when it becomes a service. All three expose an OpenAI-compatible API, so switching usually means changing a base URL and model name, not rewriting code.
Frequently Asked Questions
Q: Is Ollama built on llama.cpp?
A: Ollama runs GGUF models using the ggml/llama.cpp stack, and it has also added its own engine work, including MLX on Apple Silicon. In our benchmark on the same GGUF file, the differences came from defaults (context size and parallel slots), not from the engine itself.
Q: Is vLLM faster than Ollama?
A: For many concurrent requests on a GPU, usually yes, because vLLM batches requests continuously and manages KV-cache memory with PagedAttention. For a single user on a laptop, the difference is small, and vLLM is much harder to run on a Mac.
Q: Can I use vLLM with GGUF files?
A: Only experimentally. vLLM’s docs describe GGUF support as highly experimental and under-optimized, and it requires the separate vllm-gguf-plugin package. For GGUF models, use llama.cpp or Ollama.
Q: Why does Ollama slow down with multiple users?
A: By default Ollama processes one request per model at a time (OLLAMA_NUM_PARALLEL defaults to 1), so extra requests queue. Raise OLLAMA_NUM_PARALLEL, keeping in mind that each parallel slot needs its own context memory.
Quick Summary:
– Ollama vs vLLM vs llama.cpp is mostly a concurrency decision, not a single-request speed decision
– On the same GGUF file, default Ollama won for one user (18.7 vs 13.7 tok/s), llama.cpp won under 4 concurrent requests, and OLLAMA_NUM_PARALLEL=4 beat both
– Ollama defaults to OLLAMA_NUM_PARALLEL=1; llama.cpp auto-sizes parallel slots; vLLM batches continuously
– vLLM takes 92 percent of GPU memory by default and treats GGUF and Apple Silicon as experimental
– All three expose an OpenAI-compatible API, so moving between them is mostly a config change
If you’re choosing between Ollama vs vLLM vs llama.cpp today, start with Ollama on your laptop, measure with the concurrency you actually expect, and move to vLLM when the queue becomes the bottleneck.