Ollama vs vLLM vs llama.cpp: Which Local LLM Engine?

ollama vs vllm vs llama.cpp comparison of local LLM engines and concurrency

Ollama vs vLLM vs llama.cpp comes down to one question: how many people will hit the model at once? For one developer on a laptop, pick Ollama or llama.cpp. For a GPU server with many concurrent users, pick vLLM. Below, I explain why. I also benchmarked Ollama and llama.cpp on the same model file, and the defaults changed the result more than the engine did.

If you haven’t installed any of these yet, start with our LM Studio and Ollama setup guide.

Ollama vs vLLM vs llama.cpp at a Glance

Ollama llama.cpp vLLM
What it is Model manager and local server C/C++ inference engine, CLI and server High-throughput serving engine
Best at Easiest local setup Running anywhere, fine control Many concurrent requests on GPUs
Model format Its own library, or import GGUF / Safetensors GGUF Hugging Face weights; GGUF “highly experimental” via a plugin
Hardware macOS, Windows, Linux; CPU, NVIDIA, AMD, Apple Silicon CPU, Metal, CUDA, HIP, Vulkan, SYCL and more NVIDIA, AMD and Intel GPUs, TPUs, CPUs; experimental on Apple Silicon
API Native API plus OpenAI-compatible OpenAI-compatible (llama serve) OpenAI-compatible, Anthropic Messages API, gRPC
Concurrency default 1 request per model Parallel slots, auto-sized Continuous batching
Setup effort Lowest Low to medium Highest

What Each Engine Actually Is

llama.cpp is the engine underneath much of the local-LLM world. It’s a plain C/C++ implementation with no dependencies. Per its README, it targets “a wide range of hardware.” It runs GGUF files and supports 1.5-bit to 8-bit quantization. As of v0.5.0 it ships a single llama command, so the old llama-server workflow is now llama serve:

# Chat with a model straight from Hugging Face
llama cli -hf ggml-org/Qwen3-4B-GGUF

# OpenAI-compatible server on port 8080
llama serve -m ./Qwen3-4B-Q4_K_M.gguf --port 8080

Ollama wraps local inference in a model manager: ollama pull, ollama run, a model library, and a background service. It exposes both its own API and an OpenAI-compatible one on port 11434. It trades some control for convenience, which is why most people start here. Our Ollama API guide covers the endpoints.

vLLM is a serving engine built for throughput. Its headline features are PagedAttention for KV-cache memory, continuous batching, chunked prefill and prefix caching. It also splits models across GPUs with tensor, pipeline and expert parallelism. So it’s what you run behind an internal API that dozens of users or agents call at once.

uv pip install vllm
vllm serve Qwen/Qwen3-4B

Ollama vs llama.cpp: A Real Benchmark on the Same Model File

To compare Ollama and llama.cpp fairly, I loaded the exact same GGUF file into both: Qwen3-4B at Q4_K_M, 2.5 GB. The machine was an Apple M1 with 16 GB of RAM. I ran Ollama 0.34.4 and llama.cpp 0.5.0 with default settings. Each engine generated 256 tokens for the same prompt. First I sent one request at a time, then four at once.

Setup Single request 4 concurrent (aggregate) Wall time for 4 requests
Ollama, default (OLLAMA_NUM_PARALLEL=1) 18.7 tok/s 18.9 tok/s 54 s
llama.cpp llama serve, default (4 slots) 13.7 tok/s 28.9 tok/s 35 s
Ollama with OLLAMA_NUM_PARALLEL=4 17.8 tok/s 35.0 tok/s 29 s

Three things stand out.

For one user, Ollama was faster. It generated about 18.7 tokens per second against 13.7 for llama.cpp on the same file. The startup logs show different defaults: Ollama picked a 4,096-token context from available VRAM, while llama serve auto-configured 4 slots with a 40,960-token context each. Same weights, different setup.

Under load, Ollama’s default queues. With four requests at once, default Ollama served them one after another. Aggregate throughput stayed at about 19 tokens per second, and the last user waited nearly a minute. llama.cpp’s parallel slots split the work, so each request ran slower (8 tokens per second) but all four finished in 35 seconds.

One setting flips the result. With OLLAMA_NUM_PARALLEL=4, Ollama handled the same four requests in 29 seconds at 35 tokens per second aggregate, the best result in the test.

I also threw away an earlier run. It showed llama.cpp far ahead, but the Mac was busy (load average around 9), and the numbers didn’t hold up on a rerun. Benchmark on an idle machine, and run each test more than once.

⚠️ Note: These are small-model numbers on one laptop. They show the shape of the difference, not a universal ranking. Run the same test on your own hardware before choosing.

Ollama vs vLLM vs llama.cpp Under Concurrent Load

For a single user, the gap between engines comes mostly from configuration: context size, slot count and quantization. Model size and memory bandwidth set the ceiling, and Ollama and llama.cpp share much of the same low-level code.

Concurrency is where they split:

  • Ollama processes one request per model at a time by default. Its FAQ lists OLLAMA_NUM_PARALLEL with a default of 1, and notes that memory use scales with OLLAMA_NUM_PARALLEL × OLLAMA_CONTEXT_LENGTH. Extra requests wait in a queue, and past OLLAMA_MAX_QUEUE the server returns a 503.
  • llama.cpp runs several parallel “slots” in llama serve. The -np / --parallel flag defaults to auto.
  • vLLM batches requests continuously, adding new ones to the running batch as others finish. Combined with PagedAttention, this is what lets one GPU serve many users without the queue growing.

If Ollama is your backend and several people or agents share it, raise the parallel setting before blaming the model:

OLLAMA_NUM_PARALLEL=4 ollama serve

Each parallel slot needs its own context memory, so watch RAM or VRAM when you raise it.

The vLLM Memory Default That Surprises People

vLLM claims most of the GPU as soon as it starts. The --gpu-memory-utilization flag defaults to 0.92, meaning vLLM takes 92 percent of GPU memory for the model and its KV cache, per the engine arguments reference. That’s intentional: the spare memory becomes KV cache, which is what makes high concurrency possible.

It also means a second vLLM instance, or anything else on that GPU, fails to allocate. Run two instances on one GPU and you have to split it yourself:

vllm serve Qwen/Qwen3-4B --gpu-memory-utilization 0.45 --port 8000
vllm serve Qwen/Qwen3-1.7B --gpu-memory-utilization 0.45 --port 8001

Can vLLM Run GGUF Models or Run on a Mac?

Mostly no, for now. vLLM’s GGUF support is labelled “highly experimental and under-optimized” in its own docs, and it now needs a separate vllm-gguf-plugin package. If your models are GGUF files, llama.cpp or Ollama is the natural fit.

On Macs, vLLM’s installation docs describe experimental support for Apple Silicon that must be built from source and runs on the CPU. In practice, Apple Silicon users should pick Ollama or llama.cpp, both of which use Metal on the GPU.

Ollama vs vLLM vs llama.cpp: Which One Should You Use?

  • One developer, local machine, just want it working: Ollama. ollama run is the fastest path from nothing to a working model, and the hardware requirements guide tells you what fits.
  • Maximum control, unusual hardware, or embedding inference in your own tooling: llama.cpp. You pick the exact quantization, the context size, the slot count and the backend.
  • A shared endpoint for a team or an application with real traffic, on NVIDIA or AMD GPUs: vLLM. Its batching pays off as soon as requests overlap.
  • Small team sharing one Ollama box: Ollama can work, but set OLLAMA_NUM_PARALLEL above 1 and test under load first.

A common path is to prototype with Ollama, then move the same model to vLLM when it becomes a service. All three expose an OpenAI-compatible API, so switching usually means changing a base URL and model name, not rewriting code.

Frequently Asked Questions

Q: Is Ollama built on llama.cpp?
A: Ollama runs GGUF models using the ggml/llama.cpp stack, and it has also added its own engine work, including MLX on Apple Silicon. In our benchmark on the same GGUF file, the differences came from defaults (context size and parallel slots), not from the engine itself.

Q: Is vLLM faster than Ollama?
A: For many concurrent requests on a GPU, usually yes, because vLLM batches requests continuously and manages KV-cache memory with PagedAttention. For a single user on a laptop, the difference is small, and vLLM is much harder to run on a Mac.

Q: Can I use vLLM with GGUF files?
A: Only experimentally. vLLM’s docs describe GGUF support as highly experimental and under-optimized, and it requires the separate vllm-gguf-plugin package. For GGUF models, use llama.cpp or Ollama.

Q: Why does Ollama slow down with multiple users?
A: By default Ollama processes one request per model at a time (OLLAMA_NUM_PARALLEL defaults to 1), so extra requests queue. Raise OLLAMA_NUM_PARALLEL, keeping in mind that each parallel slot needs its own context memory.

Quick Summary:
– Ollama vs vLLM vs llama.cpp is mostly a concurrency decision, not a single-request speed decision
– On the same GGUF file, default Ollama won for one user (18.7 vs 13.7 tok/s), llama.cpp won under 4 concurrent requests, and OLLAMA_NUM_PARALLEL=4 beat both
– Ollama defaults to OLLAMA_NUM_PARALLEL=1; llama.cpp auto-sizes parallel slots; vLLM batches continuously
– vLLM takes 92 percent of GPU memory by default and treats GGUF and Apple Silicon as experimental
– All three expose an OpenAI-compatible API, so moving between them is mostly a config change

If you’re choosing between Ollama vs vLLM vs llama.cpp today, start with Ollama on your laptop, measure with the concurrency you actually expect, and move to vLLM when the queue becomes the bottleneck.

Related guides

Leave a Reply