Use Ollama when one person or a small team needs a model running quickly on a laptop or a single GPU. Use vLLM when many users hit the same model at once and you care about aggregate throughput on server GPUs.
The two tools solve different problems. Ollama optimizes for how fast you get a model running; vLLM optimizes for how many requests a GPU can serve per second. This post is part of our guide to LLM inference engines.
TL;DR
- Ollama: a model manager plus local server built on llama.cpp. Default is one parallel request per loaded model. Runs on NVIDIA, AMD and Apple hardware.
- vLLM: a serving engine built around PagedAttention and batching of concurrent requests. The quickstart targets Linux with NVIDIA CUDA, AMD ROCm and other accelerators.
- Verdict for a prototype or desktop app: Ollama.
- Verdict for a production endpoint with concurrent users: vLLM.
- Verdict for multi-GPU large models: vLLM, which has first-class tensor and pipeline parallelism.
What each tool is
Ollama downloads prepackaged, quantized models and serves them locally on port 11434. Its README names llama.cpp as the backend, and its install is one line. We cover it in what is Ollama.
vLLM is an open-source library and server for high-throughput LLM serving. Its research paper introduced PagedAttention, which manages the attention KV cache in blocks to cut wasted memory, and the abstract reports that vLLM "improves the throughput of popular LLMs by 2-4x with the same level of latency" compared with the FasterTransformer and Orca systems it was measured against (Kwon et al., 2023). That is the authors' number on their benchmark workloads, not a comparison with Ollama. Our vLLM serving guide explains the mechanism end to end.
Concurrency: the main difference
Ollama's FAQ gives its defaults:
OLLAMA_NUM_PARALLELdefaults to 1, meaning each loaded model handles one request at a time unless you raise it. Memory use scales with this value times the context length.OLLAMA_MAX_QUEUEdefaults to 512, and requests past that are rejected with a 503.OLLAMA_MAX_LOADED_MODELSdefaults to 3 times the number of GPUs.
You can raise OLLAMA_NUM_PARALLEL, and for a few users that works. But its design target is a person or a small group, with extra requests queued.
vLLM is built for the opposite case. Its feature list in the docs includes automatic prefix caching, speculative decoding, LoRA serving and chunked prefill, and its scheduling, described in our serving guide, batches requests together so many concurrent users share one GPU's memory bandwidth. This is why the usual advice is to move to vLLM once concurrent traffic, not model size, becomes your bottleneck.
Side-by-side comparison
| Ollama | vLLM | |
|---|---|---|
| Built for | Local use, development, small teams | High-concurrency serving |
| Backend | llama.cpp | Own engine with PyTorch |
| Typical model format | GGUF (quantized) | Hugging Face checkpoints, with AWQ, GPTQ, FP8, INT8 and BitsAndBytes quantization |
| Default parallel requests | 1 per loaded model | Batched by the scheduler |
| Multi-GPU | Spreads a model across GPUs only when it does not fit on one | Tensor and pipeline parallelism flags |
| Platforms | macOS, Windows, Linux, Docker | Linux (CUDA, ROCm, others); Apple Silicon through vLLM-Metal |
| API | Native plus OpenAI-compatible, port 11434 | OpenAI-compatible, port 8000 |
| Latest release (October 2026) | v0.40.1 | v0.31.0 |
Quantization support for vLLM is from its quantization docs, which list AWQ, GPTQ, FP8, INT8 and BitsAndBytes among others, and mark GGUF as supported on Volta through Hopper and on AMD GPUs but not on Intel GPUs or CPUs. For production vLLM, prefer a checkpoint quantized in a native format over GGUF.
Install and first request
Ollama:
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4
vLLM, from its quickstart:
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
vllm serve Qwen/Qwen2.5-1.5B-Instruct
The server listens on http://localhost:8000. A completions request from the docs:
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"prompt": "San Francisco is a",
"max_tokens": 7,
"temperature": 0
}'
The vLLM docs list Linux as the operating system and Python 3.11 to 3.14. With --torch-backend=auto, uv picks the PyTorch build that matches your installed CUDA driver. For containers, the docs give:
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model Qwen/Qwen3-0.6B
Because both tools expose an OpenAI-compatible API, switching is mostly a base-URL change: Ollama at http://localhost:11434/v1/, vLLM at http://localhost:8000/v1.
Multi-GPU
The two tools handle multiple GPUs very differently.
Ollama loads a model onto one GPU if it fits there, and spreads it across all available GPUs only if it does not (Ollama FAQ). That helps you fit a bigger model but is not designed to scale throughput.
vLLM's parallelism docs give explicit guidance:
- Tensor parallelism: use
tensor_parallel_sizewhen the model is too large for one GPU but fits on one node. Set it to the number of GPUs, for example 4 on a 4-GPU node. - Pipeline parallelism: when the model is too big for one node, combine tensor parallelism per node with
pipeline_parallel_sizeacross nodes. - No NVLink: the docs say pipeline parallelism offers higher throughput and lower communication overhead than tensor parallelism on GPUs without NVLink, such as the L40S.
For example, a 70B model at FP16 needs about 140 GB for weights alone (computed: 70 billion parameters times 2 bytes), so it needs at least two 80 GB cards before any KV cache. See tensor parallelism for the concept.
GPU and VRAM needs
Both tools are limited by memory first. Weights take roughly parameters times bytes per parameter: 2 bytes at FP16, 1 at INT8, about 0.5 at INT4 (computed, weights only). An 8B model is about 16 GB, 8 GB or 4 GB respectively, and the KV cache comes on top. Ollama's 4-bit GGUF files make it easy to squeeze a model onto a 24 GB consumer card such as the RTX 4090. vLLM gives you more headroom per concurrent user through PagedAttention, but it expects a server-class setup, such as the H100 or L40S, when traffic is real. The VRAM calculator does the exact math, and best GPU for LLM inference helps pick hardware.
When to pick which
Pick Ollama when:
- you are prototyping, building a local agent, or running a desktop tool;
- you are on a Mac or a Windows machine;
- you want a model running in minutes with prepackaged weights.
Pick vLLM when:
- many users or services send requests at once;
- you need tensor or pipeline parallelism across several GPUs;
- you serve LoRA adapters, long shared prompts (prefix caching) or other features in its docs;
- you are deploying an endpoint that needs to stay up and scale.
A common path is to prototype on Ollama and move the same OpenAI-style client code to vLLM when traffic grows. For other server engines, see vLLM vs TensorRT-LLM vs SGLang, and for the lower-level engine under Ollama, the llama.cpp guide.
Performance
Neither project publishes a direct comparison against the other, and we do not make one up. The vLLM paper's 2-4x figure compares it with other serving systems, not Ollama. If you need numbers, benchmark both on your hardware with the same model, the same prompts and the concurrency you expect, and record time to first token and aggregate tokens per second at that concurrency. Single-user speed on one GPU can look similar; the gap tends to appear as concurrency rises.
Run it on a cloud GPU
To try vLLM against Ollama on the same hardware, start both on one rented NVIDIA GPU and point the same client at each port. The boxes below show what is available right now.
FAQ
Is vLLM faster than Ollama?
For many concurrent requests, vLLM is designed to deliver higher aggregate throughput. For a single user on a single GPU the difference can be small. No neutral published benchmark compares the two directly.
Can I run vLLM on a Mac or Windows?
The vLLM quickstart lists Linux as the OS. Apple Silicon is covered through the separate vLLM-Metal project, which uses MLX models. Ollama runs natively on macOS and Windows.
Does Ollama support multiple GPUs?
Yes, but it loads a model onto a single GPU if it fits and only spreads it across GPUs when it does not.
Can vLLM run GGUF models?
The quantization docs mark GGUF as supported on Volta through Hopper and on AMD GPUs, not on Intel GPUs or CPUs. Native formats such as AWQ, GPTQ and FP8 are the usual choice for production.
Do both expose an OpenAI-compatible API?
Yes. Ollama serves it at http://localhost:11434/v1/ and vLLM at http://localhost:8000/v1.
Which is easier to install?
Ollama, with a one-line installer. vLLM needs a Python environment and a compatible CUDA setup, or its Docker image.
Sources
- vLLM quickstart (install, serve, requirements, hardware): https://docs.vllm.ai/en/latest/getting_started/quickstart/
- vLLM releases (v0.31.0 latest as of October 8, 2026): https://github.com/vllm-project/vllm/releases
- vLLM parallelism and scaling: https://docs.vllm.ai/en/latest/serving/parallelism_scaling/
- vLLM features: https://docs.vllm.ai/en/latest/features/
- vLLM quantization: https://docs.vllm.ai/en/latest/features/quantization/
- vLLM Docker deployment: https://docs.vllm.ai/en/latest/deployment/docker/
- PagedAttention paper, Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention": https://arxiv.org/abs/2309.06180
- Ollama README: https://github.com/ollama/ollama
- Ollama releases (v0.40.1): https://github.com/ollama/ollama/releases
- Ollama FAQ (concurrency, multi-GPU): https://docs.ollama.com/faq
- Ollama OpenAI compatibility: https://docs.ollama.com/api/openai-compatibility
Related reading: what is Ollama, Ollama vs LM Studio, what is AI inference.