The fastest way to run vLLM is the official Docker image: docker run --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface --env "HF_TOKEN=$HF_TOKEN" -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model Qwen/Qwen3-0.6B. To install it into your own Python environment, use uv pip install vllm --torch-backend=auto on Linux.
This guide walks through both routes, then covers multi-GPU serving with tensor parallelism and the quantization formats vLLM supports. Every command comes from the vLLM docs. If you want the theory (KV cache, attention backends, PagedAttention), read Serving LLMs with vLLM first; this post is the practical companion. It is part of our guide to LLM inference engines.
TL;DR
- Use Docker (
vllm/vllm-openai) on a rented GPU box: no CUDA or PyTorch version juggling. - Use
uv pip install vllm --torch-backend=autowhen you need vLLM inside your own environment. It needs Linux and Python 3.11 through 3.14. - For a model that does not fit one GPU, add
--tensor-parallel-size Nwith N equal to the GPUs in the node. Use pipeline parallelism only across nodes, or on GPUs without NVLink if the docs' guidance fits. - Quantize to fit bigger models: FP8 needs Ada or Hopper-class GPUs, AWQ and GPTQ work from Turing up, and Marlin kernels speed several formats.
- The latest stable release at the time of writing is v0.31.0 (5 October, per the GitHub releases page). Flags change between versions, so check
vllm serve --helpon your install.
Requirements
From the vLLM installation page:
- Operating system: Linux. Windows is not supported natively; the docs suggest WSL as a workaround.
- Python: 3.11 through 3.14, with 3.12 used in the docs' example environment.
- GPU: an NVIDIA GPU with a recent driver for the CUDA build. ROCm (AMD) and Intel XPU have their own images, shown below.
For Docker you also need the NVIDIA container toolkit on the host so that --runtime nvidia --gpus all can see the GPUs. Confirm with nvidia-smi before anything else.
Install vLLM with uv or pip
The docs recommend uv, which picks the matching PyTorch build for you.
uv pip install vllm --torch-backend=auto
With plain pip for CUDA 12.9:
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129
Start in a fresh virtual environment (for example uv venv --python 3.12, then activate it) so vLLM's pinned PyTorch does not collide with other projects. Once installed, serve a model with the CLI:
vllm serve Qwen/Qwen3-0.6B
That downloads the model from Hugging Face and starts an OpenAI-compatible server on port 8000. (The vllm serve form is the one the parallelism docs use throughout.)
Run vLLM with Docker
This is the command from the vLLM Docker page, verbatim:
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model Qwen/Qwen3-0.6B
What each part does:
--runtime nvidia --gpus allhands every GPU on the host to the container.-v ~/.cache/huggingface:/root/.cache/huggingfacereuses downloaded weights between runs, so a restart does not re-download.--env "HF_TOKEN=$HF_TOKEN"passes your Hugging Face token, needed for gated models. Set it in your host shell first.-p 8000:8000publishes the API port.--ipc=hostlets the container use host shared memory. The docs say you can use either--ipc=hostor--shm-size, because vLLM relies on PyTorch shared memory, especially for tensor-parallel inference.- Everything after the image name is passed to vLLM. Swap
--model Qwen/Qwen3-0.6Bfor the model you want.
Check that it works:
curl http://localhost:8000/v1/models
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [{"role": "user", "content": "Say hello in one sentence."}],
"max_tokens": 32
}'
Other accelerators and Podman
The same page documents variants. For AMD GPUs, use the ROCm image, with the tag of your choice in place of the placeholder:
docker run --rm \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
--device /dev/kfd \
--device /dev/dri \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai-rocm:<tag> \
--model Qwen/Qwen3-0.6B
For Podman on NVIDIA:
podman run --device nvidia.com/gpu=all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
docker.io/vllm/vllm-openai:latest \
--model Qwen/Qwen3-0.6B
The image names are vllm/vllm-openai (CUDA), vllm/vllm-openai-rocm (AMD, with latest and nightly tags) and vllm/vllm-openai-xpu (Intel XPU). AMD's older rocm/vllm and rocm/vllm-dev images are deprecated in favor of the official ROCm image, per the docs.
Build the image yourself
To build from the repository, the docs give:
DOCKER_BUILDKIT=1 docker build . \
--target vllm-openai \
--tag vllm/vllm-openai \
--file docker/Dockerfile
You can pass build arguments such as max_jobs and nvcc_threads to control parallelism, and an empty torch_cuda_arch_list builds only for your current GPU type. The page also covers Arm64 builds, precompiled wheels and a preview build for NVIDIA Rubin GPUs.
Pin the image tag
vllm/vllm-openai:latest moves. For a production endpoint, pin a version tag (for example one that matches the release you tested) so a restart does not pull a different engine. Record the tag together with the model revision and the flags, the same way you would pin a dependency.
Multi-GPU: tensor, pipeline and data parallelism
The parallelism page of the vLLM docs gives the rules. All commands are from it.
Tensor parallelism, for a model too large for one GPU but small enough for one node. Set the size to the number of GPUs in the node:
vllm serve facebook/opt-13b \
--tensor-parallel-size 4
With Docker, append the flag after the image name, for example vllm/vllm-openai:latest --model YOUR_MODEL --tensor-parallel-size 4. Tensor parallelism exchanges data between GPUs after every layer, so it works best on GPUs connected by NVLink. See our glossary entry on tensor parallelism.
Pipeline parallelism, for a model larger than one node, combined with tensor parallelism:
# Eight GPUs total
vllm serve gpt2 \
--tensor-parallel-size 4 \
--pipeline-parallel-size 2
For two nodes of eight GPUs each, the docs use Ray:
vllm serve /path/to/the/model/in/the/container \
--tensor-parallel-size 8 \
--pipeline-parallel-size 2 \
--distributed-executor-backend ray
The docs also suggest pipeline parallelism (with tensor parallel size 1) when the GPU count does not divide the model evenly, and recommend it over tensor parallelism for GPUs without NVLink, such as the L40S, for higher throughput and lower communication overhead.
Data parallelism: run independent replicas for more throughput. For Mixture of Experts models the docs describe data parallel attention combined with expert or tensor parallel MoE layers and point to a dedicated Data Parallel Deployment page for the commands.
Quantization
Quantization shrinks weights so a bigger model fits, or the same model leaves more room for KV cache. Weight size is parameters times bytes per parameter (computed here):
| Model size | FP16 / BF16 | FP8 / INT8 | INT4 |
|---|---|---|---|
| 8B | 16 GB | 8 GB | 4 GB |
| 32B | 64 GB | 32 GB | 16 GB |
| 70B | 140 GB | 70 GB | 35 GB |
These are weights only; KV cache and runtime overhead come on top. Our glossary covers quantization, AWQ, GPTQ and FP8.
The quantization page of the vLLM docs lists these supported methods: AutoAWQ, BitsAndBytes, GPTQModel, Intel Neural Compressor, LLM Compressor (FP8 W8A8, INT4 W4A16, INT8 W4A8, INT8 W8A8), NVIDIA Model Optimizer, online quantization, AMD Quark, quantized KV cache, TorchAO and FP8 ViT encoder attention. Its hardware table says, in summary:
| Method | NVIDIA support per the vLLM docs |
|---|---|
| AWQ | Turing to Hopper only |
| GPTQ | Yes |
| Marlin (GPTQ, AWQ, FP8, FP4) | Turing (partial) to Hopper |
| LLM Compressor INT8 W8A8 | Turing to Hopper |
| LLM Compressor FP8 W8A8 | Ada and Hopper (also AMD GPUs) |
| bitsandbytes | Yes |
| GGUF | Yes (also AMD GPUs) |
The practical rule: use a pre-quantized checkpoint from the model authors or a trusted publisher, point --model at it, and let vLLM read the quantization config. On Hopper and Ada cards, FP8 is usually the first thing to try; on older Ampere cards, AWQ or GPTQ are the common choices. The docs state that Turing does not support Marlin MXFP4. Quantization can cost some quality, so rerun your own evaluation after switching.
Useful serving flags
Three flags handle most memory problems (all covered in the vLLM docs and in our vLLM guide):
--max-model-lencaps context length and therefore KV cache per request.--gpu-memory-utilizationsets the share of each GPU vLLM may claim.--tensor-parallel-sizesplits the model across GPUs.
If the server runs out of memory at startup, lower --max-model-len first, then try a quantized checkpoint, then add GPUs. For the sizing math, use the VRAM calculator and how much VRAM you need for LLMs.
Which engine next?
If vLLM does not fit your traffic, compare it with the alternatives: SGLang for prefix-heavy and agent workloads, TensorRT-LLM for NVIDIA-specific optimization, and vLLM vs TensorRT-LLM vs SGLang for the three-way comparison.
Run it on a cloud GPU
Rent a box, check nvidia-smi, install Docker with the NVIDIA container toolkit, and run the command above. The live offers for the cards in this guide:
FAQ
How do I install vLLM?
On Linux with Python 3.11 through 3.14, run uv pip install vllm --torch-backend=auto, or pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129 for CUDA 12.9. Windows is not supported natively.
What is the vLLM Docker image?
vllm/vllm-openai on Docker Hub is the official CUDA image with an OpenAI-compatible server. AMD uses vllm/vllm-openai-rocm and Intel XPU uses vllm/vllm-openai-xpu.
Why does the docs' Docker command use --ipc=host?
vLLM relies on PyTorch shared memory, especially for tensor-parallel inference. The docs say to use either --ipc=host or --shm-size so the container can access enough host shared memory.
How do I run vLLM on multiple GPUs?
Set --tensor-parallel-size to the number of GPUs in the node, for example vllm serve facebook/opt-13b --tensor-parallel-size 4. For multi-node setups combine it with --pipeline-parallel-size and the Ray executor.
Which quantization formats does vLLM support?
AWQ, GPTQ, FP8 and INT8 through LLM Compressor, bitsandbytes, GGUF, TorchAO, NVIDIA Model Optimizer and others. Support depends on GPU generation: FP8 W8A8 needs Ada or Hopper on NVIDIA, for instance.
Can vLLM run on AMD GPUs?
Yes, through the vllm/vllm-openai-rocm image. The docs give a ROCm-specific run command with /dev/kfd and /dev/dri devices.
Sources
- vLLM Docker deployment: https://docs.vllm.ai/en/latest/deployment/docker/
- vLLM GPU installation: https://docs.vllm.ai/en/latest/getting_started/installation/gpu/
- vLLM parallelism and scaling: https://docs.vllm.ai/en/latest/serving/parallelism_scaling/
- vLLM quantization: https://docs.vllm.ai/en/latest/features/quantization/
- vLLM releases (v0.31.0): https://github.com/vllm-project/vllm/releases
- Weight-size table: computed here as parameters times bytes per parameter.