vLLM Docker and Install Guide: Multi-GPU, Quantization

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The fastest way to run vLLM is the official Docker image: docker run --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface --env "HF_TOKEN=$HF_TOKEN" -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model Qwen/Qwen3-0.6B. To install it into your own Python environment, use uv pip install vllm --torch-backend=auto on Linux.

This guide walks through both routes, then covers multi-GPU serving with tensor parallelism and the quantization formats vLLM supports. Every command comes from the vLLM docs. If you want the theory (KV cache, attention backends, PagedAttention), read Serving LLMs with vLLM first; this post is the practical companion. It is part of our guide to LLM inference engines.

TL;DR

  • Use Docker (vllm/vllm-openai) on a rented GPU box: no CUDA or PyTorch version juggling.
  • Use uv pip install vllm --torch-backend=auto when you need vLLM inside your own environment. It needs Linux and Python 3.11 through 3.14.
  • For a model that does not fit one GPU, add --tensor-parallel-size N with N equal to the GPUs in the node. Use pipeline parallelism only across nodes, or on GPUs without NVLink if the docs' guidance fits.
  • Quantize to fit bigger models: FP8 needs Ada or Hopper-class GPUs, AWQ and GPTQ work from Turing up, and Marlin kernels speed several formats.
  • The latest stable release at the time of writing is v0.31.0 (5 October, per the GitHub releases page). Flags change between versions, so check vllm serve --help on your install.

Requirements

From the vLLM installation page:

  • Operating system: Linux. Windows is not supported natively; the docs suggest WSL as a workaround.
  • Python: 3.11 through 3.14, with 3.12 used in the docs' example environment.
  • GPU: an NVIDIA GPU with a recent driver for the CUDA build. ROCm (AMD) and Intel XPU have their own images, shown below.

For Docker you also need the NVIDIA container toolkit on the host so that --runtime nvidia --gpus all can see the GPUs. Confirm with nvidia-smi before anything else.

Install vLLM with uv or pip

The docs recommend uv, which picks the matching PyTorch build for you.

uv pip install vllm --torch-backend=auto

With plain pip for CUDA 12.9:

pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129

Start in a fresh virtual environment (for example uv venv --python 3.12, then activate it) so vLLM's pinned PyTorch does not collide with other projects. Once installed, serve a model with the CLI:

vllm serve Qwen/Qwen3-0.6B

That downloads the model from Hugging Face and starts an OpenAI-compatible server on port 8000. (The vllm serve form is the one the parallelism docs use throughout.)

Run vLLM with Docker

This is the command from the vLLM Docker page, verbatim:

docker run --runtime nvidia --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=$HF_TOKEN" \
    -p 8000:8000 \
    --ipc=host \
    vllm/vllm-openai:latest \
    --model Qwen/Qwen3-0.6B

What each part does:

  • --runtime nvidia --gpus all hands every GPU on the host to the container.
  • -v ~/.cache/huggingface:/root/.cache/huggingface reuses downloaded weights between runs, so a restart does not re-download.
  • --env "HF_TOKEN=$HF_TOKEN" passes your Hugging Face token, needed for gated models. Set it in your host shell first.
  • -p 8000:8000 publishes the API port.
  • --ipc=host lets the container use host shared memory. The docs say you can use either --ipc=host or --shm-size, because vLLM relies on PyTorch shared memory, especially for tensor-parallel inference.
  • Everything after the image name is passed to vLLM. Swap --model Qwen/Qwen3-0.6B for the model you want.

Check that it works:

curl http://localhost:8000/v1/models
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-0.6B",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}],
    "max_tokens": 32
  }'

Other accelerators and Podman

The same page documents variants. For AMD GPUs, use the ROCm image, with the tag of your choice in place of the placeholder:

docker run --rm \
    --group-add=video \
    --cap-add=SYS_PTRACE \
    --security-opt seccomp=unconfined \
    --device /dev/kfd \
    --device /dev/dri \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=$HF_TOKEN" \
    -p 8000:8000 \
    --ipc=host \
    vllm/vllm-openai-rocm:<tag> \
    --model Qwen/Qwen3-0.6B

For Podman on NVIDIA:

podman run --device nvidia.com/gpu=all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
docker.io/vllm/vllm-openai:latest \
--model Qwen/Qwen3-0.6B

The image names are vllm/vllm-openai (CUDA), vllm/vllm-openai-rocm (AMD, with latest and nightly tags) and vllm/vllm-openai-xpu (Intel XPU). AMD's older rocm/vllm and rocm/vllm-dev images are deprecated in favor of the official ROCm image, per the docs.

Build the image yourself

To build from the repository, the docs give:

DOCKER_BUILDKIT=1 docker build . \
    --target vllm-openai \
    --tag vllm/vllm-openai \
    --file docker/Dockerfile

You can pass build arguments such as max_jobs and nvcc_threads to control parallelism, and an empty torch_cuda_arch_list builds only for your current GPU type. The page also covers Arm64 builds, precompiled wheels and a preview build for NVIDIA Rubin GPUs.

Pin the image tag

vllm/vllm-openai:latest moves. For a production endpoint, pin a version tag (for example one that matches the release you tested) so a restart does not pull a different engine. Record the tag together with the model revision and the flags, the same way you would pin a dependency.

Multi-GPU: tensor, pipeline and data parallelism

The parallelism page of the vLLM docs gives the rules. All commands are from it.

Tensor parallelism, for a model too large for one GPU but small enough for one node. Set the size to the number of GPUs in the node:

vllm serve facebook/opt-13b \
     --tensor-parallel-size 4

With Docker, append the flag after the image name, for example vllm/vllm-openai:latest --model YOUR_MODEL --tensor-parallel-size 4. Tensor parallelism exchanges data between GPUs after every layer, so it works best on GPUs connected by NVLink. See our glossary entry on tensor parallelism.

Pipeline parallelism, for a model larger than one node, combined with tensor parallelism:

# Eight GPUs total
vllm serve gpt2 \
     --tensor-parallel-size 4 \
     --pipeline-parallel-size 2

For two nodes of eight GPUs each, the docs use Ray:

vllm serve /path/to/the/model/in/the/container \
    --tensor-parallel-size 8 \
    --pipeline-parallel-size 2 \
    --distributed-executor-backend ray

The docs also suggest pipeline parallelism (with tensor parallel size 1) when the GPU count does not divide the model evenly, and recommend it over tensor parallelism for GPUs without NVLink, such as the L40S, for higher throughput and lower communication overhead.

Data parallelism: run independent replicas for more throughput. For Mixture of Experts models the docs describe data parallel attention combined with expert or tensor parallel MoE layers and point to a dedicated Data Parallel Deployment page for the commands.

Quantization

Quantization shrinks weights so a bigger model fits, or the same model leaves more room for KV cache. Weight size is parameters times bytes per parameter (computed here):

Model sizeFP16 / BF16FP8 / INT8INT4
8B16 GB8 GB4 GB
32B64 GB32 GB16 GB
70B140 GB70 GB35 GB

These are weights only; KV cache and runtime overhead come on top. Our glossary covers quantization, AWQ, GPTQ and FP8.

The quantization page of the vLLM docs lists these supported methods: AutoAWQ, BitsAndBytes, GPTQModel, Intel Neural Compressor, LLM Compressor (FP8 W8A8, INT4 W4A16, INT8 W4A8, INT8 W8A8), NVIDIA Model Optimizer, online quantization, AMD Quark, quantized KV cache, TorchAO and FP8 ViT encoder attention. Its hardware table says, in summary:

MethodNVIDIA support per the vLLM docs
AWQTuring to Hopper only
GPTQYes
Marlin (GPTQ, AWQ, FP8, FP4)Turing (partial) to Hopper
LLM Compressor INT8 W8A8Turing to Hopper
LLM Compressor FP8 W8A8Ada and Hopper (also AMD GPUs)
bitsandbytesYes
GGUFYes (also AMD GPUs)

The practical rule: use a pre-quantized checkpoint from the model authors or a trusted publisher, point --model at it, and let vLLM read the quantization config. On Hopper and Ada cards, FP8 is usually the first thing to try; on older Ampere cards, AWQ or GPTQ are the common choices. The docs state that Turing does not support Marlin MXFP4. Quantization can cost some quality, so rerun your own evaluation after switching.

Useful serving flags

Three flags handle most memory problems (all covered in the vLLM docs and in our vLLM guide):

  • --max-model-len caps context length and therefore KV cache per request.
  • --gpu-memory-utilization sets the share of each GPU vLLM may claim.
  • --tensor-parallel-size splits the model across GPUs.

If the server runs out of memory at startup, lower --max-model-len first, then try a quantized checkpoint, then add GPUs. For the sizing math, use the VRAM calculator and how much VRAM you need for LLMs.

Which engine next?

If vLLM does not fit your traffic, compare it with the alternatives: SGLang for prefix-heavy and agent workloads, TensorRT-LLM for NVIDIA-specific optimization, and vLLM vs TensorRT-LLM vs SGLang for the three-way comparison.

Run it on a cloud GPU

Rent a box, check nvidia-smi, install Docker with the NVIDIA container toolkit, and run the command above. The live offers for the cards in this guide:

FAQ

How do I install vLLM?

On Linux with Python 3.11 through 3.14, run uv pip install vllm --torch-backend=auto, or pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129 for CUDA 12.9. Windows is not supported natively.

What is the vLLM Docker image?

vllm/vllm-openai on Docker Hub is the official CUDA image with an OpenAI-compatible server. AMD uses vllm/vllm-openai-rocm and Intel XPU uses vllm/vllm-openai-xpu.

Why does the docs' Docker command use --ipc=host?

vLLM relies on PyTorch shared memory, especially for tensor-parallel inference. The docs say to use either --ipc=host or --shm-size so the container can access enough host shared memory.

How do I run vLLM on multiple GPUs?

Set --tensor-parallel-size to the number of GPUs in the node, for example vllm serve facebook/opt-13b --tensor-parallel-size 4. For multi-node setups combine it with --pipeline-parallel-size and the Ray executor.

Which quantization formats does vLLM support?

AWQ, GPTQ, FP8 and INT8 through LLM Compressor, bitsandbytes, GGUF, TorchAO, NVIDIA Model Optimizer and others. Support depends on GPU generation: FP8 W8A8 needs Ada or Hopper on NVIDIA, for instance.

Can vLLM run on AMD GPUs?

Yes, through the vllm/vllm-openai-rocm image. The docs give a ROCm-specific run command with /dev/kfd and /dev/dri devices.

Sources

#llm inference#inference engines#vllm#docker#tensor parallelism#quantization

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.