How to run Llama 3.2 with vLLM and llama.cpp
Engine support, Transformers version, vision input, tool-call template and EU multimodal license note for running Llama 3.2 text and vision models.
This guide covers serving the Llama 3.2 family: small text models (1B) and vision models (11B Vision shown here).
Engine support
- The 11B Vision card requires Transformers 4.45.0 or newer for conversational inference with images, and says it is compatible with vLLM.
- The 1B card lists Transformers 4.43.0 or newer, and notes quantized versions through llama.cpp, Ollama, LM Studio and Jan (see GGUF).
Vision input
Pass images inside chat messages and format them with processor.apply_chat_template() using add_generation_prompt=True.
Tool calling
vLLM documents the llama3_json parser for all Llama 3.1 and 3.2 models. Use the 3.2 template (tool_chat_template_llama3.2_json.jinja), which adds image support. vLLM says parallel tool calls are not supported for Llama 3.
Context and languages
Both cards list a 128k token context. The 1B card names English, German, French, Italian, Portuguese, Hindi, Spanish and Thai as officially supported.
EU restriction on multimodal use
The vision card states that the Section 1(a) rights are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union for the multimodal models. Check this before deploying the vision sizes.
Launch commands by Llama 3.2 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Llama-3.2-1B-Instruct (1.2B)
Needs about 2.8 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve meta-llama/Llama-3.2-1B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path meta-llama/Llama-3.2-1B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Llama-3.2-3B-Instruct (3.2B)
Needs about 7.2 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve meta-llama/Llama-3.2-3B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path meta-llama/Llama-3.2-3B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf bartowski/Llama-3.2-3B-Instruct-GGUFPrebuilt GGUF weights published at bartowski/Llama-3.2-3B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/meta-llama/Llama-3.2-11B-Vision-Instruct
- https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct
- https://docs.vllm.ai/en/latest/features/tool_calling.html
Updated 2026-10-07.