How to run Llama 3.2 with vLLM and llama.cpp

Engine support, Transformers version, vision input, tool-call template and EU multimodal license note for running Llama 3.2 text and vision models.

This guide covers serving the Llama 3.2 family: small text models (1B) and vision models (11B Vision shown here).

Engine support

  • The 11B Vision card requires Transformers 4.45.0 or newer for conversational inference with images, and says it is compatible with vLLM.
  • The 1B card lists Transformers 4.43.0 or newer, and notes quantized versions through llama.cpp, Ollama, LM Studio and Jan (see GGUF).

Vision input

Pass images inside chat messages and format them with processor.apply_chat_template() using add_generation_prompt=True.

Tool calling

vLLM documents the llama3_json parser for all Llama 3.1 and 3.2 models. Use the 3.2 template (tool_chat_template_llama3.2_json.jinja), which adds image support. vLLM says parallel tool calls are not supported for Llama 3.

Context and languages

Both cards list a 128k token context. The 1B card names English, German, French, Italian, Portuguese, Hindi, Spanish and Thai as officially supported.

EU restriction on multimodal use

The vision card states that the Section 1(a) rights are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union for the multimodal models. Check this before deploying the vision sizes.

Launch commands by Llama 3.2 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Llama-3.2-1B-Instruct (1.2B)

Needs about 2.8 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve meta-llama/Llama-3.2-1B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path meta-llama/Llama-3.2-1B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Llama-3.2-3B-Instruct (3.2B)

Needs about 7.2 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve meta-llama/Llama-3.2-3B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path meta-llama/Llama-3.2-3B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run llama3.2:3b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf bartowski/Llama-3.2-3B-Instruct-GGUF

Prebuilt GGUF weights published at bartowski/Llama-3.2-3B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.