How to run Llama 3.3 with vLLM and Transformers
Engine support, Transformers version, tool-call parser, context length and quantization notes for running Llama 3.3 70B Instruct.
This guide covers serving the Llama 3.3 family, which on the Hugging Face card is a single 70B Instruct model. Sizing and launch commands appear below.
Engine support
- The card documents Transformers 4.45.0 or newer, including conversational use with the chat template.
- It documents vLLM through
vllm servewith an OpenAI-compatible API. - The card does not discuss SGLang, llama.cpp or Ollama.
Chat template and tools
The model supports chat templates and several tool formats through Transformers. In vLLM, tool calling uses the llama3_json parser with a chat template for the 3.1 or 3.2 generation (--tool-call-parser, --chat-template). vLLM notes that parallel tool calls are not supported for Llama 3, and that the model may serialize arrays as strings in tool arguments.
Context and languages
- Context length is 128k tokens.
- Officially supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish and Thai. Meta discourages use in other languages without extra fine-tuning and controls.
Quantization
The card says 8-bit and 4-bit loading works through bitsandbytes. For other formats see GPTQ, AWQ and GGUF in the glossary.
Launch commands by Llama 3.3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Llama-3.3-70B-Instruct (70.6B)
Needs about 158 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.
vLLM
vllm serve meta-llama/Llama-3.3-70B-Instruct --tensor-parallel-size 7SGLang
sglang serve --model-path meta-llama/Llama-3.3-70B-Instruct --tp 7Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf unsloth/Llama-3.3-70B-Instruct-GGUFPrebuilt GGUF weights published at unsloth/Llama-3.3-70B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct
- https://docs.vllm.ai/en/latest/features/tool_calling.html
Updated 2026-10-07.