How to run Llama 3.3 with vLLM and Transformers

Engine support, Transformers version, tool-call parser, context length and quantization notes for running Llama 3.3 70B Instruct.

This guide covers serving the Llama 3.3 family, which on the Hugging Face card is a single 70B Instruct model. Sizing and launch commands appear below.

Engine support

  • The card documents Transformers 4.45.0 or newer, including conversational use with the chat template.
  • It documents vLLM through vllm serve with an OpenAI-compatible API.
  • The card does not discuss SGLang, llama.cpp or Ollama.

Chat template and tools

The model supports chat templates and several tool formats through Transformers. In vLLM, tool calling uses the llama3_json parser with a chat template for the 3.1 or 3.2 generation (--tool-call-parser, --chat-template). vLLM notes that parallel tool calls are not supported for Llama 3, and that the model may serialize arrays as strings in tool arguments.

Context and languages

  • Context length is 128k tokens.
  • Officially supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish and Thai. Meta discourages use in other languages without extra fine-tuning and controls.

Quantization

The card says 8-bit and 4-bit loading works through bitsandbytes. For other formats see GPTQ, AWQ and GGUF in the glossary.

Launch commands by Llama 3.3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Llama-3.3-70B-Instruct (70.6B)

Needs about 158 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.

vLLM

vllm serve meta-llama/Llama-3.3-70B-Instruct --tensor-parallel-size 7

SGLang

sglang serve --model-path meta-llama/Llama-3.3-70B-Instruct --tp 7

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run llama3.3:70b

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf unsloth/Llama-3.3-70B-Instruct-GGUF

Prebuilt GGUF weights published at unsloth/Llama-3.3-70B-Instruct-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.