How to run Llama 3.1 with vLLM, SGLang and llama.cpp
Serve Llama 3.1 Instruct with vLLM or SGLang, or run quantized builds in llama.cpp and Ollama. Version needs, chat template and 128k context notes.
The Llama 3 family hub lists sizes and per-GPU fit. This page covers the software side for the Llama 3.1 generation.
Engines
The Llama 3.1 model cards show usage with:
- Transformers, which needs version 4.43.0 or newer.
- vLLM, served through its OpenAI-compatible server.
- SGLang, which the card shows with a Docker launch.
- Quantized builds that work with llama.cpp, Ollama and LM Studio. See GGUF for the file format.
The cards do not state a minimum version for vLLM, SGLang, llama.cpp or Ollama, so use a current release of each.
Chat template and tools
Use the Instruct checkpoint for chat. Apply the template through the tokenizer (apply_chat_template) rather than hand-building prompts. The card states that tool use is also supported through chat templates in Transformers.
Llama 3.1 has no thinking toggle. It is a plain instruction-tuned model.
Context length
The cards list a 128k token context window. Long context is what drives memory use beyond the weights, so cap --max-model-len in vLLM (or the equivalent in your engine) to what your workload needs. Background on extended context is in RoPE scaling.
Pitfalls
- Official languages are English, German, French, Italian, Portuguese, Hindi, Spanish and Thai. Other languages are out of scope without fine-tuning.
- The weights are gated on Hugging Face: accept the Llama 3.1 Community License and authenticate before downloading.
- The card says the model is not designed to be deployed in isolation and should sit inside a system with safety guardrails.
- Knowledge cutoff is December 2023.
Launch commands by Llama 3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Meta-Llama-3-8B-Instruct (8.0B)
Needs about 17.9 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve meta-llama/Meta-Llama-3-8B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path meta-llama/Meta-Llama-3-8B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Meta-Llama-3-70B (70.6B)
Needs about 158 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.
vLLM
vllm serve meta-llama/Meta-Llama-3-70B --tensor-parallel-size 7SGLang
sglang serve --model-path meta-llama/Meta-Llama-3-70B --tp 7Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct
- https://huggingface.co/meta-llama/Llama-3.1-8B
Updated 2026-10-07.