How to run Llama 3.1 with vLLM, SGLang and llama.cpp

Serve Llama 3.1 Instruct with vLLM or SGLang, or run quantized builds in llama.cpp and Ollama. Version needs, chat template and 128k context notes.

The Llama 3 family hub lists sizes and per-GPU fit. This page covers the software side for the Llama 3.1 generation.

Engines

The Llama 3.1 model cards show usage with:

  • Transformers, which needs version 4.43.0 or newer.
  • vLLM, served through its OpenAI-compatible server.
  • SGLang, which the card shows with a Docker launch.
  • Quantized builds that work with llama.cpp, Ollama and LM Studio. See GGUF for the file format.

The cards do not state a minimum version for vLLM, SGLang, llama.cpp or Ollama, so use a current release of each.

Chat template and tools

Use the Instruct checkpoint for chat. Apply the template through the tokenizer (apply_chat_template) rather than hand-building prompts. The card states that tool use is also supported through chat templates in Transformers.

Llama 3.1 has no thinking toggle. It is a plain instruction-tuned model.

Context length

The cards list a 128k token context window. Long context is what drives memory use beyond the weights, so cap --max-model-len in vLLM (or the equivalent in your engine) to what your workload needs. Background on extended context is in RoPE scaling.

Pitfalls

  • Official languages are English, German, French, Italian, Portuguese, Hindi, Spanish and Thai. Other languages are out of scope without fine-tuning.
  • The weights are gated on Hugging Face: accept the Llama 3.1 Community License and authenticate before downloading.
  • The card says the model is not designed to be deployed in isolation and should sit inside a system with safety guardrails.
  • Knowledge cutoff is December 2023.

Launch commands by Llama 3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Meta-Llama-3-8B-Instruct (8.0B)

Needs about 17.9 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve meta-llama/Meta-Llama-3-8B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path meta-llama/Meta-Llama-3-8B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Meta-Llama-3-70B (70.6B)

Needs about 158 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.

vLLM

vllm serve meta-llama/Meta-Llama-3-70B --tensor-parallel-size 7

SGLang

sglang serve --model-path meta-llama/Meta-Llama-3-70B --tp 7

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.