How to run Falcon3 with vLLM, SGLang and llama.cpp

Serve Falcon3 Instruct with vLLM or SGLang, or run quantized builds in llama.cpp and Ollama. Context length, chat template and languages.

Documentation for Falcon3 is short. See the Falcon3 hub for sizes and GPU fit.

Engines

The Falcon3-7B-Instruct card lists support for:

  • Transformers.
  • vLLM, through an OpenAI-compatible API.
  • SGLang, through an OpenAI-compatible API.
  • Docker Model Runner.
  • Quantized variants usable in llama.cpp, Ollama and LM Studio (see GGUF).

The card states no minimum engine versions, so use a current release.

Chat template

Use the Instruct checkpoint and format messages with the tokenizer's apply_chat_template, with system and user roles. The card documents no thinking toggle, and no tool-call parser flags.

Context length

The cards list a 32K context window with a high RoPE value (1000042) to support long context. They do not describe any scaling flags, so no extra rope settings are documented.

Pitfalls

  • Supported languages are English, French, Spanish and Portuguese.
  • The Instruct model was post-trained on 1.2 million samples covering STEM, conversation, code, safety and function-call data.
  • The license is the TII Falcon-LLM License 2.0, which has conditions on how you share derivatives. See the fine-tuning guide for the summary.

Launch commands by Falcon3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Falcon3-1B-Instruct (1.7B)

Needs about 3.7 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve tiiuae/Falcon3-1B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path tiiuae/Falcon3-1B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Falcon3-7B-Instruct (7.5B)

Needs about 16.7 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve tiiuae/Falcon3-7B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path tiiuae/Falcon3-7B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.