How to run Falcon3 with vLLM, SGLang and llama.cpp
Serve Falcon3 Instruct with vLLM or SGLang, or run quantized builds in llama.cpp and Ollama. Context length, chat template and languages.
Documentation for Falcon3 is short. See the Falcon3 hub for sizes and GPU fit.
Engines
The Falcon3-7B-Instruct card lists support for:
- Transformers.
- vLLM, through an OpenAI-compatible API.
- SGLang, through an OpenAI-compatible API.
- Docker Model Runner.
- Quantized variants usable in llama.cpp, Ollama and LM Studio (see GGUF).
The card states no minimum engine versions, so use a current release.
Chat template
Use the Instruct checkpoint and format messages with the tokenizer's apply_chat_template, with system and user roles. The card documents no thinking toggle, and no tool-call parser flags.
Context length
The cards list a 32K context window with a high RoPE value (1000042) to support long context. They do not describe any scaling flags, so no extra rope settings are documented.
Pitfalls
- Supported languages are English, French, Spanish and Portuguese.
- The Instruct model was post-trained on 1.2 million samples covering STEM, conversation, code, safety and function-call data.
- The license is the TII Falcon-LLM License 2.0, which has conditions on how you share derivatives. See the fine-tuning guide for the summary.
Launch commands by Falcon3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Falcon3-1B-Instruct (1.7B)
Needs about 3.7 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve tiiuae/Falcon3-1B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path tiiuae/Falcon3-1B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Falcon3-7B-Instruct (7.5B)
Needs about 16.7 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve tiiuae/Falcon3-7B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path tiiuae/Falcon3-7B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.