How to run SmolLM3 with vLLM, SGLang and llama.cpp
Serve SmolLM3-3B with vLLM or SGLang, toggle /think and /no_think, extend context with YaRN, and set the tool-call parser.
SmolLM3 is a small instruct model with switchable reasoning. See the SmolLM3 hub for sizes and GPU fit.
Engines
The model card documents:
- Transformers 4.53.0 or newer.
- vLLM, with
--enable-auto-tool-choiceand--tool-call-parser=hermesfor tool calling. - SGLang, launched with its standard
launch_serverentry point. - llama.cpp, through quantized checkpoints published in the SmolLM3 collection (see GGUF).
The card does not give minimum versions for vLLM, SGLang or llama.cpp, and it does not mention Ollama.
Reasoning toggle
SmolLM3 reasons by default. Control it with flags in the system prompt:
/thinkenables reasoning (the default)./no_thinkdisables it.- In Transformers you can pass
enable_thinking=Falsetoapply_chat_template.
If both are used, the system prompt flag wins over the enable_thinking argument.
Context length
The default is 65,536 tokens (max_position_embeddings). The card shows extending to 131,072 tokens by setting a YaRN rope_scaling entry in the config with factor 2.0 and original_max_position_embeddings 65536. See RoPE scaling for what that changes.
Tool calling
The card describes two tool formats: xml_tools, which returns <tool_call> JSON, and python_tools, which returns calls inside <code> tags. For vLLM, use the hermes parser flag shown above.
Checkpoints
SmolLM3-3B is the instruction-tuned release (the card's recommended model) and SmolLM3-3B-Base is the pretrained one. The license is Apache 2.0.
Launch commands by SmolLM3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
SmolLM3-3B (3.1B)
Needs about 6.9 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve HuggingFaceTB/SmolLM3-3B --tensor-parallel-size 1SGLang
sglang serve --model-path HuggingFaceTB/SmolLM3-3B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.