How to run SmolLM3 with vLLM, SGLang and llama.cpp

Serve SmolLM3-3B with vLLM or SGLang, toggle /think and /no_think, extend context with YaRN, and set the tool-call parser.

SmolLM3 is a small instruct model with switchable reasoning. See the SmolLM3 hub for sizes and GPU fit.

Engines

The model card documents:

  • Transformers 4.53.0 or newer.
  • vLLM, with --enable-auto-tool-choice and --tool-call-parser=hermes for tool calling.
  • SGLang, launched with its standard launch_server entry point.
  • llama.cpp, through quantized checkpoints published in the SmolLM3 collection (see GGUF).

The card does not give minimum versions for vLLM, SGLang or llama.cpp, and it does not mention Ollama.

Reasoning toggle

SmolLM3 reasons by default. Control it with flags in the system prompt:

  • /think enables reasoning (the default).
  • /no_think disables it.
  • In Transformers you can pass enable_thinking=False to apply_chat_template.

If both are used, the system prompt flag wins over the enable_thinking argument.

Context length

The default is 65,536 tokens (max_position_embeddings). The card shows extending to 131,072 tokens by setting a YaRN rope_scaling entry in the config with factor 2.0 and original_max_position_embeddings 65536. See RoPE scaling for what that changes.

Tool calling

The card describes two tool formats: xml_tools, which returns <tool_call> JSON, and python_tools, which returns calls inside <code> tags. For vLLM, use the hermes parser flag shown above.

Checkpoints

SmolLM3-3B is the instruction-tuned release (the card's recommended model) and SmolLM3-3B-Base is the pretrained one. The license is Apache 2.0.

Launch commands by SmolLM3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

SmolLM3-3B (3.1B)

Needs about 6.9 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve HuggingFaceTB/SmolLM3-3B --tensor-parallel-size 1

SGLang

sglang serve --model-path HuggingFaceTB/SmolLM3-3B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.