How to run Phi-4 with vLLM, SGLang and llama.cpp
Run Phi-4 with vLLM, SGLang or llama.cpp: chat template, the 16K context, and language and code pitfalls from the Microsoft model card.
Phi-4 is a single instruction-aligned checkpoint. See the Phi-4 hub for hardware.
Engine support
- vLLM, SGLang and Transformers are covered on the card. vLLM lists Phi-3 and Phi-4 under
Phi3ForCausalLM. - llama.cpp and Ollama: the card links quantized builds (GGUF). No minimum versions are stated.
Chat template
Phi-4 uses a role-tagged format with system, user and assistant turns:
<|im_start|>system<|im_sep|> ... <|im_end|><|im_start|>user<|im_sep|> ... <|im_end|><|im_start|>assistant<|im_sep|>
Use the tokenizer's chat template rather than building this by hand.
Context length
The card states a context length of 16K tokens. No rope scaling or extension flags are documented, so keep the length flag at or below that.
Pitfalls
- Trained primarily on English; other languages perform worse.
- Most training data is Python with common packages. The card strongly recommends manually verifying API use for other languages or packages.
- It can produce plausible but wrong content.
- This is text only. Phi-4-multimodal is a separate model (vLLM lists it as
Phi4MMForCausalLM).
Launch commands by Phi-4 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Phi-4-mini-instruct (3.8B)
Needs about 8.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve microsoft/Phi-4-mini-instruct --tensor-parallel-size 1SGLang
sglang serve --model-path microsoft/Phi-4-mini-instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
phi-4 (14.7B)
Needs about 32.8 GB of VRAM at BF16. Cheapest live fit: RTX A6000 at $0.363/hr.
vLLM
vllm serve microsoft/phi-4 --tensor-parallel-size 1SGLang
sglang serve --model-path microsoft/phi-4 --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
llama.cpp
llama-server -hf bartowski/phi-4-GGUFPrebuilt GGUF weights published at bartowski/phi-4-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.