How to run Phi-4 with vLLM, SGLang and llama.cpp

Run Phi-4 with vLLM, SGLang or llama.cpp: chat template, the 16K context, and language and code pitfalls from the Microsoft model card.

Phi-4 is a single instruction-aligned checkpoint. See the Phi-4 hub for hardware.

Engine support

  • vLLM, SGLang and Transformers are covered on the card. vLLM lists Phi-3 and Phi-4 under Phi3ForCausalLM.
  • llama.cpp and Ollama: the card links quantized builds (GGUF). No minimum versions are stated.

Chat template

Phi-4 uses a role-tagged format with system, user and assistant turns:

  • <|im_start|>system<|im_sep|> ... <|im_end|>
  • <|im_start|>user<|im_sep|> ... <|im_end|>
  • <|im_start|>assistant<|im_sep|>

Use the tokenizer's chat template rather than building this by hand.

Context length

The card states a context length of 16K tokens. No rope scaling or extension flags are documented, so keep the length flag at or below that.

Pitfalls

  • Trained primarily on English; other languages perform worse.
  • Most training data is Python with common packages. The card strongly recommends manually verifying API use for other languages or packages.
  • It can produce plausible but wrong content.
  • This is text only. Phi-4-multimodal is a separate model (vLLM lists it as Phi4MMForCausalLM).

Launch commands by Phi-4 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Phi-4-mini-instruct (3.8B)

Needs about 8.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve microsoft/Phi-4-mini-instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path microsoft/Phi-4-mini-instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

phi-4 (14.7B)

Needs about 32.8 GB of VRAM at BF16. Cheapest live fit: RTX A6000 at $0.363/hr.

vLLM

vllm serve microsoft/phi-4 --tensor-parallel-size 1

SGLang

sglang serve --model-path microsoft/phi-4 --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run phi4

Verified against Ollama's own library listing. Source

llama.cpp

llama-server -hf bartowski/phi-4-GGUF

Prebuilt GGUF weights published at bartowski/phi-4-GGUF. Run with llama.cpp's llama-server or load the repo directly in LM Studio. Source

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.