llama.cpp Guide: llama-server, GPU Offload, Multi-GPU (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

llama.cpp is an open-source C/C++ engine that runs quantized large language models on almost any hardware, and its llama-server program exposes an OpenAI-compatible HTTP API. On a GPU, you control how much of the model sits in VRAM with --n-gpu-layers, and how it divides across several cards with --split-mode and --tensor-split.

Ollama and LM Studio are both built on it, so learning llama.cpp directly gives you the most control over the same engine. This guide is part of our overview of LLM inference engines.

TL;DR

  • What it is: an LLM inference library and tools written in C/C++ with no external dependencies, built on the ggml library (project README).
  • Backends: CUDA (NVIDIA), HIP (AMD), Metal (Apple), Vulkan, SYCL (Intel), OpenCL, CANN, WebGPU and more, plus CPU and CPU+GPU hybrid inference for models larger than your VRAM.
  • Server: llama-server gives OpenAI-style endpoints, continuous batching (on by default) and parallel slots.
  • Multi-GPU: layer split by default, row split available, and an experimental tensor split mode.
  • Pick it when: you want fine control, GGUF models, hybrid CPU/GPU offload, or a small dependency-free server. For many concurrent users on server GPUs, look at vLLM.

What llama.cpp is

The README describes the goal as "LLM inference in C/C++" with minimal setup across a wide range of hardware, local and cloud. It is optimized for Apple silicon and for x86 CPUs (AVX, AVX2, AVX512, AMX), supports RISC-V, and supports integer quantization from 1.5-bit through 8-bit to cut memory use and speed up inference.

The project's native model format is GGUF, a single file holding weights, tokenizer and metadata. Conversion scripts in the repository (such as convert_hf_to_gguf.py) turn Hugging Face checkpoints into GGUF, and many models are already published in GGUF form on Hugging Face with a -GGUF suffix. See the GGUF glossary entry and quantization.

The supported backends listed in the README are: BLAS, BLIS, CANN (Ascend NPU), CUDA (NVIDIA), HIP (AMD), Hexagon (Snapdragon), IBM zDNN, MUSA (Moore Threads), Metal (Apple silicon), OpenCL (Adreno), OpenVINO (in progress), RPC, SYCL (Intel GPU), VirtGPU, Vulkan, WebGPU and ZenDNN (AMD CPU). Builds are tagged continuously (b11496 was the newest tag on 8 October 2026), so expect the number to be higher by the time you read this.

Install

The README gives several routes. Its install script, and then the first commands, are:

curl -LsSf https://llama.app/install.sh | sh
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

On Windows the script is irm https://llama.app/install.ps1 | iex. The -hf flag downloads a model straight from Hugging Face. The README also lists running with Docker, downloading pre-built binaries from the releases page, and building from source.

Pre-built CUDA binaries

The release page for b11496 includes Linux x64 builds for CUDA 12.8 and CUDA 13.4, a Linux arm64 build for CUDA 13.4, and Windows x64 builds for CUDA 12.4 and 13.4. Match the CUDA version to your driver.

Build from source with CUDA

From the project's build guide:

cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release

By default the build targets the GPUs connected to the machine at that time. The guide also gives cmake -B build -DGGML_VULKAN=1 for Vulkan and cmake -S . -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1030 -DCMAKE_BUILD_TYPE=Release for AMD HIP. The binaries land in build/bin, including llama-server and llama-quantize.

Run llama-server

The server README shows a minimal start, which binds 127.0.0.1:8080:

./llama-server -m models/7B/ggml-model.gguf -c 2048

The OpenAI-compatible endpoints listed in the docs are GET /v1/models, POST /v1/completions, POST /v1/chat/completions, POST /v1/responses, POST /v1/embeddings and POST /v1/rerank. A chat request looks like this (the model field is ignored in single-model mode):

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Say this is a test"}]
  }'

Tool calling needs the --jinja flag, which the docs say is enabled by default, and some models need a --chat-template-file override to get a tool-compatible template.

Docker with a GPU

The Docker docs publish ghcr.io/ggml-org/llama.cpp:server-cuda (CUDA 12) and ghcr.io/ggml-org/llama.cpp:server-cuda13 (CUDA 13). The server README shows:

docker run -p 8080:8080 -v /path/to/models:/models --gpus all \
  ghcr.io/ggml-org/llama.cpp:server-cuda -m models/7B/ggml-model.gguf \
  -c 512 --host 0.0.0.0 --port 8080 --n-gpu-layers 99

On Linux the host needs the NVIDIA Container Toolkit. Adjust the -m path to the file inside the mounted /models directory.

GPU offload: llama cpp gpu settings

-ngl, --gpu-layers or --n-gpu-layers sets the maximum number of model layers stored in VRAM. The server docs list the default as auto, and the value can be a number, auto or all. Layers that do not fit stay on the CPU, which lets you run a model larger than your VRAM at reduced speed.

Two rules of thumb follow from the arithmetic. First, if the whole model plus its KV cache fits in VRAM, offload everything. Second, if it does not, add layers until memory is nearly full; the speed drop comes from the CPU-resident layers, not from llama.cpp itself.

Quantization sets how big those layers are. The quantize README publishes sizes for Llama-3.1-8B:

Quant typeSize (GiB)Bits per weight
Q4_K_M4.584.8944
Q5_K_M5.335.7036
Q8_07.958.5008

Those figures are the project's, for that one model. For other sizes, scale by parameter count (computed: an 8B model at 4.89 bits per weight is about 4.6 GiB, so a 70B model at the same quant is roughly 40 GiB of weights before the KV cache). The VRAM calculator adds context on top, and how much VRAM you need for LLMs walks through it.

To make your own quant from a full-precision GGUF, the README example is:

./build/bin/llama-quantize gemma-4-E2B-it-bf16.gguf gemma-4-E2B-it-Q4_K_M.gguf Q4_K_M

Context size and parallel slots

-c or --ctx-size defaults to 0, meaning the context is loaded from the model. -np or --parallel sets the number of server slots and defaults to auto. Continuous batching (-cb) is on by default, and Flash Attention (-fa) defaults to auto. More slots let several requests run together, and each slot needs its own share of the KV cache, so a bigger -np costs VRAM. Our KV cache glossary entry explains why.

Multi-GPU: llama cpp multi gpu

The server docs describe three ways to split a model across GPUs with -sm or --split-mode:

  • layer (the default): split layers and KV cache across GPUs, pipelined.
  • row: split weights across GPUs by rows, parallelized.
  • tensor: split weights and KV across GPUs, parallelized, marked EXPERIMENTAL.
  • none: use one GPU only.

Control the proportion with -ts or --tensor-split, a comma-separated list of fractions such as 3,1, so a card with more VRAM can carry more. -mg or --main-gpu picks the GPU for single-GPU mode or for intermediate results.

Example on two GPUs with an uneven split:

./llama-server -m model.gguf --n-gpu-layers 99 --split-mode layer --tensor-split 3,1

Layer split moves only small activations between cards, so it works over plain PCIe and is the safest start. The tensor mode is experimental per the docs, so test it on your model before relying on it. The concept behind the experimental mode is tensor parallelism, which is what vLLM implements as a mature production feature.

Router mode for several models

Newer builds of llama-server can run in router mode, which exposes an API for dynamically loading and unloading models and forwards each request to the right model instance. The options listed include --models-dir for a directory of models, --models-preset for an INI preset file, and --models-max for how many models may be loaded at once (default 4). In router mode the request must name the model. Autoloading is enabled by default.

llama.cpp vs Ollama vs vLLM

llama.cppOllamavLLM
LevelEngine and serverModel manager on top of llama.cppServer-class engine
ControlEvery flagDefaults plus a ModelfileMany serving flags
HardwareWidest (CPU, CUDA, HIP, Metal, Vulkan, SYCL and more)NVIDIA, AMD, Apple, VulkanLinux GPUs and other accelerators
Best forFine control, hybrid CPU/GPU, small footprintQuick startMany concurrent users

Choose llama.cpp when you want to tune offload, split a model unevenly across mismatched cards, or run on hardware that other engines skip. Choose Ollama for convenience, and vLLM for high-concurrency serving.

Performance

The llama.cpp project does not publish a single canonical benchmark table that we can cite for a given GPU, and speed varies with model, quantization, context and offload. We do not quote tokens per second. To measure, run llama-server, send realistic prompts and watch the timing the server prints. For choosing hardware, see best GPU for LLM inference.

Run it on a cloud GPU

If your own GPU cannot hold the model, rent a box, pull the server-cuda image above and tunnel port 8080. The live boxes below show what is available now.

FAQ

What is llama.cpp used for?

It runs quantized LLMs locally or on a server, on CPUs and many kinds of GPUs, and llama-server exposes an OpenAI-compatible API.

How do I run llama.cpp on a GPU?

Use a CUDA, HIP, Metal or Vulkan build, then set --n-gpu-layers so the layers sit in VRAM. The default is auto, and all offloads every layer.

How do I use multiple GPUs with llama.cpp?

Use --split-mode (layer by default, row, or experimental tensor) and --tensor-split for proportions, for example 3,1.

What port does llama-server use?

8080 by default, bound to 127.0.0.1. Use --host 0.0.0.0 to listen on all interfaces and --api-key to require a key.

Is llama.cpp faster than Ollama?

Ollama uses llama.cpp as its backend, so the engine is the same. Differences come from defaults and settings. No neutral benchmark exists that we can cite.

Does llama.cpp support continuous batching?

Yes. The server docs list -cb as enabled by default.

Sources

Related reading: what is Ollama, Ollama vs LM Studio, vLLM vs Ollama, RTX 4090.

#llm inference#inference engines#llama.cpp#gguf#quantization#multi-gpu

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.