What Is Ollama? Models, GPU Requirements, Setup (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

Ollama is a free, open-source tool that downloads large language models and runs them on your own machine behind one command and a local HTTP API. It wraps the llama.cpp inference library, so you get GPU acceleration on NVIDIA, AMD and Apple hardware without compiling anything.

If you only want to chat with a model on a laptop or a single GPU box, Ollama is the shortest path there. If you need to serve many users at once, you will outgrow it, and this guide shows where that line sits. It is part of our guide to LLM inference engines.

TL;DR

  • What it is: a local model runner and model manager. One binary pulls a model, loads it onto your GPU and serves an API on port 11434.
  • Best for: one user or a small team, development, prototyping, desktop apps, and local agents.
  • GPU support: NVIDIA compute capability 5.0 or newer, AMD through ROCm, Apple through Metal, plus a Vulkan backend on Windows and Linux (Ollama docs).
  • Not built for: high-concurrency serving. The default is one parallel request per loaded model. For that job, see vLLM vs Ollama.
  • Moving to the cloud: the same Docker image runs on a rented NVIDIA GPU when your laptop runs out of VRAM.

What Ollama is and how it works

Ollama has two parts. A background server loads models into memory and exposes an HTTP API on 127.0.0.1:11434 by default. A command-line client talks to that server so you can pull, run, list and remove models. The project's README names llama.cpp as the supported backend, which means the model files are GGUF, the same format described in our GGUF glossary entry.

When you run a model, Ollama checks how much memory it needs, places as many layers as fit onto the GPU, and keeps the rest in system RAM if necessary. The ollama ps command shows the result in a PROCESSOR column: "100% GPU" means the model sits entirely in VRAM, while a split such as "48%/52% CPU/GPU" means part of it was loaded into system memory (Ollama FAQ). A split is the classic reason a local model feels slow, and the fix is a smaller model, a lower quantization, or more VRAM.

Ollama also keeps models loaded after a request. By default a model stays in memory for 5 minutes before it is unloaded, and the OLLAMA_KEEP_ALIVE variable or the per-request keep_alive parameter changes that (Ollama FAQ).

Install and first run

The commands below are copied from the Ollama README (v0.40.1 was the latest release on GitHub when this was written).

macOS or Linux:

curl -fsSL https://ollama.com/install.sh | sh

Windows (PowerShell):

irm https://ollama.com/install.ps1 | iex

Then run a model. The README's own example is:

ollama run gemma4

The first run downloads the weights, so it takes as long as your connection needs, and every run after that starts from disk. Model files live under ~/.ollama/models on macOS, /usr/share/ollama/.ollama/models on Linux, and C:\Users\%username%\.ollama\models on Windows; set OLLAMA_MODELS to move them (Ollama FAQ).

Call the API

The README shows the native chat endpoint:

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{
    "role": "user",
    "content": "Why is the sky blue?"
  }],
  "stream": false
}'

Ollama also speaks the OpenAI wire format at http://localhost:11434/v1/. Per the compatibility docs it supports /v1/chat/completions, /v1/completions, /v1/responses (non-stateful only), /v1/models and /v1/embeddings. The client needs an API key string, but Ollama ignores its value:

from openai import OpenAI

client = OpenAI(
    base_url='http://localhost:11434/v1/',
    api_key='ollama',  # required but ignored
)

chat_completion = client.chat.completions.create(
    messages=[{'role': 'user', 'content': 'Say this is a test'}],
    model='gpt-oss:20b',
)
print(chat_completion.choices[0].message.content)

That compatibility is why most agent frameworks, chat front ends and coding tools can point at Ollama by changing a base URL.

What are Ollama models?

"Ollama models" are prebuilt, quantized model packages published in the Ollama library at ollama.com/library. Each one is a name plus a tag, for example gpt-oss:20b, where the tag usually encodes the parameter count and sometimes the quantization. Ollama downloads the matching GGUF weights plus a template and default parameters so the model behaves correctly out of the box.

You are not limited to the library. A Modelfile lets you build your own named model on top of a library model, a directory of Safetensors weights, or a GGUF file you downloaded elsewhere. This example comes straight from the Modelfile docs:

FROM llama3.2
PARAMETER num_ctx 4096
SYSTEM You are Mario from super mario bros, acting as an assistant.

Save it as Modelfile, then build and run it:

ollama create mario -f ./Modelfile
ollama run mario

To import a GGUF file, point FROM at it, for example FROM ./ollama-model.gguf. That is the path to take for a fine-tune you exported yourself, as covered in our fine-tuning overview.

Ollama GPU requirements

The GPU support rules, from the Ollama GPU documentation:

HardwareWhat Ollama needs
NVIDIACompute capability 5.0 or newer and driver version 550 or newer. Cards at compute capability 5.0 through 6.2 need driver 570 or newer.
AMD (Linux)The AMD ROCm v7 driver, on supported Radeon RX, Radeon AI PRO, Radeon PRO, Ryzen AI and Instinct parts.
AMD (Windows)Limited to certain Radeon RX 7000-series and Radeon PRO W7000-series cards, with a ROCm v7 / HIP7-capable driver stack.
AppleGPU acceleration through the Metal API.
VulkanAdds GPU support on Windows and Linux, enabled by default once the backend is installed.

How much VRAM do you need?

The documentation does not publish a per-model VRAM table, so here is the arithmetic. Weight memory is roughly parameter count times bytes per parameter. These figures are computed by us and cover weights only, not the KV cache or runtime overhead:

Model sizeFP16 (2 bytes)INT8 (1 byte)INT4 (about 0.5 byte)
8B16 GB8 GB4 GB
14B28 GB14 GB7 GB
32B64 GB32 GB16 GB
70B140 GB70 GB35 GB

As a rule of thumb from the table, a 4-bit 8B model fits an 8 GB card with little room for context, and a 32B model wants a 24 GB card once you add context (our estimate). For exact sizing against your context length, use the VRAM calculator or read how much VRAM you need for LLMs.

Context length matters. According to the Ollama docs, the default context depends on your memory: 4k tokens under 24 GiB of VRAM, 32k tokens at 24 to 48 GiB, and 256k tokens at 48 GiB or more. The docs recommend at least 64,000 tokens for web search, agents and coding tools, and warn that a larger context uses more memory. You set it with OLLAMA_CONTEXT_LENGTH, for example:

OLLAMA_CONTEXT_LENGTH=64000 ollama serve

To shrink the KV cache, OLLAMA_KV_CACHE_TYPE accepts f16 (the default), q8_0 (about half the memory) and q4_0 (about a quarter), and it applies when Flash Attention is on. Ollama turns Flash Attention on automatically when the backend and device support it. Our KV cache glossary entry explains why this cache grows with every token.

Concurrency and the limits of Ollama

Ollama's defaults tell you what it is designed for. From the FAQ:

  • OLLAMA_NUM_PARALLEL defaults to 1: each loaded model handles one request at a time unless you raise it, and memory use scales with this value times the context length.
  • OLLAMA_MAX_LOADED_MODELS defaults to 3 times the number of GPUs (or 3 for CPU inference).
  • OLLAMA_MAX_QUEUE defaults to 512, and requests beyond it are rejected with a 503.
  • OLLAMA_HOST defaults to 127.0.0.1:11434, so the server is local-only until you change it.

You can raise parallelism, and for a handful of users that is fine. For dozens of concurrent users and strict latency targets, a server built around continuous batching and paged KV memory is the better tool. We compare them directly in vLLM vs Ollama.

With several GPUs, Ollama loads a model onto a single GPU if it fits there, and only spreads it across all available GPUs when it does not (Ollama FAQ). That is a capacity feature, not a speed feature.

Run Ollama in Docker

The Ollama docs give these commands. CPU only:

docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

NVIDIA GPU (needs the NVIDIA Container Toolkit):

docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

AMD GPU, using the rocm tag:

docker run -d --device /dev/kfd --device /dev/dri -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama:rocm

Then pull and run a model inside the container with docker exec -it ollama ollama run gemma4.

To change server settings on a Linux install that uses systemd, run systemctl edit ollama.service, add an Environment line under [Service] such as Environment="OLLAMA_HOST=0.0.0.0:11434", then run systemctl daemon-reload and systemctl restart ollama. If you bind to 0.0.0.0, restrict access with a firewall or a private network, since the API as documented here has no key check.

When to pick Ollama and when not to

Pick Ollama when:

  • you are one person or a small team and want a model running in two minutes;
  • you want a stable local API for an app, an agent or a coding tool;
  • you want to try many models quickly without managing weights by hand.

Look elsewhere when:

  • you need the highest aggregate throughput on one GPU, where vLLM is built for the job;
  • you want fine control of every runtime flag, where raw llama.cpp is the lower-level tool;
  • you prefer a desktop app with a model browser, which is the comparison in Ollama vs LM Studio.

Performance

The Ollama project does not publish a benchmark table, and speed depends on the model, quantization, context length and GPU, so we do not quote tokens per second here. Measure on your own hardware: run the model with a realistic prompt, check that ollama ps shows 100% GPU, and compare. Our GPU buying guide for inference covers what to look at when choosing hardware.

Run it on a cloud GPU

When your local card cannot hold the model you want, or you need a box that stays up when your laptop closes, start the same Docker image on a rented NVIDIA GPU and tunnel port 11434 to your machine. The boxes below show what is available right now.

FAQ

Is Ollama free?

Yes. Ollama is open source and the local runner is free to install and use. The model weights you download carry their own licenses, so check each model's license before commercial use.

Does Ollama need a GPU?

No. Ollama runs on CPU, but generation is much slower. A GPU is strongly recommended, and the ollama ps PROCESSOR column tells you whether the model is on the GPU.

What GPU do I need for Ollama?

Any NVIDIA card with compute capability 5.0 or newer, a supported AMD card through ROCm, or an Apple Silicon Mac through Metal. What matters in practice is VRAM: roughly half a byte per parameter at 4-bit, plus context. See the table above.

Is Ollama the same as llama.cpp?

No. Ollama is a model manager and server that uses llama.cpp as its backend. llama.cpp is the underlying engine, which you can also run directly with more control over flags. See our llama.cpp guide.

Can Ollama serve many users at once?

It can serve several, but its default is one parallel request per model. For heavy concurrency, an engine with continuous batching such as vLLM is the better fit.

Is the Ollama API compatible with OpenAI clients?

Largely, yes. It exposes /v1/chat/completions, /v1/completions, /v1/responses (non-stateful), /v1/models and /v1/embeddings at http://localhost:11434/v1/.

Sources

Related reading: Ollama vs LM Studio, vLLM vs Ollama, llama.cpp guide, best GPU for AI, and the RTX 4090 and RTX 5090 pages.

#llm inference#inference engines#ollama#local llm#gguf#llama.cpp

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.