What is GGUF?

Abbreviated GGUF

GGUF is a binary file format for storing a machine learning model so it can be loaded for inference. According to the GGUF specification, it is designed for use with GGML and related executors, and it succeeds earlier formats named GGML, GGMF and GGJT. One GGUF file holds everything a runtime needs: the tensors (the weights) and the metadata that describes the model.

What is inside a GGUF file

The specification lists five goals for the format: single-file deployment, extensibility, memory-mapped loading (mmap), ease of use, and containing all information needed to load a model in the file itself.

  • Metadata as key-value pairs. Earlier formats stored hyperparameters as an untyped list of values. GGUF uses a key-value structure instead, which the spec calls metadata. It covers the model architecture, the quantization version, alignment, and standard fields such as author, version and license. New keys can be added without breaking models that already exist, which is what the spec means by extensibility.
  • Tensors. The weights are stored in the same file, using standardized tensor names for transformer parts such as embeddings, attention blocks and feed-forward networks.
  • Byte order. Models are little-endian by default. Big-endian is supported, with all values and tensors adjusted to match (added in version 3 of the format).

Because the file is self-contained, you download one file and point a runtime at it. There is no separate config, tokenizer or weights folder to keep in sync.

GGUF and quantization

GGUF is a container, not a quantization method. It can hold unquantized tensors (the spec lists F32 and F16) and also many quantized types. The specification lists 40 tensor types, including the Q4_0 through Q6_K families, integer types, and specialized types such as IQ2_XXS and MXFP4. When people say "a Q4 GGUF," they mean a GGUF file whose weights are stored in one of those 4-bit types. See quantization for how reduced precision trades accuracy for memory.

This is the difference from AWQ and GPTQ, which are names for quantization algorithms, and EXL2, which is both a quantization method and the file format its own runtime reads. A GGUF file's tensor types are defined by the GGML ecosystem, so the format and its quantization types travel together.

Example models on Aquanode

GGUF files are published by model authors and by third parties who convert existing models. Be aware that a repository's contents can change over time.

  • Surogate Rune 26B-A4B GGUF is a fine-tune of Google's Gemma 4 26B-A4B, a mixture-of-experts model with 8 of 128 experts active per token, according to its model card. The card says the repository currently offers bfloat16 safetensors, with earlier GGUF builds available in the revision history, so check which files the revision you pull actually contains.

That caveat is a good habit in general. The name of a repository tells you what it was meant to hold, and the file list tells you what it holds.

What GGUF does not decide

The format does not change what a model needs at runtime. Weights stored in 4-bit types still have to be dequantized for math, and the KV cache that grows with context length is held separately from the weights. A smaller file lowers the memory floor but does not remove the cache.

What it means when you pick a GPU

GGUF is mainly a runtime choice. If your serving stack is built on GGML-based software, you will be handed GGUF files, and the file size of the quantization level you pick is a good first estimate of the VRAM the weights need, with the cache and activations on top. If your stack is vLLM or another GPU-first server, you will more often meet AWQ or GPTQ repositories, and the same size-first reasoning applies. Use the VRAM calculator to check a specific model and context length, then choose a card with headroom. Aquanode rents GPUs by the hour, so you can try a quantization level on a card before committing to it.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.