What is MLX?
Abbreviated MLX
MLX is an array framework for machine learning on Apple silicon, built by Apple machine learning research. A model labeled "MLX" on a model page is a checkpoint converted to the format that framework loads, so it is meant to run on a Mac or other Apple device, not on a data-center GPU. The project describes itself in its repository as "an array framework for machine learning on Apple silicon, brought to you by Apple machine learning research."
What the framework does
According to its README, MLX has a NumPy-style Python API plus C++, C and Swift APIs, with higher-level packages that follow PyTorch conventions. It supports function transformations such as automatic differentiation, vectorization and graph optimization. Computation is lazy, so arrays are only materialized when needed, and computation graphs are built dynamically.
The feature that matters most for model sizing is the memory model. The README says arrays in MLX live in shared memory, and that operations on them can run on any supported device type without transferring data. That is the same idea as unified memory: the CPU and GPU draw on one pool, so a Mac's system memory is its model memory.
The README also lists CUDA support on Linux and CPU-only Linux installs alongside Apple silicon. The MLX-format checkpoints in the catalog are nonetheless published for the Apple path, so treat Apple silicon as the target unless the card says otherwise.
MLX models in the catalog
MLX repos are conversions of an existing model, usually with a stated precision in the name. Examples:
- Qwen3 1.7B MLX bf16. The card says it was converted to MLX format from Qwen/Qwen3-1.7B using mlx-lm 0.24.0, and shows
pip install mlx-lmfor running it, withmlx_lm.chatand an OpenAI-compatiblemlx_lm.serveras options. - Qwen3 VL 4B Instruct MLX 4bit, the same format with 4-bit quantization in the name.
- LFM2 1.2B MLX bf16 and Gemma 3n E4B it MLX bf16, small models kept at BF16.
- Parakeet TDT 0.6B v2, a speech-to-text model whose card says it was converted from nvidia/parakeet-tdt-0.6b-v2 and recommends the parakeet-mlx and mlx-audio tools. The card lists a CC-BY-4.0 license.
The precision token in the name (bf16, 4bit, 1bit) tells you how many bits each weight takes, which is what sets the file size and the memory it needs.
MLX versus GGUF
GGUF is the other common format for running models locally, and MLX checkpoints are the format that MLX and mlx-lm load. They are different file formats for different runtimes, so match the format to the software you will run.
What it means when you pick a GPU
An MLX checkpoint is a sign you are looking at a local-Mac build. To serve the same model on a rented GPU, use the original (or an FP8/AWQ/GPTQ) repo for that model instead of the MLX conversion. The MLX repo's name and card still help: they tell you the base model and a precision, which you can feed into the VRAM calculator to size a GPU. Aquanode rents GPUs by the hour for the cases where a model is too big for a laptop.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Unified Memory
Unified memory lets the CPU and GPU share one memory pool, as in Apple M-series chips and NVIDIA Grace Hopper. It adds capacity, not bandwidth.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
BF16 (bfloat16)
BF16 is a 16-bit float with FP32's 8 exponent bits but only 7 mantissa bits. It is the default for training and costs 2 bytes per model parameter.
GGUF
GGUF is a single-file binary format that packs a model's weights and metadata together, used by GGML-based runtimes such as llama.cpp.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.