What is model merging?
Model merging combines the weights of two or more trained models into a single model, without running a new training job. Because the result has the same architecture as its parents, it costs no more to serve than one of them. That is the appeal over an ensemble, which has to run every member at inference time.
Where it comes from
The most direct ancestor is weight averaging. In Model soups, Wortsman and colleagues report that averaging the weights of multiple models fine-tuned with different hyperparameters often improves accuracy and robustness, and that you can average many models without additional inference or memory costs. They suggest the fine-tuned models often lie in a single low error basin, which is why their average works. Their ViT-G soup reached 90.94% top-1 accuracy on ImageNet, which the abstract describes as a new state of the art at the time.
Plain averaging can go wrong when models disagree. TIES-Merging (Yadav, Tam, Choshen, Raffel and Bansal, NeurIPS 2023) names two sources of interference: redundant parameter values, and disagreement on the sign of a parameter across models. It resets parameters that changed little during fine-tuning, resolves sign conflicts, then merges only the parameters that agree with the chosen sign.
Tooling
MergeKit, from Arcee, is an open-source library for applying merging strategies to language models. Its paper says it combines specialized checkpoints into multitask models without additional training, can run on any hardware, and that thousands of merged models built with it have ranked among the strongest open checkpoints on the Open LLM Leaderboard at the time.
Example models in the catalog
Merged models are often named by the method rather than the word "merge", so the catalog token is sometimes different.
- LightOnOCR 2 1B bbox soup: its card says it combines OCR-improving RLVR signals with bounding-box-focused RLVR updates through joint merging, preserving OCR quality while adding image localization. It is a 1B-parameter model under Apache 2.0.
- LightOnOCR 2 1B: the sibling checkpoint in the same family, useful as a reference point for the plain model.
For related ways to change a model without a full training run, see abliteration, which edits weights to remove a direction, and LoRA adapters, which train small added weights. With distillation you get a smaller model rather than a blend.
Limits
Merging only works across models that share an architecture and usually a common ancestor, because the weights need to line up parameter by parameter. A merge also inherits every parent's license, so check each one. Quality is not guaranteed to improve, and a merge needs evaluating on your own task.
What it means when you pick a GPU
A merged model is sized like its parents, not like their sum: a merge of two 8B models is an 8B model. Use the VRAM calculator with the merged checkpoint's parameter count and dtype, and apply quantization if you need to fit a smaller card. Building a merge requires loading the parents' weights. Aquanode rents GPUs by the hour, so you can evaluate a merge against its parents on a card sized for one of them.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Distillation
Distillation trains a small student model to imitate a large teacher, so you get much of the quality at a fraction of the GPU memory.
Abliteration
Abliteration edits a language model's weights to remove its refusal direction, with no training run. How it works and what it means for GPU memory.
LoRA
LoRA fine-tunes an LLM by freezing its weights and training small low-rank matrices added to its layers, cutting trainable parameters and GPU memory.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.