How to fine-tune Olmo 3 (LoRA, QLoRA)
Base, Think and Instruct checkpoints, training stages, chat template and license notes for fine-tuning Ai2 Olmo 3 with LoRA or QLoRA.
This guide covers adapting the Olmo 3 family with LoRA or QLoRA. Estimate cost with the training cost calculator and see fine-tuning for the overview.
Which checkpoint
- Base (
Olmo-3-1025-7B) was trained on Dolma 3 in three stages: pretraining, mid-training and long-context training. Its card says you can fine-tune from the final checkpoint or from intermediate revisions namedstage1-stepXXX,stage2-stepXXXandstage3-stepXXX, using OLMo-core scripts or other recipes. - Think and Instruct models were post-trained with SFT, DPO and RLVR. The cards point to the Open-Instruct repository for the DPO and RLVR stages.
- For a light domain adapter, start from Instruct (or Think if you need reasoning). For your own full post-training recipe, start from Base.
Chat template
Use the tokenizer's template with <|im_start|> and <|im_end|>. Keep the system message style of the checkpoint you start from. Think data should keep reasoning inside <think> tags.
Target modules
The cards do not list LoRA target modules. The 7B base has 32 layers, hidden size 4096 and 32 query and key-value heads, so inspect the module names in the loaded model.
License
The code and weights are Apache 2.0, which allows derivative models. The cards add that the models are intended for research and educational use in accordance with Ai2's Responsible Use Guidelines, so read those before shipping a commercial derivative.
Memory to fine-tune Olmo 3, by size
Model state in GB before activations, for a full fine-tune, LoRA and QLoRA, with the cheapest live GPU set that has that much memory. One model per size.
| Model | Parameters | Full fine-tune | LoRA | QLoRA | Compute per 1B tokens |
|---|---|---|---|---|---|
| Olmo-3-7B-Instruct | 7.3B | 109 GBRTX A5000 × 5 · $0.880/hr | 13.6 GBRTX A5000 · $0.176/hr | 3.5 GBRTX 4070 Super · $0.121/hr | 4.4 × 10^19 FLOPs |
| Olmo-3-1125-32B | 32.2B | 480 GBRTX PRO 6000 × 6 · $8.25/hr | 60.0 GBA100 · $1.21/hr | 15.5 GBRTX A5000 · $0.176/hr | 1.9 × 10^20 FLOPs |
Full fine-tune counts 16 bytes per parameter (mixed-precision Adam, as counted in the ZeRO paper). LoRA keeps the frozen base in BF16 at 2 bytes per parameter. QLoRA stores the base in 4-bit NormalFloat with double quantization at 4.127 bits per parameter (QLoRA paper). Adapters and activations are not counted: activations depend on your batch size and sequence length, so leave headroom. Compute is the training-cost calculator's 6 × parameters × tokens for a dense model; mixture-of-experts models use fewer. These are estimates from formulas, not measurements of a run.
Estimate a full run
Turn the compute column into time and cost with the training cost calculator. For the methods themselves, read LoRA and QLoRA, and see how fine-tuning works on Aquanode at fine-tuning.
Sources
- https://huggingface.co/allenai/Olmo-3-7B-Instruct
- https://huggingface.co/allenai/Olmo-3-7B-Think
- https://huggingface.co/allenai/Olmo-3-1025-7B
Updated 2026-10-07.