Best open models for coding
How to pick an open coding model: instruct vs base, context for repo-scale work, license, and size. Picks by GPU budget with facts from model cards.
The best open coding model for you depends on three things: whether it follows instructions or only completes code, how much context it can read, and how many parameters your GPU can hold. Everything else is a tiebreaker. The catalog below the picks lists every coder model, so this page only covers how to choose.
What matters when picking a coding model
Instruct or base. An instruct model takes a request like "fix this function" and answers it, which is what chat assistants and agent tools need. A base model only continues text. The StarCoder2 card says plainly that it is not an instruction model and that a command such as "Write a function that computes the square root" will not work as a prompt. Base coders still fit editor autocomplete, where the model fills in code around your cursor.
Context length. Repo-scale work means feeding many files at once. A longer window lets the model see more of the project, but the KV cache grows with it, so long contexts cost memory on top of the weights.
Dense or mixture-of-experts. A mixture-of-experts coder stores all of its weights but only runs a fraction per token, so it answers faster than a dense model of the same total size. You still need memory for all the weights.
Precision. Weights in BF16 are the reference. A quantized version trades a little accuracy for a smaller footprint. You can estimate the fit for your card in the VRAM calculator.
License. Check it before shipping. Apache 2.0 has no use restrictions of its own. Some coders ship under licenses with usage clauses.
Picks by situation
Smaller GPUs (a 24 GB card): Qwen2.5-Coder-7B-Instruct. The card lists 7.61B parameters, a 131,072-token context, and an Apache 2.0 license. Qwen says the Qwen2.5-Coder series improves code generation, code reasoning, and code fixing. It also comes in 0.5B to 32B sizes, so you can move up or down within one family.
One 80 GB GPU: Qwen2.5-Coder-32B-Instruct. The card lists 32.5B parameters, a 131,072-token context, and Apache 2.0. Qwen describes it as the state-of-the-art open-source code model of the series.
Agentic coding and the longest context: Qwen3-Coder-30B-A3B-Instruct. It is a mixture-of-experts model with 30.5B parameters in total and 3.3B activated, under Apache 2.0. The card reports 262,144 tokens natively, extendable to 1M tokens with YaRN, and says it targets agentic coding platforms such as Qwen Code and Cline with a dedicated function-call format. The card also notes that this model supports only non-thinking mode.
Editor autocomplete: StarCoder2-7B. A 7B base model trained on 17 programming languages from The Stack v2, with a fill-in-the-middle objective, which is what autocomplete uses. Its license is BigCode OpenRAIL-M v1, so read the agreement before commercial use.
Most permissive license: the Qwen coders above are all Apache 2.0, per their model cards. StarCoder2 is the one with extra terms.
Serving notes
Coding workloads are often interactive, so first-token latency matters as much as throughput. The mixture-of-experts pick above is built for that trade. For agent loops with long prompts, budget for the KV cache rather than only the weights.
Aquanode rents GPUs by the hour, so you can run any of these on the card size they need and stop when you are done. See pricing for the current options.
Quality claims on model cards are vendor-reported on the vendor's own settings. Test two candidates on your own codebase before committing.
Open models for coding
All 32 models in the catalog for this task, grouped by size.
Under 3B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen2.5-Coder-1.5B-Instruct | 1.5B | 3.5 GB | – | RTX 4070 Super | $0.121/hr |
| Qwen2.5-Coder-1.5B | 1.5B | 3.5 GB | – | RTX 4070 Super | $0.121/hr |
| Qwen2.5-Coder-0.5B-Instruct | 494M | 1.1 GB | – | RTX 4070 Super | $0.121/hr |
| gpt_bigcode-santacoder | 1.1B | 2.5 GB | – | V100 | $0.088/hr |
| deepseek-coder-1.3b-instruct | 1.3B | 3.0 GB | – | RTX 4070 Super | $0.121/hr |
| Qwen2.5-Coder-0.5B | 494M | 1.1 GB | – | RTX 4070 Super | $0.121/hr |
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen2.5-Coder-7B-Instruct | 7.6B | 17.0 GB | – | RTX A5000 | $0.176/hr |
| deepseek-coder-7b-instruct-v1.5 | 6.9B | 15.4 GB | – | RTX A4000 | $0.167/hr |
| Qwen2.5-Coder-7B | 7.6B | 17.0 GB | – | RTX A5000 | $0.176/hr |
| CodeLlama-7b-hf | 6.7B | 15.1 GB | – | RTX A4000 | $0.167/hr |
| deepseek-coder-6.7b-instruct | 6.7B | 15.1 GB | Custom license | RTX A4000 | $0.167/hr |
| Qwen2.5-Coder-3B-Instruct | 3.1B | 6.9 GB | – | RTX 4070 Super | $0.121/hr |
| starcoder2-3b | 3.0B | 13.5 GB | – | V100 | $0.088/hr |
| deepseek-coder-6.7b-base | 6.7B | 15.1 GB | – | RTX A4000 | $0.167/hr |
| CodeLlama-7b-Instruct-hf | 6.7B | 15.1 GB | – | RTX A4000 | $0.167/hr |
| starcoder2-7b | 7.2B | 16.0 GB | – | RTX A5000 | $0.176/hr |
| granite-3b-code-base-2k | 3.5B | 7.8 GB | – | RTX 4070 Super | $0.121/hr |
| Qwen2.5-Coder-3B | 3.1B | 6.9 GB | – | RTX 4070 Super | $0.121/hr |
| sqlcoder-7b-2 | 6.7B | 15.1 GB | – | V100 | $0.088/hr |
| codegeex4-all-9b | 9.4B | 21.0 GB | – | RTX A5000 | $0.176/hr |
10B to 40B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen2.5-Coder-14B-Instruct | 14.8B | 33.0 GB | – | RTX A6000 | $0.363/hr |
| Qwen2.5-Coder-32B-Instruct | 32.8B | 73.2 GB | Apache 2.0 | A100 | $1.21/hr |
| Qwen3-Coder-30B-A3B-Instruct | 30.5B | 68.2 GB | – | A100 | $1.21/hr |
| DeepSeek-Coder-V2-Lite-Instruct | 15.7B | 35.1 GB | – | RTX A6000 | $0.363/hr |
| Qwen2.5-Coder-14B | 14.8B | 33.0 GB | – | RTX A6000 | $0.363/hr |
| Laguna-XS-2.1 | 33.4B | 74.8 GB | – | A100 | $1.21/hr |
| KAT-Coder-V2.5-Dev | 34.7B | 77.5 GB | – | A100 | $1.21/hr |
| Laguna-XS.2 | 33.4B | 74.8 GB | – | A100 | $1.21/hr |
| CodeLlama-34b-Instruct-hf | 33.7B | 75.4 GB | – | A100 | $1.21/hr |
40B to 150B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen3-Coder-Next | 79.7B | 178 GB | Apache 2.0 | RTX A5000 × 8 | $1.41/hr |
| Laguna-S-2.1 | 117.6B | 263 GB | – | RTX A6000 × 6 | $2.18/hr |
150B parameters and up
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Qwen3-Coder-480B-A35B-Instruct | 480.2B | 1073 GB | Apache 2.0 | No live fit | – |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
Sources
- https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct
- https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct
- https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct
- https://huggingface.co/bigcode/starcoder2-7b
Updated 2026-10-07.
More ways to choose a model
- Best open models for reasoning
- Best open models for chat and assistants
- Best open models for vision-language
- Best open models for OCR and document parsing
- Best open models for speech-to-text
- Best open models for text-to-speech
- Best open models for image generation
- Best open models for video generation
- Best open models for embeddings
- Best open models for translation
By GPU memory: