LLM Training Cost Calculator
Estimate training time and cost from model size, token count, and GPU choice, priced against Aquanode's live marketplace, per GPU / hr.
The math
Total FLOPs ≈ 6 × parameters × tokens (2 FLOPs/param/token forward + 4 backward, the standard dense transformer approximation; it overstates true FLOPs for mixture-of-experts models). Training time = total FLOPs ÷ (GPU's peak BF16 FLOP/s × assumed utilization), where peak throughput is the TFLOPS figure on the spec sheet. Cost = training time × the GPU's live $/hr rate.
Worked example: 7B params, 300B tokens, at assumed 40% utilization
| GPU | Rate (per GPU / hr) | Estimated time | Estimated cost |
|---|---|---|---|
| V100 | $0.088/hr | 2916.7 days | $6,160 |
| RTX 4070 Super | $0.121/hr | 3645.8 days | $10,588 |
| RTX A2000 | $0.132/hr | 3645.8 days | $11,550 |
| RTX 3070 | $0.143/hr | 3645.8 days | $12,513 |
| RTX 3060 TI | $0.162/hr | 3645.8 days | $14,149 |
All figures assume a single GPU: a real multi-GPU cluster trains faster but is not linearly modeled here (interconnect overhead varies by topology).
Try your own numbers
Formula: FLOPs = 6 × parameters × tokens. Time = FLOPs ÷ (peak GPU FLOP/s × utilization). Cost = time × live per-GPU $/hr. Assumes a single logical accelerator, a real multi-GPU cluster scales throughput roughly linearly minus interconnect overhead, which is not modeled here.
Where the rates come from
Rates are the live on-demand per-GPU prices listed on the Aquanode pricing page. To see what a GPU costs while it sits unused, try the idle cost calculator, and for how rates change over time, read the GPU price report.