LLM Training Cost Calculator

Estimate training time and cost from model size, token count, and GPU choice, priced against Aquanode's live marketplace, per GPU / hr.

The math

Total FLOPs ≈ 6 × parameters × tokens (2 FLOPs/param/token forward + 4 backward, the standard dense transformer approximation; it overstates true FLOPs for mixture-of-experts models). Training time = total FLOPs ÷ (GPU's peak BF16 FLOP/s × assumed utilization), where peak throughput is the TFLOPS figure on the spec sheet. Cost = training time × the GPU's live $/hr rate.

Worked example: 7B params, 300B tokens, at assumed 40% utilization

GPURate (per GPU / hr)Estimated timeEstimated cost
V100$0.088/hr2916.7 days$6,160
RTX 4070 Super$0.121/hr3645.8 days$10,588
RTX A2000$0.132/hr3645.8 days$11,550
RTX 3070$0.143/hr3645.8 days$12,513
RTX 3060 TI$0.162/hr3645.8 days$14,149

All figures assume a single GPU: a real multi-GPU cluster trains faster but is not linearly modeled here (interconnect overhead varies by topology).

Try your own numbers

1.26e+22
Total training FLOPs
50 TFLOP/s
Effective throughput
2916.7 days
Estimated training time
V100
GPU used
$0.088/hr
Per-GPU rate
$6,160
Estimated cost (1 GPU)

Formula: FLOPs = 6 × parameters × tokens. Time = FLOPs ÷ (peak GPU FLOP/s × utilization). Cost = time × live per-GPU $/hr. Assumes a single logical accelerator, a real multi-GPU cluster scales throughput roughly linearly minus interconnect overhead, which is not modeled here.

Where the rates come from

Rates are the live on-demand per-GPU prices listed on the Aquanode pricing page. To see what a GPU costs while it sits unused, try the idle cost calculator, and for how rates change over time, read the GPU price report.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.