LLM Training Cost Calculator

Estimate training time and cost from model size, token count, and GPU choice — priced against Aquanode's live marketplace, per GPU / hr.

The math

Total FLOPs ≈ 6 × parameters × tokens (2 FLOPs/param/token forward + 4 backward — the standard dense transformer approximation; it overstates true FLOPs for mixture-of-experts models). Training time = total FLOPs ÷ (GPU's peak BF16 FLOP/s × assumed utilization). Cost = training time × the GPU's live $/hr rate.

Worked example: 7B params, 300B tokens, at assumed 40% utilization

GPU
Provider
Rate (per GPU / hr)
Estimated time
Estimated cost
A16
Vultr
$0.059/hr
3645.8 days
$5,163
V100
SimplePod
$0.070/hr
2916.7 days
$4,900
RTX 3060
SimplePod
$0.070/hr
3645.8 days
$6,125
A40
Vultr
$0.075/hr
3645.8 days
$6,563
RTX 4070
SimplePod
$0.080/hr
3645.8 days
$7,000

All figures assume a single GPU — a real multi-GPU cluster trains faster but is not linearly modeled here (interconnect overhead varies by topology).

Try your own numbers

1.26e+22
Total training FLOPs
40 TFLOP/s
Effective throughput
3645.8 days
Estimated training time
A16 (Vultr)
GPU used
$0.059/hr
Per-GPU rate
$5,163
Estimated cost (1 GPU)

Formula: FLOPs = 6 × parameters × tokens. Time = FLOPs ÷ (peak GPU FLOP/s × utilization). Cost = time × live per-GPU $/hr. Assumes a single logical accelerator — a real multi-GPU cluster scales throughput roughly linearly minus interconnect overhead, which is not modeled here.

Ready when you are

Stop paying for
idle GPUs.

Sign up in 60 seconds. Pay only for the GPU minutes you actually use.

Aquanode LogoAquanode

Your GPU environment, preserved. Pause it, move it, come back to it.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.