NVIDIA V100 GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA V100 (16GB or 32GB HBM2) launched in 2017. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the V100 rents from $0.088 per GPU per hour on Aquanode.
How much VRAM does the V100 have?
The V100 has 16GB or 32GB HBM2, with 900 GB/s of peak memory bandwidth.
- VRAM: V100 16GB or 32GB HBM2 vs A100 80GB HBM2e
- VRAM: V100 16GB or 32GB HBM2 vs T4 16GB GDDR6
What fits in 16GB or 32GB HBM2 of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3 8B | FP16 | roughly 16GB. Fits on the 32GB card, not the 16GB one. |
| Mistral 7B | FP16 | ~14GB. Fits on either card, tightly on the 16GB one. |
| Llama 3.1 70B | FP16 | ~140GB. Needs 5+ cards even at 32GB each. |
| Llama 3.1 70B | INT4 | ~35-40GB. The INT4 kernels do not run on Volta at all, so this is not an option regardless of VRAM. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the V100 run?
Popular open models from small to frontier scale, with the memory each needs and how many V100 cards (16GB or 32GB HBM2 each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen3-32B 32.8B | BF16: not supported | FP8: not supported | INT4: not supported |
| Qwen/Qwen-72B 72.3B | BF16: not supported | FP8: not supported | INT4: not supported |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: not supported | – | INT4: not supported |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: not supported | – | INT4: not supported |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the V100 VRAM calculator.
V100 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 16 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
V100 specs
| Architecture | NVIDIA, launched 2017 |
| VRAM | 16GB or 32GB HBM2 |
| Memory bandwidth | 900 GB/s |
| FP16 / BF16 tensor throughput | 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) (peak, dense) |
| Interconnect | NVLink, 300 GB/s (SXM2 only) |
| TDP | 300W (SXM2), 250W (PCIe) |
| Form factor | SXM2, PCIe, full height/length |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: A100 vs V100, and The best GPUs for AI, ranked.
GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe
V100 AI performance
HBM2 bandwidth (900 GB/s) at the cheapest hourly rate of any HBM card on this marketplace, and the SXM2 version has real NVLink at 300 GB/s. For FP16 training and classical HPC it is still a lot of memory bandwidth per dollar.
- Dense FP16/BF16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe)
- Memory bandwidth: 900 GB/s
No BF16, no FP8, and compute capability 7.0. Below the 7.5 floor AWQ/GPTQ/Marlin INT4 kernels require, so quantized serving does not run on it at all. Most current LLM training recipes assume BF16, which means a V100 needs an FP16 recipe with loss-scaling or it does not run.
See how it stacks up against other cards in the GPU benchmarks and specs table.
The cheapest V100 offer right now ($0.088/GPU/hr) is about 53% below the market median of $0.187/GPU/hr.
V100 price: what does it cost?
Renting. The cheapest current on-demand rate for the V100 on Aquanode is $0.088/GPU/hr (no on-demand supply right now; this is the cheapest offer overall). Live rates range from $0.088 to $0.253 per GPU per hour, with a median of $0.187/GPU/hr. Billed by the offer's own terms; the table below shows every live rate.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| United States | $0.088 | 2 | 16 GB | 6 | 30 GB |
| – | $0.209 | 0 | 16 GB | – | – |
| Washington, Us | $0.213 | 4 | 32 GB | 12 | 47.3 GB |
| Lis, Pt | $0.277 | 1 | 32 GB | 725 | 1774 GB |
V100 price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the V100 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate was $0.088 on both 2026-10-06 and 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $0.088 | $0.187 | – | 70 |
| 2026-10-09 | $0.187 | $0.209 | – | 36 |
| 2026-10-08 | $0.088 | $0.187 | – | 76 |
| 2026-10-07 | $0.088 | $0.187 | – | 74 |
| 2026-10-06 | $0.088 | $0.209 | – | 27 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
V100 vs A100: how do they compare?
- VRAM: V100 16GB or 32GB HBM2 vs A100 80GB HBM2e
- Memory bandwidth: 900 GB/s vs 2,039 GB/s
- Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 312 TFLOPS
On Aquanode right now, V100 starts at $0.088/GPU/hr against A100's $0.991/GPU/hr, about 91% less.
V100 vs T4: how do they compare?
- VRAM: V100 16GB or 32GB HBM2 vs T4 16GB GDDR6
- Memory bandwidth: 900 GB/s vs 320+ GB/s
- Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 65 TFLOPS
On Aquanode right now, V100 starts at $0.088/GPU/hr against T4's $0.185/GPU/hr, about 52% less.
Compare the V100 with other GPUs
Compare V100 with
- B200 vs V100
- H200 vs V100
- H100 vs V100
- A100 vs V100
- DGX A100 vs V100
- AMD MI300X vs V100
- L40S vs V100
- L40 vs V100
- A40 vs V100
- RTX A6000 vs V100
- RTX A5000 vs V100
- RTX A4000 vs V100
- L4 vs V100
- T4 vs V100
- RTX 6000 Ada vs V100
- RTX PRO 6000 vs V100
- RTX PRO 6000 WS vs V100
- RTX PRO 6000 SE vs V100
- RTX PRO 5000 vs V100
- RTX 5090 vs V100
- RTX 5080 vs V100
- RTX 5070 Ti vs V100
- RTX 5070 vs V100
- RTX 5060 Ti vs V100
- RTX 4090 vs V100
- RTX 4080 Super vs V100
- RTX 4080 vs V100
- RTX 4070 Ti vs V100
- RTX 4070 Super vs V100
- RTX 4070 vs V100
- RTX 4060 Ti vs V100
- RTX 3090 vs V100
- RTX 3080 vs V100
- RTX 3070 vs V100
- RTX 3060 vs V100
- B300 vs V100
- AMD MI355X vs V100
- AMD MI325X vs V100
- H100 NVL vs V100
- RTX PRO 4500 vs V100
- RTX PRO 4500 SE vs V100
- RTX PRO 4000 vs V100
- RTX 5880 Ada vs V100
- RTX 5000 Ada vs V100
- RTX 4000 Ada vs V100
- RTX 4000 SFF Ada vs V100
- RTX 2000 Ada vs V100
- RTX 4070 Ti Super vs V100
- RTX A4500 vs V100
- RTX 2080 Ti vs V100
- RTX 3080 Ti vs V100
- RTX 3070 Ti vs V100
- RTX 3060 Ti vs V100
- GTX 1080 Ti vs V100
- RTX 4060 vs V100
- RTX 5060 vs V100
- GTX 1660 Super vs V100
- RTX 2060 vs V100
- Radeon RX 9070 XT vs V100
- Radeon RX 6700 XT vs V100
- RTX 6000 vs V100
- A16 vs V100
- P40 vs V100
- P4 vs V100
- Quadro P2000 vs V100
- Quadro M4000 vs V100
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a V100 on Aquanode
- On-demand instances from $0.088/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 75 live offers across 4 regions today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available V100 offer appears.
- Compare every V100 offer side by side, or browse the full multi-provider GPU marketplace.
Good for
HBM2 bandwidth (900 GB/s) at the cheapest hourly rate of any HBM card on this marketplace, and the SXM2 version has real NVLink at 300 GB/s. For FP16 training and classical HPC it is still a lot of memory bandwidth per dollar.
Not good for
No BF16, no FP8, and compute capability 7.0. Below the 7.5 floor AWQ/GPTQ/Marlin INT4 kernels require, so quantized serving does not run on it at all. Most current LLM training recipes assume BF16, which means a V100 needs an FP16 recipe with loss-scaling or it does not run.
V100 FAQs
How much VRAM does the V100 have?
The V100 has 16GB or 32GB HBM2, with 900 GB/s of peak memory bandwidth.
What is the NVIDIA V100?
The NVIDIA V100 is a GPU released in 2017, with 16GB or 32GB HBM2 of memory and a 300W (SXM2), 250W (PCIe) power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a V100?
Live V100 rental prices currently range from $0.088 to $0.253 per GPU per hour, with a median of $0.187 per GPU per hour.
What's the cheapest V100 rate?
The lowest current V100 rate on Aquanode is $0.088 per GPU per hour in United States.
How does the V100 compare to the A100?
VRAM: V100 16GB or 32GB HBM2 vs A100 80GB HBM2e Memory bandwidth: 900 GB/s vs 2,039 GB/s Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 312 TFLOPS On Aquanode right now, V100 starts at $0.088/GPU/hr against A100's $0.991/GPU/hr, about 91% less. See the full V100 vs A100 comparison for a shared-provider price breakdown.
How does the V100 compare to the T4?
VRAM: V100 16GB or 32GB HBM2 vs T4 16GB GDDR6 Memory bandwidth: 900 GB/s vs 320+ GB/s Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 65 TFLOPS On Aquanode right now, V100 starts at $0.088/GPU/hr against T4's $0.185/GPU/hr, about 52% less. See the full V100 vs T4 comparison for a shared-provider price breakdown.
What is the V100 good for?
HBM2 bandwidth (900 GB/s) at the cheapest hourly rate of any HBM card on this marketplace, and the SXM2 version has real NVLink at 300 GB/s. For FP16 training and classical HPC it is still a lot of memory bandwidth per dollar.
What are the V100's limitations?
No BF16, no FP8, and compute capability 7.0. Below the 7.5 floor AWQ/GPTQ/Marlin INT4 kernels require, so quantized serving does not run on it at all. Most current LLM training recipes assume BF16, which means a V100 needs an FP16 recipe with loss-scaling or it does not run.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented V100 at $0.088/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting.
How is the V100 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
V100 price by region
Related guides
Other models in the same generation, then the rest of the GPU index.