NVIDIA V100 GPU: Specs, VRAM, Price & Benchmarks (2026)

The NVIDIA V100 (16GB or 32GB HBM2) launched in 2017. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.

Short answer on cost: the V100 rents from $0.088 per GPU per hour on Aquanode.

How much VRAM does the V100 have?

The V100 has 16GB or 32GB HBM2, with 900 GB/s of peak memory bandwidth.

  • VRAM: V100 16GB or 32GB HBM2 vs A100 80GB HBM2e
  • VRAM: V100 16GB or 32GB HBM2 vs T4 16GB GDDR6

What fits in 16GB or 32GB HBM2 of VRAM

ModelPrecisionFits?
Llama 3 8BFP16roughly 16GB. Fits on the 32GB card, not the 16GB one.
Mistral 7BFP16~14GB. Fits on either card, tightly on the 16GB one.
Llama 3.1 70BFP16~140GB. Needs 5+ cards even at 32GB each.
Llama 3.1 70BINT4~35-40GB. The INT4 kernels do not run on Volta at all, so this is not an option regardless of VRAM.

Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.

What can the V100 run?

Popular open models from small to frontier scale, with the memory each needs and how many V100 cards (16GB or 32GB HBM2 each) that takes.

ModelAs publishedFP8INT4
Qwen/Qwen3-8B 8.2BBF16: not supportedFP8: not supportedINT4: not supported
Qwen/Qwen2.5-14B-Instruct 14.8BBF16: not supportedFP8: not supportedINT4: not supported
Qwen/Qwen3-32B 32.8BBF16: not supportedFP8: not supportedINT4: not supported
Qwen/Qwen-72B 72.3BBF16: not supportedFP8: not supportedINT4: not supported
MiniMaxAI/MiniMax-M2.7 228.7BFP8: not supported–INT4: not supported
deepseek-ai/DeepSeek-R1 684.5BFP8: not supported–INT4: not supported

Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the V100 VRAM calculator.

V100 VRAM calculator: check which models fit in its memory at each precision.

All models that fit in 16 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.

$0.088/GPU/hr
Lowest / GPU / hr
$0.187/GPU/hr
Median / GPU / hr
$0.253/GPU/hr
p90 / GPU / hr
75
Live offers
Last updated: 2026-10-10 21:05:01 UTCRefreshes hourly0 offer(s) excluded from this snapshot

V100 specs

ArchitectureNVIDIA, launched 2017
VRAM16GB or 32GB HBM2
Memory bandwidth900 GB/s
FP16 / BF16 tensor throughput125 TFLOPS (SXM2), 112 TFLOPS (PCIe) (peak, dense)
InterconnectNVLink, 300 GB/s (SXM2 only)
TDP300W (SXM2), 250W (PCIe)
Form factorSXM2, PCIe, full height/length

Specs sourced from the vendor's public datasheet/product page. See the source.

Related reading: A100 vs V100, and The best GPUs for AI, ranked.

GPU Glossary: What is VRAM?, HBM, Tensor Cores, CUDA Cores, TFLOPS, NVLink vs PCIe

V100 AI performance

HBM2 bandwidth (900 GB/s) at the cheapest hourly rate of any HBM card on this marketplace, and the SXM2 version has real NVLink at 300 GB/s. For FP16 training and classical HPC it is still a lot of memory bandwidth per dollar.

  • Dense FP16/BF16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe)
  • Memory bandwidth: 900 GB/s

No BF16, no FP8, and compute capability 7.0. Below the 7.5 floor AWQ/GPTQ/Marlin INT4 kernels require, so quantized serving does not run on it at all. Most current LLM training recipes assume BF16, which means a V100 needs an FP16 recipe with loss-scaling or it does not run.

See how it stacks up against other cards in the GPU benchmarks and specs table.

The cheapest V100 offer right now ($0.088/GPU/hr) is about 53% below the market median of $0.187/GPU/hr.

V100 price: what does it cost?

Renting. The cheapest current on-demand rate for the V100 on Aquanode is $0.088/GPU/hr (no on-demand supply right now; this is the cheapest offer overall). Live rates range from $0.088 to $0.253 per GPU per hour, with a median of $0.187/GPU/hr. Billed by the offer's own terms; the table below shows every live rate.

Region$/GPU/hrAvailableVRAMvCPURAM
United States$0.088216 GB630 GB
–$0.209016 GB––
Washington, Us$0.213432 GB1247.3 GB
Lis, Pt$0.277132 GB7251774 GB

V100 price history

Aquanode stores one snapshot of its GPU price index per UTC day. For the V100 that is 5 days so far, 2026-10-06 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate was $0.088 on both 2026-10-06 and 2026-10-10.

Day (UTC)Lowest per GPU hourMedian per GPU hourData-center lowest per GPU hourOffers
2026-10-10$0.088$0.187–70
2026-10-09$0.187$0.209–36
2026-10-08$0.088$0.187–76
2026-10-07$0.088$0.187–74
2026-10-06$0.088$0.209–27

Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.

How this price is calculated

All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.

No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.

Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.

V100 vs A100: how do they compare?

  • VRAM: V100 16GB or 32GB HBM2 vs A100 80GB HBM2e
  • Memory bandwidth: 900 GB/s vs 2,039 GB/s
  • Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 312 TFLOPS

On Aquanode right now, V100 starts at $0.088/GPU/hr against A100's $0.991/GPU/hr, about 91% less.

Full V100 vs A100 price comparison

V100 vs T4: how do they compare?

  • VRAM: V100 16GB or 32GB HBM2 vs T4 16GB GDDR6
  • Memory bandwidth: 900 GB/s vs 320+ GB/s
  • Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 65 TFLOPS

On Aquanode right now, V100 starts at $0.088/GPU/hr against T4's $0.185/GPU/hr, about 52% less.

Full V100 vs T4 price comparison

Compare the V100 with other GPUs

Compare V100 with

Get notified when the price drops

GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.

One email per matching alert. Unsubscribe any time.

Rent a V100 on Aquanode

  • On-demand instances from $0.088/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
  • 75 live offers across 4 regions today.
  • Set a price/availability alert above to hear the moment a cheaper or newly-available V100 offer appears.
  • Compare every V100 offer side by side, or browse the full multi-provider GPU marketplace.

Good for

HBM2 bandwidth (900 GB/s) at the cheapest hourly rate of any HBM card on this marketplace, and the SXM2 version has real NVLink at 300 GB/s. For FP16 training and classical HPC it is still a lot of memory bandwidth per dollar.

Not good for

No BF16, no FP8, and compute capability 7.0. Below the 7.5 floor AWQ/GPTQ/Marlin INT4 kernels require, so quantized serving does not run on it at all. Most current LLM training recipes assume BF16, which means a V100 needs an FP16 recipe with loss-scaling or it does not run.

V100 FAQs

How much VRAM does the V100 have?

The V100 has 16GB or 32GB HBM2, with 900 GB/s of peak memory bandwidth.

What is the NVIDIA V100?

The NVIDIA V100 is a GPU released in 2017, with 16GB or 32GB HBM2 of memory and a 300W (SXM2), 250W (PCIe) power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.

How much does it cost to rent a V100?

Live V100 rental prices currently range from $0.088 to $0.253 per GPU per hour, with a median of $0.187 per GPU per hour.

What's the cheapest V100 rate?

The lowest current V100 rate on Aquanode is $0.088 per GPU per hour in United States.

How does the V100 compare to the A100?

VRAM: V100 16GB or 32GB HBM2 vs A100 80GB HBM2e Memory bandwidth: 900 GB/s vs 2,039 GB/s Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 312 TFLOPS On Aquanode right now, V100 starts at $0.088/GPU/hr against A100's $0.991/GPU/hr, about 91% less. See the full V100 vs A100 comparison for a shared-provider price breakdown.

How does the V100 compare to the T4?

VRAM: V100 16GB or 32GB HBM2 vs T4 16GB GDDR6 Memory bandwidth: 900 GB/s vs 320+ GB/s Dense FP16 tensor throughput: 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) vs 65 TFLOPS On Aquanode right now, V100 starts at $0.088/GPU/hr against T4's $0.185/GPU/hr, about 52% less. See the full V100 vs T4 comparison for a shared-provider price breakdown.

What is the V100 good for?

HBM2 bandwidth (900 GB/s) at the cheapest hourly rate of any HBM card on this marketplace, and the SXM2 version has real NVLink at 300 GB/s. For FP16 training and classical HPC it is still a lot of memory bandwidth per dollar.

What are the V100's limitations?

No BF16, no FP8, and compute capability 7.0. Below the 7.5 floor AWQ/GPTQ/Marlin INT4 kernels require, so quantized serving does not run on it at all. Most current LLM training recipes assume BF16, which means a V100 needs an FP16 recipe with loss-scaling or it does not run.

Is renting cheaper than buying?

Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented V100 at $0.088/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting.

How is the V100 price calculated?

All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.

V100 price by region

Related guides

Other models in the same generation, then the rest of the GPU index.

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.