V100 vs A100: Specs and Is the V100 Worth It in 2026

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

The NVIDIA V100 is a 2017-era Volta data center GPU with 640 Tensor Cores, 16 or 32 GB of HBM2, 900 GB/s of memory bandwidth and up to 125 TFLOPS of Tensor performance. The A100 that followed it has up to 80 GB, about 2.3 times the memory bandwidth of the V100 (2,039 against 900 GB/s) and about 2.5 times the dense FP16 Tensor throughput (312 against 125 TFLOPS) on NVIDIA's datasheets. In 2026 the V100 is hard to recommend for new work, mainly because of its memory size and the end of current CUDA toolkit support, covered below.

TL;DR

  • V100 (Volta): 5,120 CUDA cores, 640 Tensor Cores, 16 GB or 32 GB HBM2 at 900 GB/s, 300 GB/s NVLink, 250 to 300 W.
  • A100 (Ampere): up to 80 GB HBM2e at up to 2,039 GB/s, 600 GB/s NVLink, TF32 and BF16 support, and MIG partitioning into up to 7 instances.
  • Gap on paper: about 2.3x memory bandwidth, 2.5x dense FP16 Tensor throughput (312 vs 125 TFLOPS), 2x NVLink, and 2.5x to 5x the memory capacity.
  • Software: NVIDIA removed Volta support from the CUDA 13.0 toolkit; CUDA 12.x remains the path for V100.
  • Verdict: use a V100 only if you already have one or need a very cheap box for small models in FP16 on CUDA 12. For anything new, start at A100 or newer.

What is the V100?

The Tesla V100 was NVIDIA's flagship data center GPU when it launched with the Volta architecture. NVIDIA's datasheet (March 2018 revision) describes it as built to accelerate AI, HPC and graphics, and introduces Tensor Cores: 640 of them, delivering "125 teraFLOPS of deep learning performance," which NVIDIA says is 12x the Tensor FLOPS for deep learning training and 6x for inference compared with its Pascal GPUs.

It came in two form factors: V100 PCIe (250 W, 32 GB/s PCIe Gen3) and V100 SXM2 (300 W, NVLink at 300 GB/s). Both were offered with 16 GB or 32 GB of HBM2. The 32 GB version "doubles the memory of the standard 16GB offering," in NVIDIA's words. If you see "v100" in a listing, check which of these four combinations it is.

V100 specs

From NVIDIA's V100 datasheet:

SpecV100 PCIeV100 SXM2
ArchitectureVoltaVolta
CUDA cores5,1205,120
Tensor Cores640640
FP647 TFLOPS7.8 TFLOPS
FP3214 TFLOPS15.7 TFLOPS
Tensor performance112 TFLOPS125 TFLOPS
Memory16 GB or 32 GB HBM216 GB or 32 GB HBM2
Memory bandwidth900 GB/s900 GB/s
InterconnectPCIe Gen3, 32 GB/sNVLink, 300 GB/s
Max power250 W300 W

The datasheet lists a single Tensor figure per card. It does not list TF32, BF16, FP8 or INT8 Tensor rates or structured sparsity, because those came later.

V100 vs A100 spec comparison

The A100 column is the 80 GB SXM version from NVIDIA's A100 page. The original 40 GB A100 had 1,555 GB/s of bandwidth, per NVIDIA's Hopper architecture blog.

SpecV100 SXM2A100 80GB SXM
ArchitectureVoltaAmpere
Memory16 or 32 GB HBM280 GB HBM2e
Memory bandwidth900 GB/s2,039 GB/s
FP647.8 TFLOPS9.7 TFLOPS (PCIe listing)
FP64 Tensor Corenot listed19.5 TFLOPS (PCIe listing)
FP3215.7 TFLOPS19.5 TFLOPS (per NVIDIA's Hopper blog)
FP16 Tensor Core125 TFLOPS312 dense, 624 with sparsity
BF16 Tensor Corenot listed312 dense, 624 with sparsity
TF32 Tensor Corenot listed156 dense, 312 with sparsity
INT8 Tensor Corenot listed624 dense, 1,248 with sparsity
NVLink300 GB/s600 GB/s
PCIeGen3, 32 GB/sGen4, 64 GB/s
MIGnoup to 7 instances at 10 GB
Max power300 W400 W

"Not listed" means the V100 datasheet does not give a figure; we are not saying the capability is absent in every sense. Two notes on the A100 rows: NVIDIA prints its A100 Tensor numbers with sparsity and the dense value is half, and the A100 page does not give an FP64 figure for the SXM version separately, so we show the PCIe listing. MIG does not appear on the V100 datasheet, which is why we say "no" there; confirm for your exact system. Our ratios in the intro (2.3x bandwidth, 2.5x dense FP16) are arithmetic on these listed peaks, not measured speedups.

For a live side-by-side on this site, see A100 vs V100, plus /gpu/nvidia-a100 and /gpu/nvidia-v100.

What the A100 added

  • More memory, faster. Up to 80 GB against 32 GB at most, and 2,039 GB/s against 900 GB/s. Capacity decides which models fit; see VRAM and HBM.
  • TF32 and BF16. The A100 page lists TF32, BF16, FP16 and INT8 Tensor Core rates. BF16 is the common training format for current LLM code. NVIDIA says TF32 delivers "up to 20X higher performance over the NVIDIA Volta" with "zero code changes," though the A100 page does not name the exact V100 baseline for the headline. Read it as NVIDIA's best case for a TF32 workload.
  • Structured sparsity. The A100's higher peak figures assume sparsity. Most users run dense, so the dense column is the fair comparison.
  • MIG. One A100 can be split into as many as seven isolated instances, which is useful for sharing a GPU across small inference jobs. See tensor cores for background on the compute units.
  • Faster links. NVLink doubles from 300 to 600 GB/s, and PCIe goes from Gen3 to Gen4.

Is the V100 still worth using in 2026?

Short answer: rarely, for new work. The reasons come straight from the specs and NVIDIA's software notes.

  1. Memory. The largest V100 has 32 GB. That is not enough for modern 70B-class models even when quantized, and it constrains batch size and context length for mid-sized ones. An A100 80 GB fits far more. Compare with the H100 vs A100 post for the generation after.
  2. Precision. The V100 datasheet lists FP16 Tensor performance only. Current training recipes often assume BF16, which the A100 and later list. FP8 and FP4 are further away: FP8 arrived with Hopper (Transformer Engine and FP8) and FP4 with Blackwell (NVFP4 vs MXFP4).
  3. Toolkit support. NVIDIA's CUDA features archive records that CUDA 13.0 removed offline compilation and library support for Maxwell, Pascal and Volta. You can keep building for the V100 with CUDA 12.x toolkits, which NVIDIA continues to document, but new framework releases will increasingly assume newer toolkits. That is an ecosystem risk, not an immediate break.
  4. Bandwidth. At 900 GB/s, LLM decoding on a V100 is slow compared with the A100's 2,039 GB/s, because decoding is largely limited by memory bandwidth.

When a V100 still works:

  • Classic computer vision, small transformer fine-tunes, or research code already written for FP16 on CUDA 11 or 12.
  • You are learning or prototyping and the lowest possible hourly rate matters more than throughput.
  • You already own the hardware and the workload fits in 16 or 32 GB.

When to move on:

Cost: measure throughput, then price it

We do not type rental prices. A fair comparison is cost per job. Run a fixed, short workload on each GPU, record steps per second or tokens per second, and divide the live hourly price by it. NVIDIA's peak ratios suggest an A100 can do roughly 2.3x to 2.5x the work of a V100 on bandwidth-bound and dense FP16 work. If an A100 costs less than that multiple of a V100 per hour, the A100 should cost less per job; if your model does not fit on the V100 at all, the comparison is moot.

Rent today

Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; "None right now" means there is no offer at the moment.

What's next

Ampere led to Hopper and Blackwell. For the current range, see H200 vs B200 vs GB200 and, further out, Rubin vs Blackwell vs Hopper.

FAQ

What is the NVIDIA V100?

A Volta-architecture data center GPU with 5,120 CUDA cores, 640 Tensor Cores and 16 or 32 GB of HBM2 at 900 GB/s, offered as a PCIe card (250 W) and an SXM2 module (300 W).

How much faster is the A100 than the V100?

On NVIDIA's datasheet peaks, about 2.3x the memory bandwidth and 2.5x the dense FP16 Tensor throughput. Measured gains depend on whether your job is compute bound, bandwidth bound or limited elsewhere.

How much memory does a V100 have?

16 GB or 32 GB of HBM2, depending on the configuration.

Does the V100 support BF16?

NVIDIA's V100 datasheet does not list BF16 Tensor Core support. The A100 page lists BF16 at 312 TFLOPS dense (624 with sparsity).

Does CUDA still support the V100?

CUDA 12.x toolkits do. NVIDIA's CUDA features archive says CUDA 13.0 removed support for Volta, so you cannot target it with CUDA 13 or later toolkits.

Is the V100 good for LLMs?

For small models in FP16 that fit in 16 or 32 GB, it can work. For larger models, long context or BF16, the A100 or newer parts fit better.

Sources

#datacenter gpu#nvidia hopper#nvidia ampere#v100#a100 vs v100#nvidia volta

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.