The NVIDIA V100 is a 2017-era Volta data center GPU with 640 Tensor Cores, 16 or 32 GB of HBM2, 900 GB/s of memory bandwidth and up to 125 TFLOPS of Tensor performance. The A100 that followed it has up to 80 GB, about 2.3 times the memory bandwidth of the V100 (2,039 against 900 GB/s) and about 2.5 times the dense FP16 Tensor throughput (312 against 125 TFLOPS) on NVIDIA's datasheets. In 2026 the V100 is hard to recommend for new work, mainly because of its memory size and the end of current CUDA toolkit support, covered below.
TL;DR
- V100 (Volta): 5,120 CUDA cores, 640 Tensor Cores, 16 GB or 32 GB HBM2 at 900 GB/s, 300 GB/s NVLink, 250 to 300 W.
- A100 (Ampere): up to 80 GB HBM2e at up to 2,039 GB/s, 600 GB/s NVLink, TF32 and BF16 support, and MIG partitioning into up to 7 instances.
- Gap on paper: about 2.3x memory bandwidth, 2.5x dense FP16 Tensor throughput (312 vs 125 TFLOPS), 2x NVLink, and 2.5x to 5x the memory capacity.
- Software: NVIDIA removed Volta support from the CUDA 13.0 toolkit; CUDA 12.x remains the path for V100.
- Verdict: use a V100 only if you already have one or need a very cheap box for small models in FP16 on CUDA 12. For anything new, start at A100 or newer.
What is the V100?
The Tesla V100 was NVIDIA's flagship data center GPU when it launched with the Volta architecture. NVIDIA's datasheet (March 2018 revision) describes it as built to accelerate AI, HPC and graphics, and introduces Tensor Cores: 640 of them, delivering "125 teraFLOPS of deep learning performance," which NVIDIA says is 12x the Tensor FLOPS for deep learning training and 6x for inference compared with its Pascal GPUs.
It came in two form factors: V100 PCIe (250 W, 32 GB/s PCIe Gen3) and V100 SXM2 (300 W, NVLink at 300 GB/s). Both were offered with 16 GB or 32 GB of HBM2. The 32 GB version "doubles the memory of the standard 16GB offering," in NVIDIA's words. If you see "v100" in a listing, check which of these four combinations it is.
V100 specs
From NVIDIA's V100 datasheet:
| Spec | V100 PCIe | V100 SXM2 |
|---|---|---|
| Architecture | Volta | Volta |
| CUDA cores | 5,120 | 5,120 |
| Tensor Cores | 640 | 640 |
| FP64 | 7 TFLOPS | 7.8 TFLOPS |
| FP32 | 14 TFLOPS | 15.7 TFLOPS |
| Tensor performance | 112 TFLOPS | 125 TFLOPS |
| Memory | 16 GB or 32 GB HBM2 | 16 GB or 32 GB HBM2 |
| Memory bandwidth | 900 GB/s | 900 GB/s |
| Interconnect | PCIe Gen3, 32 GB/s | NVLink, 300 GB/s |
| Max power | 250 W | 300 W |
The datasheet lists a single Tensor figure per card. It does not list TF32, BF16, FP8 or INT8 Tensor rates or structured sparsity, because those came later.
V100 vs A100 spec comparison
The A100 column is the 80 GB SXM version from NVIDIA's A100 page. The original 40 GB A100 had 1,555 GB/s of bandwidth, per NVIDIA's Hopper architecture blog.
| Spec | V100 SXM2 | A100 80GB SXM |
|---|---|---|
| Architecture | Volta | Ampere |
| Memory | 16 or 32 GB HBM2 | 80 GB HBM2e |
| Memory bandwidth | 900 GB/s | 2,039 GB/s |
| FP64 | 7.8 TFLOPS | 9.7 TFLOPS (PCIe listing) |
| FP64 Tensor Core | not listed | 19.5 TFLOPS (PCIe listing) |
| FP32 | 15.7 TFLOPS | 19.5 TFLOPS (per NVIDIA's Hopper blog) |
| FP16 Tensor Core | 125 TFLOPS | 312 dense, 624 with sparsity |
| BF16 Tensor Core | not listed | 312 dense, 624 with sparsity |
| TF32 Tensor Core | not listed | 156 dense, 312 with sparsity |
| INT8 Tensor Core | not listed | 624 dense, 1,248 with sparsity |
| NVLink | 300 GB/s | 600 GB/s |
| PCIe | Gen3, 32 GB/s | Gen4, 64 GB/s |
| MIG | no | up to 7 instances at 10 GB |
| Max power | 300 W | 400 W |
"Not listed" means the V100 datasheet does not give a figure; we are not saying the capability is absent in every sense. Two notes on the A100 rows: NVIDIA prints its A100 Tensor numbers with sparsity and the dense value is half, and the A100 page does not give an FP64 figure for the SXM version separately, so we show the PCIe listing. MIG does not appear on the V100 datasheet, which is why we say "no" there; confirm for your exact system. Our ratios in the intro (2.3x bandwidth, 2.5x dense FP16) are arithmetic on these listed peaks, not measured speedups.
For a live side-by-side on this site, see A100 vs V100, plus /gpu/nvidia-a100 and /gpu/nvidia-v100.
What the A100 added
- More memory, faster. Up to 80 GB against 32 GB at most, and 2,039 GB/s against 900 GB/s. Capacity decides which models fit; see VRAM and HBM.
- TF32 and BF16. The A100 page lists TF32, BF16, FP16 and INT8 Tensor Core rates. BF16 is the common training format for current LLM code. NVIDIA says TF32 delivers "up to 20X higher performance over the NVIDIA Volta" with "zero code changes," though the A100 page does not name the exact V100 baseline for the headline. Read it as NVIDIA's best case for a TF32 workload.
- Structured sparsity. The A100's higher peak figures assume sparsity. Most users run dense, so the dense column is the fair comparison.
- MIG. One A100 can be split into as many as seven isolated instances, which is useful for sharing a GPU across small inference jobs. See tensor cores for background on the compute units.
- Faster links. NVLink doubles from 300 to 600 GB/s, and PCIe goes from Gen3 to Gen4.
Is the V100 still worth using in 2026?
Short answer: rarely, for new work. The reasons come straight from the specs and NVIDIA's software notes.
- Memory. The largest V100 has 32 GB. That is not enough for modern 70B-class models even when quantized, and it constrains batch size and context length for mid-sized ones. An A100 80 GB fits far more. Compare with the H100 vs A100 post for the generation after.
- Precision. The V100 datasheet lists FP16 Tensor performance only. Current training recipes often assume BF16, which the A100 and later list. FP8 and FP4 are further away: FP8 arrived with Hopper (Transformer Engine and FP8) and FP4 with Blackwell (NVFP4 vs MXFP4).
- Toolkit support. NVIDIA's CUDA features archive records that CUDA 13.0 removed offline compilation and library support for Maxwell, Pascal and Volta. You can keep building for the V100 with CUDA 12.x toolkits, which NVIDIA continues to document, but new framework releases will increasingly assume newer toolkits. That is an ecosystem risk, not an immediate break.
- Bandwidth. At 900 GB/s, LLM decoding on a V100 is slow compared with the A100's 2,039 GB/s, because decoding is largely limited by memory bandwidth.
When a V100 still works:
- Classic computer vision, small transformer fine-tunes, or research code already written for FP16 on CUDA 11 or 12.
- You are learning or prototyping and the lowest possible hourly rate matters more than throughput.
- You already own the hardware and the workload fits in 16 or 32 GB.
When to move on:
- Your model needs more than 32 GB, or you want BF16, or you want to use current inference stacks. Start with the A100 and look at newer chips in A100 vs H100 and H200 vs B200 vs GB200. The full range is in the datacenter GPU guide.
Cost: measure throughput, then price it
We do not type rental prices. A fair comparison is cost per job. Run a fixed, short workload on each GPU, record steps per second or tokens per second, and divide the live hourly price by it. NVIDIA's peak ratios suggest an A100 can do roughly 2.3x to 2.5x the work of a V100 on bandwidth-bound and dense FP16 work. If an A100 costs less than that multiple of a V100 per hour, the A100 should cost less per job; if your model does not fit on the V100 at all, the comparison is moot.
Rent today
Aquanode manages and optimizes GPUs for training and inference workloads, and you can rent the GPUs below on demand. The box shows live availability; "None right now" means there is no offer at the moment.
What's next
Ampere led to Hopper and Blackwell. For the current range, see H200 vs B200 vs GB200 and, further out, Rubin vs Blackwell vs Hopper.
FAQ
What is the NVIDIA V100?
A Volta-architecture data center GPU with 5,120 CUDA cores, 640 Tensor Cores and 16 or 32 GB of HBM2 at 900 GB/s, offered as a PCIe card (250 W) and an SXM2 module (300 W).
How much faster is the A100 than the V100?
On NVIDIA's datasheet peaks, about 2.3x the memory bandwidth and 2.5x the dense FP16 Tensor throughput. Measured gains depend on whether your job is compute bound, bandwidth bound or limited elsewhere.
How much memory does a V100 have?
16 GB or 32 GB of HBM2, depending on the configuration.
Does the V100 support BF16?
NVIDIA's V100 datasheet does not list BF16 Tensor Core support. The A100 page lists BF16 at 312 TFLOPS dense (624 with sparsity).
Does CUDA still support the V100?
CUDA 12.x toolkits do. NVIDIA's CUDA features archive says CUDA 13.0 removed support for Volta, so you cannot target it with CUDA 13 or later toolkits.
Is the V100 good for LLMs?
For small models in FP16 that fit in 16 or 32 GB, it can work. For larger models, long context or BF16, the A100 or newer parts fit better.
Sources
- NVIDIA Tesla V100 datasheet (March 2018): https://images.nvidia.com/content/technologies/volta/pdf/tesla-volta-v100-datasheet-letter-fnl-web.pdf
- NVIDIA A100 product page: https://www.nvidia.com/en-us/data-center/a100/
- NVIDIA technical blog, Hopper architecture in depth (A100 SXM4 comparison table): https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/
- NVIDIA CUDA features archive (CUDA 13.0 architecture support removal): https://docs.nvidia.com/cuda/cuda-features-archive/index.html