What is cuDNN?

Abbreviated cuDNN

cuDNN (the NVIDIA CUDA Deep Neural Network library) is NVIDIA's GPU-accelerated library of building blocks for deep learning. It provides tuned implementations of the operations that neural networks use constantly, such as convolution, scaled dot-product attention, matrix multiplication, normalization, softmax and pooling, so that frameworks do not have to hand-write CUDA kernels for each of them on each GPU generation.

You rarely call it yourself. NVIDIA lists PyTorch, JAX, TensorFlow and others among the frameworks that cuDNN accelerates, and they call it whenever your model runs a convolution or one of its other supported operations on an NVIDIA GPU.

How it relates to cuBLAS and CUTLASS

Three NVIDIA libraries are easy to confuse:

  • cuBLAS is general linear algebra: matrix multiplication and its relatives. Most of a transformer's matrix math ends up there.
  • CUTLASS is an open-source template library for writing your own matrix-multiply kernels.
  • cuDNN is the deep learning layer: operations specific to neural networks, including convolution and attention, plus the ability to fuse several of them.

NVIDIA describes cuDNN as providing kernels that target Tensor Cores whenever that makes sense, heuristics for choosing the right kernel for a given problem size, and fusion of compute-bound with memory-bound operations. Computations are expressed as a graph of operations on tensors. NVIDIA's pages list scaled dot-product attention among its supported operations, the same operation that FlashAttention accelerates.

Why convolution needs its own library

A convolution is not a plain matrix multiplication. The same small filter slides across an image, and there are many valid ways to compute it, each fastest for different shapes and data types. cuDNN holds several algorithms and picks one. In PyTorch you can see this in the torch.backends.cudnn settings: setting benchmark to True makes cuDNN try several convolution algorithms and keep the fastest, deterministic restricts it to deterministic algorithms, and enabled switches it off altogether.

Worked example: the first layer of ResNet-50

The ResNet paper (He et al., 2015) defines ResNet-50's first layer, conv1, as a 7x7 convolution with 64 filters and stride 2, taking a 224x224 RGB image to a 112x112 output.

  • Multiply-accumulates per image: 112 x 112 outputs x 64 filters x (3 channels x 7 x 7) = 118,013,952, about 118 million.
  • At 2 FLOPs per multiply-accumulate, that is about 236 million FLOPs for this one layer, for one image.
  • A batch of 256 images costs about 60 billion FLOPs in this layer's forward pass alone.

The output of this layer is 112 x 112 x 64 values per image, which is 1.6 MB in 16-bit. A bias add and a ReLU that follow it do almost no arithmetic per value, so run as separate steps each one reads and writes that whole tensor again: about 6.4 MB of extra memory traffic per image for the two, or 1.6 GB across a batch of 256. This is what fusion removes. NVIDIA says cuDNN can fuse compute-bound and memory-bound operations, using runtime-generated kernels for common patterns, so the bias and ReLU are applied while the convolution's results are still on the chip.

Only 3 input channels feed 64 filters, so the arithmetic per byte moved is very different from a late layer with hundreds of channels. A hand-written kernel tuned for one shape can be slow on another, which is the gap a library with many algorithms and a selection heuristic is built to cover. These counts are arithmetic from the published architecture, not measurements of any GPU.

What it means when you pick a GPU

cuDNN is NVIDIA-only. On AMD GPUs the same frameworks use a different library stack (see CUDA vs ROCm), and code that calls cuDNN directly will not run there.

Three practical points follow:

  1. Version matching. cuDNN is installed separately from the CUDA toolkit, and its packages name the CUDA version they target. NVIDIA's quick-install commands specify one explicitly, for example cuda-version=12 for conda, and its container tags carry both, such as 12.8.1-cudnn-devel-ubuntu22.04. Pick an image whose CUDA and cuDNN versions match what your framework build expects.
  2. Precision follows the GPU generation. PyTorch documents that TensorFloat-32 Tensor Core use in cuDNN convolutions applies on Ampere or newer GPUs. The V100 and T4 predate Ampere, so for convolution-heavy vision models an A100 or newer can use TF32 Tensor Core convolutions, which those cards cannot. Check per-card precision support on the A100 page and the H100 page.
  3. Profile before blaming the card. nvidia-smi shows whether the GPU is busy, and Nsight Systems shows whether the time is in cuDNN kernels, in data loading or in communication.

For transformer-only workloads most of the time is in matrix multiplies, so VRAM and memory bandwidth decide the GPU choice more than cuDNN does. Aquanode rents GPUs by the hour, so a short convolution benchmark on two cards costs little.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.