What is CUDA?
CUDA, short for Compute Unified Device Architecture, is NVIDIA's parallel computing platform and programming model for running general-purpose code on its GPUs. It lets a program hand thousands of small pieces of work to the GPU at once, which is what AI and scientific workloads need, and it is the software layer most deep learning frameworks are built on.
NVIDIA first released CUDA in 2007, and it runs only on NVIDIA GPUs. Nearly two decades of tooling, libraries and tuned code now sit on top of it, which is why the software question often decides which GPU you can use.
What CUDA includes
CUDA is a stack, not a single program:
- Language extensions. CUDA C++ adds keywords such as
__global__to mark functions that run on the GPU, and the<<<blocks, threads>>>syntax to launch them. A function that runs on the GPU is a kernel, and how threads, blocks and memory are organized is the CUDA programming model. - A compiler. The
nvcccompiler turns that code into instructions for a specific NVIDIA GPU architecture. Those instructions run on the GPU's CUDA cores and other execution units. - Runtime and driver APIs. The runtime API handles memory allocation, copies between CPU and GPU, and kernel launches. The lower-level driver API gives finer control. The driver API is part of NVIDIA's GPU driver, while the toolkit (compiler, headers, libraries) is installed separately.
- Libraries. cuBLAS provides matrix math, and cuDNN provides convolutions and other neural network primitives. Most code never calls the GPU directly; it calls these libraries.
Why AI frameworks depend on CUDA
PyTorch, TensorFlow and JAX all ship CUDA backends and call cuBLAS and cuDNN underneath, so a matrix multiply in your model code ends up as a CUDA library call. Hand-written CUDA kernels for fast attention and low-precision math, plus profilers and debuggers, hang off the same platform.
CUDA vs ROCm
ROCm is AMD's open-source GPU software stack, with its own libraries and a CUDA-like programming interface called HIP. PyTorch and most mainstream training and inference frameworks have ROCm builds, so common workloads run on AMD GPUs. The gaps are CUDA-only libraries and custom CUDA kernels, which have to be ported or do not run. CUDA vs ROCm covers the trade-offs in detail.
What it means when you pick a GPU
CUDA turns the NVIDIA-versus-AMD choice into a software question before it is a hardware one. Check three things before you rent.
Does your stack need CUDA specifically? Custom kernels, CUDA-only libraries or a pinned CUDA build point to NVIDIA, such as the H100. If your framework has a ROCm build you trust, AMD becomes an option. The MI300X has 192GB of memory against the H100's 80GB, so a 70B-class model at 16-bit precision (70 x 2 bytes = 140GB of weights) fits on one MI300X but needs at least two H100s. The cost is the ROCm risk above.
Is the GPU new enough? Each NVIDIA architecture has a compute capability number, and libraries set a floor. The V100 is compute capability 7.0, below the 7.5 floor that AWQ, GPTQ and Marlin INT4 kernels require, so 4-bit serving with those kernels does not run on it at all, however much memory it has. Checking before you rent saves debugging time on a box you are paying for.
Does the CUDA version match the machine? A framework build or container compiled for a newer CUDA release can refuse to start when the machine's driver is too old. Check the CUDA version the driver supports before you pull an image, and pin the framework build to it.
Aquanode rents GPUs by the hour, with current rates on the pricing page. Pick the card from your software requirements first, then from memory capacity and bandwidth.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
CUDA Programming Model
The CUDA programming model organizes GPU code around a nested hierarchy of threads and memory. The three abstractions from NVIDIA's own programming guide, and why they let one program get faster on every new GPU without a rewrite.
CUDA Core
A CUDA Core is the unit inside a Streaming Multiprocessor that executes scalar arithmetic, one instruction issued to a whole group at a time. What separates it from a Tensor Core, whether more of them means a faster GPU, and where they fit in AI training and inference.
Kernel
A CUDA kernel is the function a GPU programmer writes and launches, executed once per thread across thousands of threads at once. How kernels map onto the thread and memory hierarchy, with two worked matrix-multiply examples.
cuBLAS
cuBLAS is NVIDIA's tuned, closed-source implementation of the BLAS standard for GPUs, and the default engine under most matrix math in PyTorch and similar frameworks. The column-major memory layout quirk that trips up almost everyone who calls it directly.
NCCL (NVIDIA Collective Communications Library)
NCCL is NVIDIA's library for all-reduce and other multi-GPU communication. PyTorch uses it to sync GPUs, and NVLink vs PCIe decides how fast it runs.
FlashAttention
FlashAttention is an exact attention algorithm that tiles the computation so the full attention matrix is never written to GPU memory, cutting HBM traffic.