NVIDIA T4 GPU: Specs, VRAM, Price & Benchmarks (2026)
The NVIDIA T4 (16GB GDDR6) launched in 2018. This guide covers its full specs, VRAM, SXM/PCIe differences where they apply, AI performance, and live per-GPU rental pricing on Aquanode.
Short answer on cost: the T4 rents from $0.185 per GPU per hour on Aquanode.
How much VRAM does the T4 have?
The T4 has 16GB GDDR6, with 320+ GB/s of peak memory bandwidth.
- VRAM: T4 16GB GDDR6 vs V100 16GB or 32GB HBM2
- VRAM: T4 16GB GDDR6 vs A100 80GB HBM2e
What fits in 16GB GDDR6 of VRAM
| Model | Precision | Fits? |
|---|---|---|
| Llama 3.2 3B | FP16 | roughly 6GB. Fits comfortably. |
| Llama 3 8B | INT4 | ~5-6GB. Fits, with room for a modest batch. |
| Mistral 7B | FP16 | ~14GB. Fits, but with almost no KV-cache headroom. |
| Llama 3.1 70B | INT4 | ~35-40GB. Does not fit on a single 16GB card. |
Approximate, based on published parameter counts and standard bytes-per-parameter rules of thumb (FP16 ≈ 2 bytes/param, INT4 ≈ 0.5-0.6 bytes/param). Real footprint also depends on KV-cache size and framework overhead.
What can the T4 run?
Popular open models from small to frontier scale, with the memory each needs and how many T4 cards (16GB GDDR6 each) that takes.
| Model | As published | FP8 | INT4 |
|---|---|---|---|
| Qwen/Qwen3-8B 8.2B | BF16: not supported | FP8: not supported | INT4: ~4.6 GB, 1 GPU |
| Qwen/Qwen2.5-14B-Instruct 14.8B | BF16: not supported | FP8: not supported | INT4: ~8.3 GB, 1 GPU |
| Qwen/Qwen3-32B 32.8B | BF16: not supported | FP8: not supported | INT4: ~18.3 GB, 2 GPUs |
| Qwen/Qwen-72B 72.3B | BF16: not supported | FP8: not supported | INT4: ~40.4 GB, 3 GPUs |
| MiniMaxAI/MiniMax-M2.7 228.7B | FP8: not supported | – | INT4: ~128 GB, 8 GPUs |
| deepseek-ai/DeepSeek-R1 684.5B | FP8: not supported | – | INT4: ~383 GB, 24 GPUs |
Estimates: weights at the stated precision plus a flat 20% for KV cache and overhead, at a moderate context length. A dash means the precision is not offered for that model (it is already published at that size). INT4 needs a published quantized checkpoint. Open any model for a per-GPU breakdown, or use the T4 VRAM calculator.
T4 VRAM calculator: check which models fit in its memory at each precision.
All models that fit in 16 GB: the open models whose weights and overhead fit, at native, FP8 and INT4 precision.
T4 specs
| Architecture | NVIDIA, launched 2018 |
| VRAM | 16GB GDDR6 |
| Memory bandwidth | 320+ GB/s |
| FP16 / BF16 tensor throughput | 65 TFLOPS (peak, dense) |
| TDP | 70W |
| Form factor | PCIe, low-profile, single-slot |
Specs sourced from the vendor's public datasheet/product page. See the source.
Related reading: The NVIDIA Inception program, explained, Google Colab alternatives for dedicated GPU access, Free GPU credits for students and researchers, and RunPod volume disk vs. network volume, compared.
GPU Glossary: What is VRAM?, Tensor Cores, CUDA Cores, TFLOPS
T4 AI performance
A 70W single-slot card that runs INT8/INT4 quantized inference cheaply. It is the first generation whose compute capability (7.5) clears the AWQ/GPTQ kernel floor, so small quantized models genuinely run on it. Useful when the job is high-volume small-model serving or video/AI pipelines and the hourly rate matters more than latency.
- Dense FP16/BF16 tensor throughput: 65 TFLOPS
- Memory bandwidth: 320+ GB/s
It is a 2018 part: no BF16 and no FP8 at all, 16GB of VRAM, and 320+ GB/s of bandwidth. Anything trained in BF16 has to be converted, most modern serving stacks assume BF16 or FP8, and nothing above ~13B fits even quantized.
See how it stacks up against other cards in the GPU benchmarks and specs table.
T4 price: what does it cost?
Buying. $2,299 (launch MSRP (2018)).
Renting. The cheapest current on-demand rate for the T4 on Aquanode is $0.185/GPU/hr (no on-demand supply right now; this is the cheapest offer overall). Live rates range from $0.185 to $0.185 per GPU per hour, with a median of $0.185/GPU/hr. Billed by the offer's own terms; the table below shows every live rate.
| Region | $/GPU/hr | Available | VRAM | vCPU | RAM |
|---|---|---|---|---|---|
| Lis, Pt | $0.185 | 1 | 16 GB | 725 | 1774 GB |
T4 price history
Aquanode stores one snapshot of its GPU price index per UTC day. For the T4 that is 2 days so far, 2026-10-09 to 2026-10-10, so this is a short history, not a long-run trend. The lowest per-GPU rate was $0.185 on both 2026-10-09 and 2026-10-10.
| Day (UTC) | Lowest per GPU hour | Median per GPU hour | Data-center lowest per GPU hour | Offers |
|---|---|---|---|---|
| 2026-10-10 | $0.185 | $0.185 | – | 1 |
| 2026-10-09 | $0.185 | $0.185 | – | 1 |
Each row is the stored daily snapshot of the live index, copied as recorded. A dash means that day stored no figure. The same series is available as JSON and summarised in the monthly GPU price report.
How this price is calculated
All prices on this page are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, so that raw price is divided by the number of GPUs it actually covers; some offers report price as already per-GPU, so those are used as-listed. An offer with a missing, zero, or invalid GPU count is excluded entirely rather than published at a guessed rate.
No offers were excluded from this snapshot for a missing or invalid price. No offers were dropped as price outliers in this snapshot.
Only the cheapest qualifying offer per provider is shown in the table above. This page regenerates at most once per hour.
T4 vs V100: how do they compare?
- VRAM: T4 16GB GDDR6 vs V100 16GB or 32GB HBM2
- Memory bandwidth: 320+ GB/s vs 900 GB/s
- Dense FP16 tensor throughput: 65 TFLOPS vs 125 TFLOPS (SXM2), 112 TFLOPS (PCIe)
On Aquanode right now, V100 starts at $0.088/GPU/hr against T4's $0.185/GPU/hr, about 52% less.
T4 vs A100: how do they compare?
- VRAM: T4 16GB GDDR6 vs A100 80GB HBM2e
- Memory bandwidth: 320+ GB/s vs 2,039 GB/s
- Dense FP16 tensor throughput: 65 TFLOPS vs 312 TFLOPS
On Aquanode right now, T4 starts at $0.185/GPU/hr against A100's $0.991/GPU/hr, about 81% less.
Compare the T4 with other GPUs
Compare T4 with
- B200 vs T4
- H200 vs T4
- H100 vs T4
- A100 vs T4
- DGX A100 vs T4
- T4 vs V100
- AMD MI300X vs T4
- L40S vs T4
- L40 vs T4
- A40 vs T4
- RTX A6000 vs T4
- RTX A5000 vs T4
- RTX A4000 vs T4
- L4 vs T4
- RTX 6000 Ada vs T4
- RTX PRO 6000 vs T4
- RTX PRO 6000 WS vs T4
- RTX PRO 6000 SE vs T4
- RTX PRO 5000 vs T4
- RTX 5090 vs T4
- RTX 5080 vs T4
- RTX 5070 Ti vs T4
- RTX 5070 vs T4
- RTX 5060 Ti vs T4
- RTX 4090 vs T4
- RTX 4080 Super vs T4
- RTX 4080 vs T4
- RTX 4070 Ti vs T4
- RTX 4070 Super vs T4
- RTX 4070 vs T4
- RTX 4060 Ti vs T4
- RTX 3090 vs T4
- RTX 3080 vs T4
- RTX 3070 vs T4
- RTX 3060 vs T4
- B300 vs T4
- AMD MI355X vs T4
- AMD MI325X vs T4
- H100 NVL vs T4
- RTX PRO 4500 vs T4
- RTX PRO 4500 SE vs T4
- RTX PRO 4000 vs T4
- RTX 5880 Ada vs T4
- RTX 5000 Ada vs T4
- RTX 4000 Ada vs T4
- RTX 4000 SFF Ada vs T4
- RTX 2000 Ada vs T4
- RTX 4070 Ti Super vs T4
- RTX A4500 vs T4
- RTX 2080 Ti vs T4
- RTX 3080 Ti vs T4
- RTX 3070 Ti vs T4
- RTX 3060 Ti vs T4
- GTX 1080 Ti vs T4
- RTX 4060 vs T4
- RTX 5060 vs T4
- GTX 1660 Super vs T4
- RTX 2060 vs T4
- Radeon RX 9070 XT vs T4
- Radeon RX 6700 XT vs T4
- RTX 6000 vs T4
- A16 vs T4
- P40 vs T4
- P4 vs T4
- Quadro P2000 vs T4
- Quadro M4000 vs T4
Get notified when the price drops
GPU supply moves hourly. Tell us what you're waiting for and we'll email you when a matching offer appears across any provider we track.
Rent a T4 on Aquanode
- On-demand instances from $0.185/GPU/hr, billed by the provider's own terms, with no hardware procurement or long-term commitment.
- 1 live offer across 1 region today.
- Set a price/availability alert above to hear the moment a cheaper or newly-available T4 offer appears.
- Compare every T4 offer side by side, or browse the full multi-provider GPU marketplace.
Good for
A 70W single-slot card that runs INT8/INT4 quantized inference cheaply. It is the first generation whose compute capability (7.5) clears the AWQ/GPTQ kernel floor, so small quantized models genuinely run on it. Useful when the job is high-volume small-model serving or video/AI pipelines and the hourly rate matters more than latency.
Not good for
It is a 2018 part: no BF16 and no FP8 at all, 16GB of VRAM, and 320+ GB/s of bandwidth. Anything trained in BF16 has to be converted, most modern serving stacks assume BF16 or FP8, and nothing above ~13B fits even quantized.
T4 FAQs
How much VRAM does the T4 have?
The T4 has 16GB GDDR6, with 320+ GB/s of peak memory bandwidth.
What is the NVIDIA T4?
The NVIDIA T4 is a GPU released in 2018, with 16GB GDDR6 of memory and a 70W power envelope. See the full spec table above for interconnect, form factor and tensor-throughput details.
How much does it cost to rent a T4?
Live T4 rental prices currently range from $0.185 to $0.185 per GPU per hour, with a median of $0.185 per GPU per hour.
What's the cheapest T4 rate?
The lowest current T4 rate on Aquanode is $0.185 per GPU per hour in Lis, Pt.
How does the T4 compare to the V100?
VRAM: T4 16GB GDDR6 vs V100 16GB or 32GB HBM2 Memory bandwidth: 320+ GB/s vs 900 GB/s Dense FP16 tensor throughput: 65 TFLOPS vs 125 TFLOPS (SXM2), 112 TFLOPS (PCIe) On Aquanode right now, V100 starts at $0.088/GPU/hr against T4's $0.185/GPU/hr, about 52% less. See the full T4 vs V100 comparison for a shared-provider price breakdown.
How does the T4 compare to the A100?
VRAM: T4 16GB GDDR6 vs A100 80GB HBM2e Memory bandwidth: 320+ GB/s vs 2,039 GB/s Dense FP16 tensor throughput: 65 TFLOPS vs 312 TFLOPS On Aquanode right now, T4 starts at $0.185/GPU/hr against A100's $0.991/GPU/hr, about 81% less. See the full T4 vs A100 comparison for a shared-provider price breakdown.
What is the T4 good for?
A 70W single-slot card that runs INT8/INT4 quantized inference cheaply. It is the first generation whose compute capability (7.5) clears the AWQ/GPTQ kernel floor, so small quantized models genuinely run on it. Useful when the job is high-volume small-model serving or video/AI pipelines and the hourly rate matters more than latency.
What are the T4's limitations?
It is a 2018 part: no BF16 and no FP8 at all, 16GB of VRAM, and 320+ GB/s of bandwidth. Anything trained in BF16 has to be converted, most modern serving stacks assume BF16 or FP8, and nothing above ~13B fits even quantized.
Is renting cheaper than buying?
Renting avoids the upfront hardware cost and lets you match spend to actual usage. A rented T4 at $0.185/hr only costs money while it's running, whereas buying ties up capital in hardware that keeps depreciating whether it's in use or not. Buying outright runs $2,299 (launch MSRP (2018)). Which is cheaper depends on how continuously you'd run it; short or bursty workloads usually favor renting.
How is the T4 price calculated?
All prices are normalized to a per-GPU hourly rate using each offer's authoritative GPU count, which the raw price is divided by; some offers report price as already per-GPU. Offers whose price can't be safely normalized, or whose rate is an extreme outlier against the rest of the market, are excluded.
T4 price by region
Related guides
Other models in the same generation, then the rest of the GPU index.