The NVIDIA Jetson Orin Nano is a credit-card-sized module with an Ampere GPU, 8 GB of shared LPDDR5 memory and up to 67 INT8 TOPS, sold in a developer kit that NVIDIA cut from $499 to $249. It runs small language models (up to roughly 8B parameters at 4-bit), vision models and robotics stacks on 7 to 25 watts, but it is a deployment target, not a place to train anything large.
TL;DR
- Specs: 8 GB 128-bit LPDDR5, 102 GB/s, 1024 CUDA cores, 32 tensor cores, 67 INT8 TOPS (sparse), 7W to 25W, per NVIDIA's product page.
- Price: the developer kit is $249, down from $499, per NVIDIA's announcement. Production-module pricing is set by NVIDIA's partners and we did not find a list price.
- What fits: 8 GB is shared between the CPU, GPU and operating system, so 4-bit models up to about 8B are the ceiling. NVIDIA publishes Llama 3.1 8B at 19.1 tokens per second in its INT4 test.
- The sensible workflow: fine-tune or train on a datacenter GPU, quantize, then ship the result to the Jetson. Do not fine-tune an LLM on the device.
- Verdict: buy it to prototype and deploy edge inference and vision. Rent a cloud GPU for the heavy lifting.
What the Orin Nano actually is
Jetson is NVIDIA's embedded line: a system-on-module with CPU, GPU and memory on one board, designed for robots, cameras and other devices that cannot hold a desktop card. The Orin Nano uses the Ampere GPU architecture, the same generation as the RTX 3090, scaled down to 1024 CUDA cores and 32 tensor cores.
The decisive design choice is memory. A desktop GPU has its own VRAM. The Orin Nano has a single pool of LPDDR5 that the CPU and GPU both draw from, which is why NVIDIA lists one 8 GB figure rather than separate system and graphics memory. Our explainers on unified memory and dedicated vs shared GPU memory cover why that matters. In practice, the operating system, your application and the model all share the same 8 GB, so the model gets less than 8 GB.
Specs
All figures below are from NVIDIA's Jetson Orin Nano Super Developer Kit page and the JetPack 6.2 technical blog.
| Spec | Orin Nano 8GB (Super mode) | Before Super mode |
|---|---|---|
| GPU | Ampere, 1024 CUDA cores, 32 tensor cores | same |
| AI performance | 67 INT8 TOPS (sparse) | 40 TOPS |
| CPU | 6-core Arm Cortex-A78AE, 1.7 GHz | 1.5 GHz |
| Memory | 8 GB 128-bit LPDDR5 | same |
| Memory bandwidth | 102 GB/s | 68 GB/s |
| GPU clock | 1020 MHz | 625 MHz |
| Power | 7W to 25W | 7W, 15W |
| Storage | SD card slot and external NVMe | same |
NVIDIA also sells a 4 GB Orin Nano module. Its bandwidth went from 34 GB/s to 51 GB/s with Super mode, and it is the right choice only for small vision workloads.
TOPS here means INT8 operations per second with structured sparsity. It is a peak figure for quantized inference, not a speed for any specific model, and it is not comparable to the TFLOPS numbers on datacenter cards (see our TFLOPS glossary entry).
Power modes
JetPack 6.2 added new power modes to the Orin Nano. For the 8 GB module, NVIDIA's table lists the old modes (7W, 15W) and the new set (15W, 25W and MAXN SUPER). MAXN SUPER is uncapped: if module power goes over the thermal design budget, the module throttles itself. The new modes need a new flashing configuration, so a device that was set up earlier may need reflashing. NVIDIA says existing Orin Nano Developer Kit owners get the performance increase through a software update, with no hardware change.
Power mode is the main tuning knob on Jetson. A battery-powered robot runs the low modes and accepts lower throughput. A wall-powered kiosk runs 25W or MAXN SUPER.
What runs on 8 GB
Start with arithmetic, then look at what NVIDIA measured.
Computed (parameters times bytes per parameter, ignoring KV cache and overhead):
| Model size | FP16 weights | INT8 weights | INT4 weights |
|---|---|---|---|
| 3B | 6 GB | 3 GB | 1.5 GB |
| 7B | 14 GB | 7 GB | 3.5 GB |
| 8B | 16 GB | 8 GB | 4 GB |
| 9B | 18 GB | 9 GB | 4.5 GB |
With the operating system and your application taking a share of the same 8 GB, an 8B model at INT4 (about 4 GB) leaves room for a modest context. An 8B model at FP16 does not fit at all. See how much VRAM do I need for LLMs for the full formula, including the KV cache that grows with context length.
Published by NVIDIA. The JetPack 6.2 blog reports tokens per second for INT4 models served through the MLC API on the Orin Nano 8GB, before and after Super mode. These are NVIDIA's numbers, measured by NVIDIA, not ours:
| Model | Before Super mode | With Super mode |
|---|---|---|
| Llama 3.2 3B | 27.7 | 43.1 |
| Llama 3.1 8B | 14.0 | 19.1 |
| Qwen 2.5 7B | 14.2 | 21.8 |
| Gemma 2 2B | 21.5 | 35.0 |
| Gemma 2 9B | 7.2 | 9.2 |
| Phi-3.5 3.8B | 24.7 | 38.1 |
| SmolLM2 1.7B | 41.0 | 64.5 |
NVIDIA also lists vision-language models on the same module, and these are much slower: for example Qwen2 VL 2B at 4.4 tokens per second and PaliGemma2 3B at 21.6 in Super mode, while the 7B and 8B VLAs and VLMs (VILA 1.5 8B, LLaVA 1.6 7B) sit below one token per second. The VLM tables in the blog carry inconsistent labels, so treat the model ranking as indicative and read the original before relying on a specific cell.
The vision transformer results are where the board is comfortable: NVIDIA lists clip-vit-base-patch32 at 314 and DINOv2 base at 126 in Super mode, run at FP16 through TensorRT.
Why the speeds look like this
Decoding a token reads the active weights once, so single-stream speed is capped by memory bandwidth divided by model size. This is a computed ceiling, not a benchmark: 102 GB/s divided by roughly 4 GB of INT4 8B weights is about 25 tokens per second. NVIDIA's published 19.1 sits below that ceiling, as expected. It also explains why the 9B model at 9.2 tokens per second is so slow: it moves more bytes per token across the same 102 GB/s. This is the same bandwidth-bound behavior described in our best GPU for AI guide, just at roughly 3 percent of an H100's 3.35 TB/s (102 / 3,350, computed).
What to run on it
- Small LLMs (1B to 8B, 4-bit): assistants, command parsers, summarizers and on-device RAG. Pick 3B-class models if you need responsiveness.
- Vision: detection, segmentation and embedding models (CLIP, DINOv2, SAM2) run well. This is the Jetson's traditional strength.
- Robotics and cameras: the board is built for sensor input, and Super mode raised the CPU to 1.7 GHz for the surrounding application.
- Speech: small speech models fit comfortably. See Whisper variants compared for sizes.
What it is not for: 13B and larger LLMs at usable speed, anything with long contexts, and training of anything but tiny models. For the bigger siblings of this module, see the Jetson Thor guide.
JetPack and software
Jetson software ships as JetPack. NVIDIA's Super mode announcement ties the Orin Nano's new modes to JetPack 6.2. JetPack bundles Jetson Linux, CUDA, cuDNN and TensorRT, so models built with standard PyTorch and exported for TensorRT run without a port. NVIDIA's Jetson AI Lab publishes tutorials and prebuilt containers for running LLMs, vision-language models and robotics stacks on Jetson, and is the first place to check for a current recipe for a given model. We did not find a per-model support matrix there, so confirm that a specific model has a Jetson build before you commit to it.
Train in the cloud, deploy on the Jetson
The honest division of labor for an edge product:
- Train or fine-tune on a datacenter or workstation GPU. By the arithmetic in our VRAM guide, QLoRA on an 8B model fits a 24 GB card such as the RTX 4090, and full fine-tuning needs far more. The Orin Nano's 8 GB shared pool and 102 GB/s bandwidth make even LoRA on a 7B model impractical.
- Quantize the result to INT4 or INT8 with a method such as AWQ or GPTQ so it fits the memory budget. Our quantization entry explains the quality trade.
- Convert and test on the device with the runtime you will ship (TensorRT or MLC are the ones NVIDIA benchmarks).
- Deploy and measure on the target power mode, because MAXN SUPER and 15W give very different results.
Fine-tuning is the step that benefits from rented hardware: you need a big GPU for a few hours, not a card on your desk. For the method side, see what is AI model fine-tuning.
Run it on a cloud GPU
When you are preparing a model for a Jetson, the cloud side is the fine-tune and the quantization pass. Aquanode manages and optimizes GPUs for training and inference workloads; the box below shows live availability.
FAQ
Is the Jetson Orin Nano good for running LLMs?
For small ones, yes. NVIDIA publishes 19.1 tokens per second for Llama 3.1 8B at INT4 and 43.1 for Llama 3.2 3B in Super mode. Bigger models do not fit in 8 GB.
How much does the Jetson Orin Nano cost?
NVIDIA states the Super developer kit price as $249, down from $499, as of the announcement. We did not find a current NVIDIA list price for the production modules, so check a distributor.
Can I fine-tune a model on the Orin Nano?
Not practically for LLMs. Fine-tune on a larger GPU and deploy the quantized result. Small vision models are a different matter.
What is Super mode?
A JetPack 6.2 software update that raised GPU, CPU and memory clocks and added the 25W and MAXN SUPER power modes. NVIDIA reports up to a 1.7x generative AI performance gain on the developer kit.
Orin Nano or Jetson Thor?
Orin Nano for cost-sensitive vision and small-model edge devices. Thor if you need a large model or many concurrent models on the device; see the Jetson Thor guide.
Is it better than an RTX card for AI?
Different job. A desktop card such as the RTX 4090 has far more memory bandwidth and speed, but draws 450 W. The Jetson is for devices that must run on 7 to 25 W. Compare broader options in consumer GPUs for AI in 2026.
Sources
- NVIDIA Jetson Orin Nano Super Developer Kit: specs, power range, bandwidth
- NVIDIA blog: Jetson generative AI supercomputer announcement: $249 price, 67 INT8 TOPS, availability
- NVIDIA technical blog: JetPack 6.2 brings Super mode: power modes, clocks, LLM, VLM and ViT tokens per second
- NVIDIA Jetson AI Lab: tutorials and containers
- NVIDIA H100 datacenter page: 3.35 TB/s comparison
- How much VRAM do I need for LLMs: weight-size arithmetic used above