NVIDIA Jetson Thor Guide: Specs, Price, Models (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

NVIDIA Jetson Thor is a Blackwell-based robotics computer with 128 GB of memory and 2,070 sparse FP4 TFLOPS, sold as the $3,499 Jetson AGX Thor Developer Kit. It is the first Jetson that can hold a 70B-class model on the device, at a 40 to 130 watt budget and 273 GB/s of memory bandwidth.

TL;DR

  • Specs: 2560-core Blackwell GPU with 96 fifth-generation tensor cores, 128 GB 256-bit LPDDR5X at 273 GB/s, 2,070 FP4 TFLOPS (sparse), 40W to 130W (NVIDIA developer blog).
  • Price: developer kit $3,499 per the same NVIDIA post. NVIDIA lists no price for the T5000 module.
  • T4000: a cut-down module with 1,536 cores, 64 GB and 1,200 FP4 TFLOPS at 40W to 70W, supported from JetPack 7.1.
  • Fit: 128 GB of shared memory holds a 70B model at 8-bit by arithmetic, but 273 GB/s limits how fast it decodes.
  • Verdict: a robotics and physical-AI platform, not a cheap LLM server. If you only want to run big models at home, compare a DGX Spark or a cloud GPU.

What Jetson Thor is

Thor is the Blackwell generation of NVIDIA's embedded line, succeeding Jetson AGX Orin. The module pairs a 14-core Arm Neoverse-V3AE CPU with a Blackwell GPU and one pool of LPDDR5X memory shared by both, which is the same unified memory idea as the Jetson Orin Nano, at 16 times the capacity. The GPU supports multi-instance GPU partitioning (MIG), so one module can split into isolated slices for separate workloads, such as a perception stack and a language model.

It also has native FP4 support, with a Transformer Engine that switches dynamically between FP4 and FP8. Our explainers on FP4 and NVFP4 vs MXFP4 cover the formats.

Specs and the three variants

From NVIDIA's Jetson Thor page:

SpecAGX Thor Dev KitT5000 moduleT4000 module
AI performance (FP4, sparse)2,070 TFLOPS2,070 TFLOPS1,200 TFLOPS
GPU2560-core Blackwell, MIG with 10 TPCs2560-core Blackwell, MIG with 10 TPCs1536-core Blackwell, MIG with six TPCs
Memory128 GB 256-bit LPDDR5X128 GB 256-bit LPDDR5X64 GB 256-bit LPDDR5X
Memory bandwidth273 GB/s273 GB/s273 GB/s
Power40W to 130W40W to 130W40W to 70W

Notice that the T4000 keeps the same 273 GB/s. Token generation is mostly bandwidth-bound, so the T4000 is the smaller part on compute and capacity but not on the number that governs single-stream decode speed.

The developer kit is $3,499 according to NVIDIA's Thor announcement post. NVIDIA's product page lists no prices for the kit or either module and sends buyers to partners. Retail prices and stock vary by distributor, so check one before budgeting.

Software: JetPack 7.1

JetsonHacks reports that JetPack 7.1 is the first production release of Jetson Linux and the JetPack libraries for AGX Thor. It packages Jetson Linux 38.4 with Linux kernel 6.8 and an Ubuntu 24.04 LTS root file system, adds support for the Jetson T4000, and adds TensorRT Edge-LLM support. After flashing, the compute stack installs with:

sudo apt update
sudo apt install nvidia-jetpack

Those two commands are from the JetsonHacks post. The same source says the Orin release is planned for JetPack 7.2 per a reply in its comments, and that NGC containers exist for PyTorch, TensorRT, SGLang, vLLM and Triton. See our guides to SGLang and vLLM for what those engines do; whether a given version runs well on Thor is something to confirm in NVIDIA's container notes.

What runs on-device

Computed (weights only, parameters times bytes): 128 GB shared with the operating system and your robot software is still a large budget.

ModelFP16FP84-bit
8B dense16 GB8 GB4 GB
32B dense64 GB32 GB16 GB
70B dense140 GB70 GB35 GB

A 70B model at FP16 (140 GB) does not fit in 128 GB at all. At FP8 (70 GB) or 4-bit (35 GB) it does, with room for a KV cache. The T4000's 64 GB takes a 32B model at FP8 or a 70B at 4-bit (see our VRAM sizing guide).

Computed bandwidth ceiling. Single-stream decode reads the model once per token, so the upper bound is bandwidth divided by weight bytes: 273 GB/s over 70 GB of FP8 weights is about 3.9 tokens per second, and 273 over 35 GB of 4-bit weights is about 7.8. These are ceilings from arithmetic, not measurements. NVIDIA's own number below is higher because it measures batched throughput, not single-stream speed.

Published by NVIDIA. NVIDIA's Thor post compares Thor against Jetson AGX Orin in tokens per second, in MAXN power mode on both, with sequence length 2048, output length 128 and max concurrency 8. The batched setup is the reason these exceed single-stream ceilings. These are NVIDIA's numbers:

ModelThorAGX Orin
Llama 3.1 8B150.8112.33
Llama 3.3 70B12.647.38
Qwen3-30B-A3B226.4276.69
Qwen3-32B79.116.84
DeepSeek-R1-Distill-Qwen-32B82.6316.96
Qwen2.5-VL-7B252154.02
GR00T N1.5 (robot VLA)41.515.2

NVIDIA headlines up to 7.5x the AI compute and up to 3.5x the energy efficiency of AGX Orin, and says Qwen2.5-VL-7B reached up to 3.5x faster with FP4 and Eagle speculative decoding against Orin running W4A16. The table shows the practical picture: gains vary from 1.3x on an 8B model to 4.9x on 32B reasoning models. These are NVIDIA's measurements restated; we have not run Thor ourselves.

The mixture-of-experts row stands out: Qwen3-30B-A3B reads only about 3 billion active parameters per token, so it decodes quickly on modest bandwidth. See our mixture-of-experts entry for why capacity and speed decouple.

Vision-language-action models for robots are the headline workload: NVIDIA lists GR00T N1 and N1.5. NVIDIA's Jetson AI Lab showcases community projects on Thor, including a 10.5B-parameter vision-language-action model (Alpamayo R1) running in a retrofitted car and two robot arms driven by GR00T N1.5.

Thor vs the alternatives

  • Vs DGX Spark: both pair 128 GB with 273 GB/s of bandwidth. Spark is a desktop dev box; Thor is built for embedded power budgets, I/O for sensors and robots.
  • Vs a Mac Studio: Apple's unified memory suits pure LLM use, but it is not a CUDA robotics platform. See the comparison post for numbers.
  • Vs a discrete GPU such as the RTX 5090: the RTX 5090 has 32 GB, so it is faster for models that fit, while Thor holds models four times larger in shared memory. The wider tradeoffs are in consumer GPUs for AI in 2026.
  • Vs the Orin Nano: 8 GB and 67 TOPS for $249 against 128 GB for $3,499. They are different products for different devices.

Train in the cloud, deploy on Thor

Thor can run big models, but it is not where you train them. The workflow:

  1. Fine-tune on a datacenter or workstation GPU, where the optimizer state, gradients and activations of full or LoRA training need far more than 273 GB/s of bandwidth to be practical. Our fine-tuning explainer lists the methods.
  2. Quantize to FP8 or 4-bit so the model fits and decodes quickly.
  3. Convert for the runtime you will ship (TensorRT Edge-LLM, vLLM or SGLang containers) and measure on the target power mode.
  4. Deploy, then repeat the cloud step as your data improves.

The cloud part is a few hours of rented compute rather than a purchase.

Run it on a cloud GPU

Aquanode manages and optimizes GPUs for training and inference workloads; the box below shows the cloud GPUs you can rent for the training and quantization side, with live prices.

FAQ

How much does Jetson AGX Thor cost?

NVIDIA's announcement lists the developer kit at $3,499. NVIDIA's pages list no price for the T5000 or T4000 modules.

How much memory does Jetson Thor have?

128 GB of LPDDR5X on the dev kit and T5000, 64 GB on the T4000, all at 273 GB/s.

Can Jetson Thor run a 70B model?

Memory-wise yes at FP8 or 4-bit, not at FP16. NVIDIA publishes 12.64 tokens per second for Llama 3.3 70B in its batched test.

Which JetPack does Thor need?

JetPack 7.1 is the first production release for AGX Thor and adds T4000 support, per JetsonHacks.

Is Thor good for training?

No. Fine-tune on a bigger or faster GPU and deploy the result to Thor.

Thor or DGX Spark?

Spark for a desktop dev box, Thor for a robot or an embedded device. Specs are similar on memory and bandwidth.

Sources

#consumer gpu#edge ai#jetson#jetson thor#robotics#edge inference

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.