gpt-oss-120b
A 117B (MoE) reasoning model whose MXFP4 weights fit a single 80GB GPU. View on Hugging Face.
What gpt-oss-120b is
gpt-oss-120b is a 117B-parameter mixture-of-experts language model published by OpenAI on Hugging Face, one of two open-weight gpt-oss models built for reasoning, agentic tasks and developer use. Its MoE weights are post-trained in MXFP4 precision, which the model card says is what lets a 117B-parameter model run on a single 80GB GPU (NVIDIA H100 or AMD MI300X). It supports configurable reasoning effort (low, medium, high), function calling, web browsing and Python code execution, and is released under Apache 2.0.
License note: fully permissive, no gating and no usage restrictions. Facts on this page are sourced from gpt-oss-120b's Hugging Face model card, not benchmarked by Aquanode.
What it's used for
- Reasoning and chain-of-thought tasks
- Agentic tool use and function calling
- Web browsing and code-execution agents
Deploy gpt-oss-120b
Needs 1× A100 at its published (MXFP4) precision. Aquanode doesn't run a hosted inference API for gpt-oss-120b. Spin up a GPU pod sized for it below and install your own inference stack (vLLM, SGLang, TGI, llama.cpp, Ollama, or whatever you prefer) on it.
Cheapest current fit at gpt-oss-120b's published (MXFP4) precision: 1× A100, at $0.851/hr per GPU ($0.851/hr total).
More models
DeepSeek-R1-0528
Reasoning
Llama 3.3 70B Instruct
LLM
Qwen2.5 7B Instruct
LLM
FLUX.1 [dev]
Image generation
deepseek-coder-6.7b-instruct
Coding
DeepSeek-OCR
Vision
DeepSeek-R1-Distill-Llama-70B
Reasoning
DeepSeek-R1-Distill-Llama-8B-abliterated
Reasoning
DeepSeek-R1-Distill-Qwen-1.5B
Reasoning
DeepSeek-R1-Distill-Qwen-14B-abliterated-v2
Reasoning
DeepSeek-V3
LLM
DeepSeek-V3-0324
LLM
DeepSeek-V3.1
LLM
DeepSeek-V3.2-Exp
LLM
Dia-1.6B
Speech / TTS
Dolphin3.0-R1-Mistral-24B
Reasoning
EXAONE-3.5-32B-Instruct
LLM
FLUX.1-Kontext-dev
Image generation
FLUX.1-schnell
Image generation
FLUX.2-dev
Image generation
gemma-3-12b-it
Vision
gemma-3-4b-it
Vision
gemma-4-31B-it
Vision
GLM-4.5-Air
LLM
GLM-4.5V
Vision
GLM-4.6
LLM
GLM-4.7-Flash
LLM
GLM-5
LLM
GLM-5.1
LLM
GLM-5.2
LLM
granite-3.2-8b-instruct
LLM
Hermes-3-Llama-3.1-405B-FP8
LLM
Hermes-3-Llama-3.1-70B
LLM
Hermes-3-Llama-3.1-8B
LLM
HiDream-I1-Full
Image generation
InternVL3-78B
Vision
Kimi-K2-Instruct
LLM
Kimi-K2-Instruct-0905
LLM
Krea-2-Raw
Image generation
turn-detector
Speech / TTS
Llama-3.1-405B-Instruct
LLM
Meta-Llama-3.1-70B-Instruct-FP8
LLM
Llama-3.1-8B-Instruct
LLM
Llama-3.1-8B-Lexi-Uncensored-V2
LLM
Llama-3.1-Nemotron-Nano-8B-v1
LLM
Llama-3.2-3B-Instruct
LLM
Llama-4-Maverick-17B-128E-Instruct-FP8
Vision
Llama-4-Scout-17B-16E-Instruct
Vision
Llama-Guard-3-8B
LLM
Llama-Guard-4-12B
LLM
LTX-2
Video generation
MiniMax-H3
LLM
MiniMax-M2.5
LLM
MiniMax-M2.7
LLM
Mistral-Small-24B-Instruct-2501
LLM
Mixtral-8x7B-Instruct-v0.1
LLM
mochi-1-preview
Video generation
multilingual-e5-large-instruct
Embedding
Nanonets-OCR-s
Vision
Nous-Hermes-2-Mistral-7B-DPO
LLM
orpheus-3b-0.1-ft
Speech / TTS
parakeet-tdt-0.6b-v3
Speech / TTS
Phi-3.5-mini-instruct
LLM
Phi-3-mini-4k-instruct
LLM
phi-4
LLM
Qwen-Image
Image generation
Qwen2.5-14B-Instruct
LLM
Qwen2.5-3B-Instruct
LLM
Qwen2.5-Coder-32B-Instruct
Coding
Qwen3-1.7B
LLM
Qwen3-1.7B-Base
LLM
Qwen3-14B-Base
LLM
Qwen3-235B-A22B-Instruct-2507-FP8
LLM
Qwen3-235B-A22B-Thinking-2507
Reasoning
Qwen3-30B-A3B
LLM
Qwen3-32B
LLM
Qwen3-4B
LLM
Qwen3-4B-Base
LLM
Qwen3.5-122B-A10B
LLM
Qwen3.5-397B-A17B
LLM
Qwen3.5-9B
LLM
Qwen3.6-35B-A3B
LLM
Qwen3.8-27B-FP8
LLM
Qwen3-8B
LLM
Qwen3-Coder-480B-A35B-Instruct
Coding
Qwen3-Coder-Next
Coding
Qwen3-Embedding-0.6B
Embedding
Qwen3-Next-80B-A3B-Instruct
Reasoning
Qwen3-Next-80B-A3B-Thinking
Reasoning
Qwen3-VL-32B-Instruct
Vision
ReaderLM-v2
LLM
RealVisXL_V5.0
Image generation
SmolLM2-360M-Instruct
LLM
stable-diffusion-xl-base-1.0
Image generation
Trinity-Large-Preview
LLM
Wayfarer-12B
LLM
whisper-large-v3
Speech / TTS
gpt-oss-20b
Reasoning
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
LLM
MiniMax-M2
Coding