NVIDIA: Nemotron 3 Ultra
NVIDIA Developer Architecture Profile
Intelligence (ELO)1419Chatbot Arena Verified
Max Context Limit262,144Max Out: 16,384 tokens
Prompt Caching80% OFF$0.10 / 1M cached
Standard API / 1M$2.70🧠 Deep Reasoner
Model Capabilities & Contract SLAs
- Reasoning
- Coding & Logic
- Fictional
- 🧠 Extended Test-Time Reasoning
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Granular Pricing Matrix
Input Tokens (Standard Prompt)$0.50 / 1M
⚡ Cached Input (Prompt Cache Read)80% OFF$0.10 / 1M
Output Tokens (Completion)$2.20 / 1M
Pricing data via OpenRouter. Sync: 10/2/2026
Evaluate Competitors
VS Engine MatchupNVIDIA: Nemotron 3 Ultra vs AionLabs: Aion 3.5 MiniVS Engine MatchupNVIDIA: Nemotron 3 Ultra vs DeepSeek: DeepSeek V4 Flash Vision ExpVS Engine MatchupNVIDIA: Nemotron 3 Ultra vs AionLabs: Aion-3.0-MiniVS Engine MatchupNVIDIA: Nemotron 3 Ultra vs SpaceXAI: Grok 4.20VS Engine MatchupNVIDIA: Nemotron 3 Ultra vs Mistral: Ministral 3 14B 2512VS Engine MatchupNVIDIA: Nemotron 3 Ultra vs Anthropic: Claude Haiku 4.5