← Back to Directory

NVIDIA: Nemotron 3 Ultra

NVIDIA Developer Architecture Profile

Intelligence (ELO)1419Chatbot Arena Verified
Max Context Limit262,144Max Out: 16,384 tokens
Prompt Caching80% OFF$0.10 / 1M cached
Standard API / 1M$2.70🧠 Deep Reasoner

Model Capabilities & Contract SLAs

  • Reasoning
  • Coding & Logic
  • Fictional
  • 🧠 Extended Test-Time Reasoning
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Granular Pricing Matrix

Input Tokens (Standard Prompt)$0.50 / 1M
⚡ Cached Input (Prompt Cache Read)80% OFF$0.10 / 1M
Output Tokens (Completion)$2.20 / 1M

Pricing data via OpenRouter. Sync: 10/2/2026

Evaluate Competitors