← Back to Directory

NVIDIA: Nemotron 3.5 Lightning

NVIDIA Developer Architecture Profile

Intelligence (ELO)1045Chatbot Arena Verified
Max Context Limit262,144Max Out: 32,768 tokens
Prompt Caching50% OFF$0.03 / 1M cached
Standard API / 1M$0.22Blended Standard Rate

Model Capabilities & Contract SLAs

    NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

    Granular Pricing Matrix

    Input Tokens (Standard Prompt)$0.06 / 1M
    ⚡ Cached Input (Prompt Cache Read)50% OFF$0.03 / 1M
    Output Tokens (Completion)$0.16 / 1M

    Pricing data via OpenRouter. Sync: 10/2/2026

    Evaluate Competitors