nvidia/nemotron-3-super-120b-a12b
NVIDIA Nemotron 3 Super 120B A12B
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. It delivers up to 7x higher throughput, providing fast, cost-efficient inference for agentic tasks. Additionally, a long context window gives the model long-term memory, preventing AI agents from losing focus on long, multi-step tasks and ensuring high-accuracy results. Fully open with weights, datasets, and recipes, Super allows easy customization and secure deployment anywhere.
Price comparison
USD per one million tokens, using current provider data.
Provider latency
Last-hour p50 in milliseconds; lower is better.
Provider uptime
Reported availability over the last 24 hours; not a contractual SLA.
Inference providers
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Bedrock | 256K | 32K | 255 ms | 94.0 t/s | 100.000% | $0.15 | $0.65 |
What can this model do?
Available modalities and capabilities across inference providers.
Technical profile
- Model type
- Language model
- Release date
- Mar 11, 2026
- Knowledge cutoff
- —
- Regions
- —
- API specifications
- V2, V3, V4
- Temperature control
- Supported
Model identifier
nvidia/nemotron-3-super-120b-a12bRelated models
More options from the same creator or model category.
