nvidia/nemotron-nano-9b-v2
Nvidia Nemotron Nano 9B V2
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.\
Price comparison
USD per one million tokens, using current provider data.
Provider latency
Last-hour p50 in milliseconds; lower is better.
Provider uptime
Reported availability over the last 24 hours; not a contractual SLA.
Inference providers
Compare current pricing, performance, and privacy across every available provider.
| Provider | Maximum context | Max output | p50 latency | p50 throughput | 24h uptime | Input / 1M tokens | Output / 1M tokens | Privacy |
|---|---|---|---|---|---|---|---|---|
Bedrock | 131.1K | 131.1K | 156 ms | 143.0 t/s | 100.000% | $0.06 | $0.23 | |
Deepinfra | 131.1K | 131.1K | 262 ms | 96.0 t/s | 100.000% | $0.04 | $0.16 |
What can this model do?
Available modalities and capabilities across inference providers.
Technical profile
- Model type
- Language model
- Release date
- Aug 18, 2025
- Knowledge cutoff
- —
- Regions
- —
- API specifications
- V2, V3, V4
- Temperature control
- Supported
Model identifier
nvidia/nemotron-nano-9b-v2Related models
More options from the same creator or model category.
