Nemotron 3 Nano (30B A3B)
Nemotron 3 Nano is a 31.6B hybrid MoE model optimized for fast, long‑context agentic reasoning. It mixes Mamba‑2 and Transformer layers with a sparse MoE router (~3.6B active params per token) to deliver up to 4× higher throughput than Nemotron 2 and strong accuracy across math, coding, and tools. It supports a 1M‑token context window, offers Reasoning ON/OFF and a thinking‑budget to control costs, and ships with open weights, data, and RL tooling (NeMo Gym/RL). Released Dec 15, 2025 under the NVIDIA Open Model License, it’s built as the efficient backbone for multi‑agent systems at scale.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DeepInfra | available | $0.05/Mtok | $0.20/Mtok | 262K tokens context | 99.2% | 1,247 ms p50 TTFT | fp4 |
| Novita | available | $0.05/Mtok | $0.20/Mtok | 262K tokens context | 100.0% | 545 ms p50 TTFT | fp4 |
| Nebius | available | $0.06/Mtok | $0.24/Mtok | 262K tokens context | 96% | 299 ms p50 TTFT | fp8 |