Nemotron 3 Super (120B A12B)
Nemotron 3 Super is a 120B total / 12B active parameter hybrid Mamba-Attention Mixture-of-Experts model optimized for agentic reasoning, coding, planning, tool calling, and long-context analysis. It introduces LatentMoE (projecting tokens into a compressed latent space for expert routing, enabling 4x more experts at the same inference cost), Multi-Token Prediction for native speculative decoding (up to 3x faster generation), and native NVFP4 pretraining on Blackwell. The hybrid architecture interleaves Mamba-2 layers for linear-time sequence processing with strategically placed Transformer attention layers as global anchors, supporting a 1M-token context window. Pre-trained on 25 trillion tokens and post-trained with multi-environment RL across 21 configurations using NeMo Gym/RL with 1.2 million rollouts. Achieves up to 5x higher throughput than previous Nemotron Super and 2.2x higher throughput than GPT-OSS-120B while maintaining comparable accuracy.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DigitalOcean | available | $0.21/Mtok | $0.45/Mtok | 1.0M tokens context | 95% | 8,787 ms p50 TTFT | |
| Nebius | available | $0.30/Mtok | $0.90/Mtok | 262K tokens context | 99.9% | 841 ms p50 TTFT | fp4 |
| DekaLLM | -2 | $0.08/Mtok | $0.45/Mtok | 262K tokens context | 87% | 17,104 ms p50 TTFT | fp8 |
| DeepInfra | -2 | $0.08/Mtok | $0.40/Mtok | 262K tokens context | 90% | 5,574 ms p50 TTFT | bf16 |