Qwen3-Next-80B-A3B-Instruct
Qwen3-Next-80B-A3B-Instruct is the first in the Qwen3-Next series, featuring groundbreaking architectural innovations. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared) achieving extreme low activation ratio, and Multi-Token Prediction for improved performance and faster inference. With 80B total parameters and only 3B activated, it outperforms Qwen3-32B-Base with 10% training cost and 10x throughput for 32K+ contexts. The model performs on par with Qwen3-235B-A22B-Instruct-2507 while excelling at ultra-long-context tasks up to 256K tokens (extensible to 1M with YaRN). Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)).
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DeepInfra | available | $0.09/Mtok | $1.10/Mtok | 262K tokens context | 98% | 533 ms p50 TTFT | fp8 |
| Alibaba | available | $0.10/Mtok | $0.78/Mtok | 131K tokens context | 100.0% | 564 ms p50 TTFT | |
| Parasail | available | $0.10/Mtok | $1.10/Mtok | 262K tokens context | 100.0% | 534 ms p50 TTFT | fp8 |
| available | $0.15/Mtok | $1.20/Mtok | 262K tokens context | 100.0% | 584 ms p50 TTFT | ||
| Novita | available | $0.15/Mtok | $1.50/Mtok | 131K tokens context | 99.5% | 976 ms p50 TTFT | bf16 |