Qwen3.6-35B-A3B
Qwen3.6-35B-A3B is the first open-weight variant of the Qwen3.6 series, a multimodal Mixture-of-Experts model with 35B total parameters and 3B activated. It pairs a vision encoder with a hybrid 40-layer language model that interleaves Gated DeltaNet linear-attention blocks and Gated Attention blocks (10 × (3 × DeltaNet + 1 × Attention)) over 256 experts (8 routed + 1 shared, expert dim 512). The release prioritizes stability and real-world utility, with substantial gains in agentic coding (frontend workflows, repo-level reasoning) and a new option to preserve reasoning context across turns. Native context length is 262K tokens, extensible to ~1M via YaRN, and the model thinks by default.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| AkashML | available | $0.14/Mtok | $1.00/Mtok | 262K tokens context | 99.9% | 761 ms p50 TTFT | fp8 |
| Parasail | available | $0.15/Mtok | $1.00/Mtok | 262K tokens context | 99.9% | 576 ms p50 TTFT | fp8 |
| AtlasCloud | available | $0.19/Mtok | $1.11/Mtok | 262K tokens context | 99.8% | 1,229 ms p50 TTFT | fp8 |
| SiliconFlow | available | $0.20/Mtok | $1.60/Mtok | 262K tokens context | 99.5% | 1,290 ms p50 TTFT | fp8 |
| WandB | available | $0.25/Mtok | $1.25/Mtok | 262K tokens context | 99.9% | 336 ms p50 TTFT | fp8 |
| DekaLLM | -2 | $0.13/Mtok | $1.00/Mtok | 262K tokens context | 94% | 704 ms p50 TTFT |