Llama 4 Scout
Llama 4 Scout is a natively multimodal model capable of processing both text and images. It features a 17 billion activated parameter (109B total) mixture-of-experts (MoE) architecture with 16 experts, supporting a wide range of multimodal tasks such as conversational interaction, image analysis, and code generation. The model includes a 10 million token context window.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DeepInfra | available | $0.10/Mtok | $0.30/Mtok | 328K tokens context | 100.0% | 444 ms p50 TTFT | fp8 |
| Groq | available | $0.11/Mtok | $0.34/Mtok | 131K tokens context | 99.9% | 371 ms p50 TTFT | |
| Novita | available | $0.18/Mtok | $0.59/Mtok | 131K tokens context | 100.0% | 614 ms p50 TTFT | bf16 |
| available | $0.25/Mtok | $0.70/Mtok | 1.3M tokens context | 99% | 981 ms p50 TTFT |