Llama 4 Maverick
Llama 4 Maverick is a natively multimodal model capable of processing both text and images. It features a 17 billion active parameter mixture-of-experts (MoE) architecture with 128 experts, supporting a wide range of multimodal tasks such as conversational interaction, image analysis, and code generation. The model includes a 1 million token context window.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DeepInfra | available | $0.20/Mtok | $0.80/Mtok | 1.0M tokens context | 99.9% | 272 ms p50 TTFT | fp8 |
| DigitalOcean | available | $0.25/Mtok | $0.87/Mtok | 128K tokens context | 100.0% | 342 ms p50 TTFT | |
| Novita | available | $0.27/Mtok | $0.85/Mtok | 1.0M tokens context | 99.9% | 464 ms p50 TTFT | fp8 |
| Parasail | available | $0.35/Mtok | $1.00/Mtok | 524K tokens context | 100.0% | 345 ms p50 TTFT | fp8 |
| available | $0.35/Mtok | $1.15/Mtok | 524K tokens context | 99.2% | 566 ms p50 TTFT |