Llama 3.2 3B Instruct
Llama 3.2 3B Instruct is a large language model that supports a context length of 128K tokens and are state-of-the-art in their class for on-device use cases like summarization, instruction following, and rewriting tasks running locally at the edge.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| Parasail | available | $0.05/Mtok | $0.33/Mtok | 131K tokens context | 100.0% | 212 ms p50 TTFT | bf16 |
| Cloudflare | available | $0.05/Mtok | $0.34/Mtok | 80K tokens context | 99.9% | 212 ms p50 TTFT |