Gemma 3 4B
Gemma 3 4B is a 4-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DeepInfra | available | $0.05/Mtok | $0.10/Mtok | 131K tokens context | 100.0% | 217 ms p50 TTFT | bf16 |