GPT-4.1 mini GPT-4.1 mini provides a balance between intelligence, speed, and cost. It's a significant leap in small model performance, even beating GPT-4o in many benchmarks while reducing latency and cost.
Benchmark results Service providers 88.4%
87.5%
84.1%
78.5%
73.1%
72.7%
67.0%
65.0%
61.7%
60.5%
56.8%
55.8%
54.6%
49.6%
49.3%
47.2%
45.1%
42.2%
40.2%
36.0%
35.8%
35.0%
34.7%
33.3%
31.6%
23.6%
15.0%
11.0%
3.7%
Pricing, uptime, and speed via OpenRouter — updated Jul 17, 2026, 04:19 AM.
Provider Status Input Output Limits Uptime Speed Notes OpenAI available $0.40/Mtok cache $0.10/Mtok $1.60/Mtok 1.0M tokens context 33K tokens max output 99.6% 5m 99.8% 650 ms p50 TTFT 32 tok/s p50 $0.01/web search Azure available $0.44/Mtok cache $0.11/Mtok $1.76/Mtok 1.0M tokens context 33K tokens max output — 827 ms p50 TTFT 56 tok/s p50 cache $0.01/web search