GLM-4.7-Flash
GLM-4.7-Flash is a high-speed, cost-efficient variant of GLM-4.7 optimized for fast inference and lower latency. It retains the coding-centric capabilities of GLM-4.7 including thinking before acting, preserved reasoning across turns, and per-request thinking control for speed or accuracy trade-offs. Ideal for applications requiring quick responses while maintaining strong performance on coding, agentic workflows, and general reasoning tasks.
| Provider | Status | Input | Output | Limits | Uptime | Speed | Notes |
|---|---|---|---|---|---|---|---|
| DeepInfra | available | $0.06/Mtok | $0.40/Mtok | 203K tokens context | 99.8% | 453 ms p50 TTFT | bf16 |
| Cloudflare | available | $0.06/Mtok | $0.40/Mtok | 131K tokens context | 98% | 420 ms p50 TTFT | |
| Venice | available | $0.13/Mtok | $0.50/Mtok | 128K tokens context | 99% | 625 ms p50 TTFT | fp8 |
| Novita | -2 | $0.07/Mtok | $0.40/Mtok | 200K tokens context | 82% | 1,170 ms p50 TTFT | bf16 |