MMBench reasoning A bilingual benchmark for assessing multi-modal capabilities of vision-language models through multiple-choice questions in both English and Chinese, providing systematic evaluation across diverse vision-language tasks with robust metrics. Leaderboard Showing 9 of 9 results Export CSV Graph Rank Step3-VL-10B 91.8% iQwen2.5 VL 72B Instruct 88.0% iPhi-4-multimodal-instruct 86.7% iQwen2-VL-72B-Instruct 86.5% iQwen2.5 VL 7B Instruct 84.3% iPhi-3.5-vision-instruct 81.9% iDeepSeek VL2 Small 80.3% iDeepSeek VL2 79.6% iDeepSeek VL2 Tiny 69.2% i