MLVU multimodal A comprehensive benchmark for multi-task long video understanding that evaluates multimodal large language models on videos ranging from 3 minutes to 2 hours across 9 distinct tasks including reasoning, captioning, recognition, and summarization. Leaderboard Showing 9 of 9 results Export CSV Graph Rank Qwen3.5-122B-A10B 87.3% iQwen3.6 Plus 86.7% iQwen3.6-27B 86.6% iQwen3.6-35B-A3B 86.2% iQwen3.5-27B 85.9% iQwen3.5-35B-A3B 85.6% iQwen3 VL 235B A22B Instruct 84.3% iQwen3 VL 235B A22B Thinking 83.8% iQwen2.5 VL 7B Instruct 70.2% i