MME

reasoning official site →

A comprehensive evaluation benchmark for Multimodal Large Language Models measuring both perception and cognition abilities across 14 subtasks. Features manually designed instruction-answer pairs to avoid data leakage and provides systematic quantitative assessment of MLLM capabilities.

Methodology

Imported from llm-stats public benchmark metadata. Modality: multimodal. Max score: 1. Categories: multimodal, reasoning, vision. Language: en. Verified by llm-stats: no.

Leaderboard

  1. DeepSeek VL2 self-reported llm-stats
    22.5%
  2. DeepSeek VL2 Small self-reported llm-stats
    21.2%
  3. DeepSeek VL2 Tiny self-reported llm-stats
    19.1%