MMLU Chat

math official site →

Chat-format variant of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US history, computer science, law, and other professional and academic subjects. This version uses conversational prompting format for model evaluation.

Methodology

Imported from llm-stats public benchmark metadata. Modality: text. Max score: 1. Categories: finance, general, healthcare, language, legal, math, reasoning. Language: en. Verified by llm-stats: no.

Leaderboard

  1. Llama 3.1 Nemotron 70B Instruct self-reported llm-stats
    80.6%