MultiChallenge (o3-mini grader)

reasoning

A realistic multi-turn conversation evaluation benchmark that challenges frontier LLMs across four key areas: instruction retention, inference memory, reliable versioned editing, and self-coherence. Despite near-perfect scores on existing benchmarks, frontier models achieve less than 50% accuracy on MultiChallenge.

Leaderboard

  1. 69.6%
  2. 50.2%
  3. 50.1%
  4. 46.2%
  5. 42.2%
  6. 39.9%
  7. 31.1%