OpenAI-MRCR: 2 needle 128k

reasoning

Multi-round Co-reference Resolution (MRCR) benchmark for evaluating an LLM's ability to distinguish between multiple needles hidden in long context. Models are given a long, multi-turn synthetic conversation and must retrieve a specific instance of a repeated request, requiring reasoning and disambiguation skills beyond simple retrieval.

Leaderboard

  1. 95.2%
  2. 76.1%
  3. 73.4%
  4. 57.2%
  5. 47.2%
  6. 38.5%
  7. 36.6%
  8. 31.9%
  9. 18.7%