Graphwalks BFS <128k

reasoning

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length under 128k tokens, returning nodes reachable at specified depths.

Leaderboard

  1. 94.0%
  2. 93.0%
  3. 78.3%
  4. 76.3%
  5. 73.4%
  6. 72.3%
  7. 61.7%
  8. 61.7%
  9. 51.0%
  10. 41.7%
  11. 25.0%