ComplexFuncBench

reasoning

ComplexFuncBench is a benchmark designed to evaluate large language models' capabilities in handling complex function calling scenarios. It encompasses multi-step and constrained function calling tasks that require long-parameter filling, parameter value reasoning, and managing contexts up to 128k tokens. The benchmark includes 1,000 samples across five real-world scenarios.

Leaderboard

  1. 66.5%
  2. 65.5%
  3. 65.2%
  4. 63.0%
  5. 49.3%
  6. 17.6%
  7. 5.7%