GroundUI-1K

multimodal

A subset of GroundUI-18K for UI grounding evaluation, where models must predict action coordinates on screenshots based on single-step instructions across web, desktop, and mobile platforms.

Leaderboard

  1. 81.4%
  2. 80.2%