VLMsAreBlind
reasoning
A vision-language benchmark that probes blind spots and brittle reasoning in multimodal models.
Methodology
Imported from llm-stats public benchmark metadata. Modality: multimodal. Max score: 1. Categories: multimodal, reasoning, vision. Language: en. Verified by llm-stats: no.