Dataset for evaluating the visual perception capabilities of LVLMs.
-
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
Paper • 2412.00947 • Published • 7 -
ryokamoi/VisOnlyQA_Eval_Real
Viewer • Updated • 500 • 515 • 2 -
ryokamoi/VisOnlyQA_Eval_Synthetic
Viewer • Updated • 700 • 538 • 2 -
ryokamoi/VisOnlyQA_Train
Viewer • Updated • 70k • 899 • 2