Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning
打破似曾相识:通过视觉语言推理对视觉场所识别进行独立审计
Sania Waheed, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan
机构
*
School of Electronics and Computer Science, University of Southampton(南安普顿大学电子与计算机科学学院)
;
School of Electrical Engineering and Computer Science, Queensland University of Technology(昆士兰科技大学电气工程与计算机科学学院)
;
School of Computer Science and Electronic Engineering, University of Essex(埃塞克斯大学计算机科学与电子工程学院)
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Hunan University(湖南大学)
;
ETH Zürich(苏黎世联邦理工学院)
;
INSAIT, Sofia University ``St. Kliment Ohridski''(INSAIT,索菲亚大学『圣克莱门特·奥赫里迪斯』)
;
RAI Institute(RAI研究所)
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
PixelEyes: 解耦感知与推理以实现精准视觉证据查找
Dengxian Gong, Yuanzheng Wu, Haobo Yuan, Zhengdong Hu, Tao Zhang, Yikang Zhou, Shihao Chen, Quanzhu Niu, Kai Wang, Jason Li, Haochen Wang, Lu Qi, Shunping Ji, Ming-Hsuan Yang