Comments19 pages, 6 figures, 2 tables. Includes appendix with supporting figures and per-subgroup fairness detail. Code and data: https://github.com/bibeshpyakurel/SLAPBench
机构
*
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
;
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院)
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
专题命中
视觉定位与Grounding
:MLLM(title,summary_cn);multimodal large language model(abstract);分类 cs.CV
Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs
机构
*
The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
;
Eisai Inc.(卫材株式会社)
;
IISc, Bangalore(印度科学研究所班加罗尔分校)
;
Cohere Labs Community(Cohere实验室社区)
专题命中
视觉定位与Grounding
:vision language model(title,abstract);grounding(title,abstract);分类 cs.CV
ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination
ReGround:通过自诊断与视觉重检验恢复多步推理中的视觉接地
Lei Peng, Shuai Lv, Wei Hu
机构
*
University of Science and Technology of China(中国科学技术大学)
;
School of Artificial Intelligence and Data Science(人工智能与数据科学学院)
;
State Key Laboratory of Precision and Intelligent Chemistry(精准与智能化学国家重点实验室)
GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision
GeoSearcher: 基于锚点引导的渐进推理遥感视觉定位与过程监督
Dianyu Wang, Peirong Zhang, Xuyang Li, Xiaoxuan Liu, Lei Wang
机构
*
Key Laboratory of Target Cognition and Application Technology (TCAT), Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)
专题命中
视觉定位与Grounding
:grounding(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos
面向遥感图像与视频的无训练开放词汇视觉定位
Ke Li, Di Wang, Yongshan Zhu, Ting Wang, Weiping Ni, Tao Lei, Quan Wang, Xinbo Gao
机构
*
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
;
Interdisciplinary Institute of Artificial Intelligence, Xidian University(西安电子科技大学跨学科人工智能研究院)
;
School of Artificial Intelligence, Xidian University(西安电子科技大学人工智能学院)
;
Northwest Institute of Nuclear Technology(西北核技术研究所)
;
School of Physics and Information Engineering, Fuzhou University(福州大学物理与信息工程学院)
3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
3D-缺陷基准:用于细粒度3D生成缺陷的视觉语言模型评估管道的控制因子研究
Zhenyu Zhao, Nanshan Jia, Jihyeon Je, Yifu Tang, Alvin Chan, Michael Spedden, Michael V. Palleschi, Sui Huang, Jingshen Wang, Zeyu Zheng
机构
*
Roblox Corporation(罗布乐思公司)
;
Berkeley AI Research Lab & Department of Industrial Engineering and Operations Research, University of California, Berkeley(加州大学伯克利分校伯克利人工智能研究实验室及工业工程与运筹学系)
;
Computer Science Department, Stanford University(斯坦福大学计算机科学系)
;
Division of Biostatistics, University of California, Berkeley(加州大学伯克利分校生物统计学系)