Visual serial processing deficits explain divergences in human and VLM reasoning
视觉序列处理缺陷解释了人类与VLM推理之间的差异
Nicholas Budny, Kia Ghods, Declan Campbell, Raja Marjieh, Amogh Joshi, Sreejan Kumar, Jonathan D. Cohen, Taylor W. Webb, Thomas L. Griffiths
机构
*
Princeton Neuroscience Institute(普林斯顿神经科学研究所)
;
Department of Psychology, Princeton University(普林斯顿大学心理学系)
;
Department of Psychology, Université de Montréal(蒙特利尔大学心理学系)
;
Mila - Quebec AI Institute(魁北克AI研究所)
;
Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割
Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
机构
*
Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系)
;
Department of Radiology, University of Florida(佛罗里达大学放射学系)
;
Research Computing, University of Florida(佛罗里达大学研究计算中心)
;
Department of Medicine, University of Florida(佛罗里达大学医学系)
;
Department of Radiology, UC San Francisco(旧金山大学放射学系)
机构
*
Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院)
;
Department of Computer Science, Emory University(埃默里大学计算机科学系)
;
Department of Computer Science, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系)
;
Tsinghua University(清华大学)
;
Innovation and Research and Development Department, BoardWare Information System Company(BoardWare信息系统公司创新与研发部)
;
School of Mathematics, Peking University(北京大学数学学院)
Do Multi-Agents Solve Better Than Single? Evaluating Agentic Frameworks for Diagram-Grounded Geometry Problem Solving and Reasoning
多智能体表现更优吗?评估用于图示引导几何问题解决与推理的智能体框架
Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Mohammad Nehad Alam, Proma Hossain Progga, Swakkhar Shatabda
机构
*
BRAC University(布拉克大学)
;
United International University(国际联合大学)
;
Center for Computational & Data Sciences(计算与数据科学中心)
;
Independent University, Bangladesh(孟加拉国独立大学)
;
Spectrum Software & Consulting Ltd(Spectrum软件与咨询有限公司)
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
通过程序化数据合成提升大规模语言模型的空间推理能力
Zhi Helu, Huang Jingjing, Xu Wang, Xu Yangbin, Zhang Wanyue, Jiang Baoyang, Deng Shirui, Zhu Liang, Li Fangfang, Zhao Tiejun, Lin Yankai, Yao Yuan
机构
*
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
;
Tsinghua University(清华大学)
;
Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院电子学研究所)
;
Renmin University of China(中国人民大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Central South University(中南大学)
GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
GLACIA:基于多模态大语言模型的冰湖分割实例感知位置推理
Lalit Maurya, Saurabh Kaushik, Beth Tellman
机构
*
Portsmouth AI and Data Science Centre (PAIDS), School of Computing, University of Portsmouth(普森大学计算学院)
;
Center for Sustainability and the Global Environment (SAGE), University of Wisconsin–Madison(威斯康星大学麦迪逊分校可持续性与全球环境中心)
机构
*
Fudan University(复旦大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Beihang University(北京航空航天大学)
;
Shanghai Jiao Tong University(上海交通大学)