机构
*
Alibaba Group(阿里巴巴集团)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
University of Alberta(阿尔伯塔大学)
;
Zhejiang University(浙江大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
大型视觉语言模型(LVLMs)能否揭示视觉错觉背后的真相?感知与推理能力分析
Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang, Rong Yin, Yulan Hu, Da Li, Jianing Li
机构
*
Adelaide University(阿德莱德大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
Amap, Alibaba Group(阿里巴巴集团高德地图)
;
Beihang University(北京航空航天大学)
专题命中
视觉推理
:vision language model(abstract);分类 cs.AI
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
评估具有挑战性的真实世界临床病例中的多轮多模态诊断推理
Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicolás Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu, Yifan Peng
机构
*
Center for Biomedical Data Science, Duke-NUS Medical School(生物医学数据科学中心,杜克 - 新加坡国立大学医学院)
;
Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School(杜克 - 新加坡国立大学人工智能与医学科学计划,杜克 - 新加坡国立大学医学院)
;
Department of Population Health Sciences, Weill Cornell Medicine(人口健康科学系,威尔康奈尔医学院)
;
System Engineering, College of Engineering, Cornell University(系统工程,康奈尔大学工程学院)
;
Graduate School of Frontier Sciences, The University of Tokyo(前沿科学研究生院,东京大学)
;
RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心)
;
Department of Biostatistics and Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学)
;
Perelman School of Medicine, University of Pennsylvania(佩雷尔曼医学院,宾夕法尼亚大学)
;
School of Clinical Medicine, University of Cambridge(临床医学学院,剑桥大学)
;
Hospital of the University of Pennsylvania(宾夕法尼亚大学医院)
;
Children’s Hospital of Philadelphia (CHOP)(费城儿童医院)
;
Department of Anatomic Pathology and Laboratory Medicine, Hospital of the University of Pennsylvania(解剖病理学与检验医学系,宾夕法尼亚大学医院)
;
Department of Pathology and Laboratory Medicine, University of California, San Francisco(病理学与检验医学系,加州大学旧金山分校)
;
Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件学院)
;
Harbin Institute of Technology Shenzhen(哈尔滨工业大学(深圳))
;
City University of Macau(澳门城市大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
ViSTR-Bench:多模态大语言模型能否从动态场景中的连续视觉线索进行推理?
Han Li, Si Liu, Zehao Huang, Dongxin Lyu, Longfei Xu, Jiahui Fu, Daxin Tian, Yuliang Xiu, Naiyan Wang
机构
*
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
Zhongguancun Academy(中关村科学城)
;
School of Engineering, Westlake University(西湖大学工学院)
;
School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
机构
*
School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
School of Computer Science, Peking University(北京大学计算机科学学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV