Comments8 pages, 4 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). Project page: https://radar-iros.netlify.app/
机构
*
South China University of Technology(华南理工大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Peking University(北京大学)
;
University of Electronic Science and Technology of China(电子科技大学)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系)
;
Department of Radiology, Renmin Hospital of Wuhan University(武汉大学仁民医院放射科)
;
Shanghai Sixth People’s Hospital Affiliated to Shanghai Jiao Tong University(上海交通大学附属第六人民医院)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
DeepTumorVQA: 一种分层的3D CT基准,用于分阶段评估医学视觉语言模型和工具增强代理
Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi, Boyan Wang, Liang He, Xinze Zhou, Sezgin Er, Ibrahim Ethem Hamamci, Zongwei Zhou, Alan Yuille
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
University of Bologna(博洛尼亚大学)
;
Istanbul Medipol University(伊斯坦布尔梅迪波尔大学)
;
Center for Biomolecular Nanotechnologies, Istituto Italiano di Tecnologia(生物分子纳米技术中心,意大利技术研究院)
;
The First Affiliated Hospital, Sun Yat-Sen University(中山大学第一附属医院)
;
Tongji University(同济大学)
When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
当话语压力冲突时:视觉-语言模型输出中的信息结构
Marcell Fekete, Johannes Bjerva, Tamás Káldi
机构
*
Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系)
;
Department of Psycholinguistics and Neurolinguistics, ELTE Research Centre for Linguistics(ELTE语言研究中心心理学语言学与神经语言学系)
;
ELTE Bárczi Gusztáv Faculty of Special Needs Education(ELTE巴尔茨吉斯塔夫特殊教育学院)
Protecting multimodal large language models against misleading visualizations
保护多模态大语言模型免受误导性可视化的影响
Jonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna Gurevych
机构
*
Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普遍知识处理实验室(UKP实验室)、计算机科学系,德累斯顿技术大学及应用网络安全国家研究中心ATHENE)
;
Department of Electrical Engineering, KU Leuven(电气工程系,鲁文大学)
;
Department of Computer Science, KU Leuven(计算机科学系,鲁文大学)
专题命中
视觉问答
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
Varun Srivastava, Fan Lei, Srija Mukhopadhyay, Vivek Gupta, Ross Maciejewski
机构
*
School of Computing and Augmented Intelligence(计算与增强智能学院)
;
Arizona State University(亚利桑那州立大学)
;
Department of Computer Science(计算机科学系)
;
International Institute of Information Technology(国际信息科技研究所)
专题命中
视觉问答
:multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
CommentsPublished as a conference paper at COLM 2025
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
Qi Yang, Chenghao Zhang, Lubin Fan, Kun Ding, Jieping Ye, Shiming Xiang
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences, China(中国科学院大学人工智能学院)
;
MAIS, Institute of Automation, Chinese Academy of Sciences, China(中国科学院自动化研究所MAIS部)
;
Alibaba Cloud Computing, China(阿里巴巴云计算)
专题命中
视觉问答
:vision-language model(title);vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering
Songtao Jiang, Chenyi Zhou, Yan Zhang, Yeying Jin, Zuozhu Liu
机构
*
Zhejiang University(浙江大学)
;
Byte Dance(字节跳动)
;
National University of Singapore(新加坡国立大学)
;
Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江省医学影像人工智能重点实验室)
专题命中
视觉问答
:visual question answering(title,abstract);multimodal large language model(abstract);MLLM(abstract)