机构
*
Tsinghua University(清华大学)
;
Hefei University of Technology(合肥工业大学)
;
University of Arizona(亚利桑那大学)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学计算机软件新技术国家重点实验室)
;
School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院)
;
AI Business, Alibaba Group(阿里巴巴集团智能计算研究院)
;
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, China(上海交通大学自动化与智能感知学院)
;
School of Intelligent Systems Engineering, Sun Yat-sen University, Shenzhen, China(中山大学智能工程学院(深圳))
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院)
;
School of Software Technology, Zhejiang University, China(浙江大学软件学院)
;
School of Computing, National University of Singapore, Singapore, Singapore(新加坡国立大学计算机学院)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding
ChemVLR:在化学视觉语言理解中优先考虑推理
Xuanle Zhao, Xinyuan Cai, Xiang Cheng, Xiuyi Chen, Bo Xu
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
Evidence-Based Actor-Verifier Reasoning for Echocardiographic Agents
基于证据的Actor-验证者推理用于超声波代理
Peng Huang, Yiming Wang, Yineng Chen, Liangqiao Gui, Hui Guo, Bo Peng, Shu Hu, Xi Wu, Tsao Connie, Hongtu Zhu, Balakrishnan Prabhakaran, Xin Wang
机构
*
University at Albany, SUNY(纽约州立大学奥尔巴尼分校)
;
Southwest Jiaotong University(西南交通大学)
;
Purdue University(普渡大学)
;
Chengdu University of Information Technology(成都信息工程大学)
;
Harvard Medical School(哈佛医学院)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中
视觉推理
:VLM(abstract);visual language model(abstract);分类 cs.CV
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
多模态大语言模型中的忠实优先推理、规划与行动
Junxian Li, Xinyue Xu, Sai Ma, Di Zhang, Sichao Li
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
The Australian National University(澳大利亚国立大学)
;
Fudan University(复旦大学)
;
University of Sydney(悉尼大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
CommentsThis version is withdrawn to consolidate the submission under the corresponding author's primary account. The most recent and maintained version of this work can be found at arXiv:2603.09246
Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework
聚焦低空治理中真正重要的问题:一种以管理为中心的多模态基准与隐式协调的视觉-语言推理框架
Hao Chang, Zhihui Wang, Lingxiang Wu, Wei An, Boyang Li, Zaiping Lin, Weidong Sheng, Jinqiao Wang
机构
*
National University of Defense Technology(国防科技大学)
;
Zidong Taichu (Beijing) Technology Co., Ltd.(紫东太初(北京)科技有限公司)
;
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Wuhan AI Research(武汉人工智能研究院)