RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning
RuCL:基于分层评分标准的课程学习用于多模态大语言模型推理
Yukun Chen, Jiaming Li, Longze Chen, Ze Gong, Jingpeng Li, Zhen Qin, Hengyu Chang, Ancheng Xu, Zhihao Yang, Hamid Alinejad-Rokny, Qiang Qu, Bo Zheng, Min Yang
机构
*
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Alibaba Group(阿里巴巴集团)
;
School of Biomedical Engineering, UNSW Sydney(新南威尔士大学生物医学工程学院)
专题命中
视觉推理
:multimodal large language model(title,abstract);visual reasoning(abstract)
COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning
COVLM-RL:基于VLM引导强化学习的自动驾驶中关键对象导向推理
Lin Li, Yuxin Cai, Jianwu Fang, Jianru Xue, Chen Lv
机构
*
School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院)
;
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)
Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
Yunqi Hong, Johnson Kao, Liam Edwards, Nein-Tzu Liu, Chung-Yen Huang, Alex Oliveira-Kowaleski, Cho-Jui Hsieh, Neil Y. C. Lin
机构
*
Computer Science Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校计算机科学系)
;
Mechanical and Aerospace Engineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校机械与航空航天工程系)
;
Department of Pathology, Tri-Service General Hospital, National Defense Medical Center, Taipei, Taiwan(台湾国防医学院三军总医院病理部)
;
Department of Pathology, National Taiwan University Hospital, Taipei, Taiwan(台湾国立台湾大学医院病理部)
;
Department of Pathology, David Geffen School of Medicine, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校大卫·Geffen医学院病理部)
;
Bioengineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校生物工程系)
;
Institute for Quantitative and Computational Biosciences, University of California, CA, USA(加州大学定量与计算生物科学研究所)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
Wenhao Zhou, Hao Zheng, Rong Zhao
机构
*
Center for Brain-Inspired Computing Research (CBICR)(脑启发计算研究中心)
;
Department of Precision Instruments(精密仪器系)
;
IDG/McGovern Institute for Brain Research(IDG/麦戈文脑研究学院)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li
机构
*
Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院)
;
School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学)
;
Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)
;
School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学)
;
Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))
专题命中
视觉推理
:MLLM(title,abstract);multimodal large language model(abstract)
Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning
Zengqi Peng, Yusen Xie, Yubin Wang, Rui Yang, Qifeng Chen, Jun Ma
机构
*
Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou)(机器人与自主系统方向,香港科学与技术大学(广州))
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学)
;
Cheng Kar-Shun Robotics Institute, The Hong Kong University of Science and Technology(陈家骏机器人研究所,香港科学与技术大学)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
Zhichen Lou, Kechun Xu, Zhongxiang Zhou, Rong Xiong
机构
*
State Key Laboratory of Industrial Control Technology(工业控制技术国家重点实验室)
;
Institute of Cyber-Systems and Control(网络系统与控制研究所)
;
Zhejiang University(浙江大学)
;
Zhejiang Humanoid Robot Innovation Center Co., Ltd.(浙江人形机器人创新中心有限公司)
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
Yang Yao, Lingyu Li, Jiaxin Song, Chiyu Chen, Zhenqi He, Yixu Wang, Xin Wang, Tianle Gu, Jie Li, Yan Teng, Yingchun Wang
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The University of Hong Kong(香港大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Fudan University(复旦大学)
;
Tsinghua University(清华大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
Wenjie Li, Yujie Zhang, Haoran Sun, Yueqi Li, Fanrui Zhang, Mengzhe Xu, Victoria Borja Clausich, Sade Mellin, Renhao Yang, Chenrun Wang, Jethro Zih-Shuo Wang, Shiyi Yao, Gen Li, Yidong Xu, Hanyu Wang, Yilin Huang, Angela Lin Wang, Chen Shi, Yin Zhang, Jianan Guo, Luqi Yang, Renxuan Li, Yang Xu, Jiawei Liu, Yao Zhang, Lei Liu, Carlos Gutiérrez SanRomán, Lei Wang
机构
*
College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心)
;
School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院)
;
Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科)
;
MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室)
;
Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院)
;
Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系)
;
Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院)
;
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室)
;
Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科)
;
Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心)
;
Ningbo Institute of Technology, Beihang University(北航宁波理工学院)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving
Haichao Liu, Haoren Guo, Pei Liu, Benshan Ma, Yuxiang Zhang, Jun Ma, Tong Heng Lee
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
National University of Singapore(新加坡国立大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)