Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education
视觉推理基准:评估多模态大语言模型在小学课堂真实视觉问题上的能力
Mohamed Huti, Alasdair Mackintosh, Amy Waldock, Dominic Andrews, Maxime Lelièvre, Moritz Boos, Tobias Murray, Paul Atherton, Robin A. A. Ince, Oliver G. B. Garrod
机构
*
Fab AI
专题命中
视觉推理
:visual reasoning(title,abstract);multimodal large language model(abstract);分类 cs.AI
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
Ground-R1: 通过强化学习激励基于地面的视觉推理
Meng Cao, Haoze Zhao, Can Zhang, Xiaojun Chang, Ian Reid, Xiaodan Liang
机构
*
Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学)
;
Peking University(北京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Sun Yat-sen University(中山大学)
Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis
Med3D-R1: 促进3D医学视觉-语言模型的临床推理
Haoran Lai, Zihang Jiang, Kun Zhang, Qingsong Yao, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Wei Wei, Shaohua Kevin Zhou
机构
*
School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China(生物医学工程学院,生命科学与医学系,中国科学技术大学)
;
Suzhou Institute for Advanced Research, University of Science and Technology of China(先进研究所,中国科学技术大学)
;
Stanford University(斯坦福大学)
;
Medical Business Department, iFlytek Co.Ltd(iFlytek公司医学业务部)
;
The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine University of Science and Technology of China(中国科学技术大学第一附属医院,生命科学与医学系)
Investigating the Development of Task-Oriented Communication in Vision-Language Models
探究视觉-语言模型中以任务为导向的交流发展
Boaz Carmeli, Orr Paradise, Shafi Goldwasser, Yonatan Belinkov, Ron Meir
机构
*
Technion – Israel Institute of Technology(技术ion–以色列理工学院)
;
EPFL(苏黎世联邦理工学院)
;
University of California, Berkeley(加州大学伯克利分校)
;
Harvard University(哈佛大学)
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
MIT-IBM Watson AI Lab, IBM Research(MIT-IBM沃森人工智能实验室,IBM研究)
;
Stony Brook University(石溪大学)
;
Brookhaven National Laboratory(布鲁赫斯国家实验室)
专题命中
视觉推理
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
VisualQuest: 一个用于多模态大语言模型(MLLMs)抽象视觉推理的基准数据集
Kelaiti Xiao, Liang Yang, Dongyu Zhang, Paerhati Tulajiang, Hongfei Lin
机构
*
School of Computer Science and Technology, Dalian University of Technology, Dalian, China(大连理工大学计算机科学与技术学院)
;
School of Foreign Languages, Dalian University of Technology, Dalian, China(大连理工大学外语学院)
;
School of Computer Science and Technology, Xinjiang Normal University, Urumqi, China(新疆师范大学计算机科学与技术学院)
专题命中
视觉推理
:visual reasoning(title,abstract);multimodal large language model(abstract);分类 cs.CV
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
R4:在4D时空空间中为视觉语言模型引入检索增强推理
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Esslingen University of Applied Sciences(埃斯林根应用科学大学)
;
Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业协会)
;
University of Michigan(密歇根大学)
;
Voxel51 Inc.(Voxel51公司)
Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
合成血管和病理增强视觉-语言模型推理
Chenjun Li, Cheng Wan, Laurin Lux, Alexander Berger, Richard B. Rosen, Martin J. Menten, Johannes C. Paetzold
机构
*
Cornell University(康奈尔大学)
;
Weill Cornell Medicine(韦尔·康奈尔医学)
;
Technical University of Munich(慕尼黑技术大学)
;
New York Eye and Ear Infirmary of Mount Sinai(圣文森特医院)
;
Cornell Tech(康奈尔科技)
Comments6 pages, 3 figures. Code and data: https://github.com/Amiton7/Tri-Bench. Accepted to the AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
MindDrive: 一个整合世界模型和视觉-语言模型的全功能框架,用于端到端自动驾驶
Bin Sun, Yaoguang Cao, Yan Wang, Rui Wang, Jiachen Shang, Xiejie Feng, Jiayi Lu, Jia Shi, Shichun Yang, Xiaoyu Yan, Ziying Song
机构
*
School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
;
State Key Laboratory of Intelligent Transportation System, Beihang University(北京航空航天大学智能交通系统国家重点实验室)
;
Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新院)
;
Contemporary Amperex Technology Co., Limited (CATL)(当代电动车技术有限公司(CATL))
;
Research Institute of Aero-Engine, Beihang University(北京航空航天大学航空发动机研究院)
;
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
;
China Automotive Engineering Research Institute Co., Ltd.(中国汽车工程研究院股份有限公司)