机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
MIT-IBM Watson AI Lab, IBM Research(MIT-IBM沃森人工智能实验室,IBM研究)
;
Stony Brook University(石溪大学)
;
Brookhaven National Laboratory(布鲁赫斯国家实验室)
专题命中
视觉推理
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
VisualQuest: 一个用于多模态大语言模型(MLLMs)抽象视觉推理的基准数据集
Kelaiti Xiao, Liang Yang, Dongyu Zhang, Paerhati Tulajiang, Hongfei Lin
机构
*
School of Computer Science and Technology, Dalian University of Technology, Dalian, China(大连理工大学计算机科学与技术学院)
;
School of Foreign Languages, Dalian University of Technology, Dalian, China(大连理工大学外语学院)
;
School of Computer Science and Technology, Xinjiang Normal University, Urumqi, China(新疆师范大学计算机科学与技术学院)
专题命中
视觉推理
:visual reasoning(title,abstract);multimodal large language model(abstract);分类 cs.CV
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
R4:在4D时空空间中为视觉语言模型引入检索增强推理
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Esslingen University of Applied Sciences(埃斯林根应用科学大学)
;
Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业协会)
;
University of Michigan(密歇根大学)
;
Voxel51 Inc.(Voxel51公司)
Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
合成血管和病理增强视觉-语言模型推理
Chenjun Li, Cheng Wan, Laurin Lux, Alexander Berger, Richard B. Rosen, Martin J. Menten, Johannes C. Paetzold
机构
*
Cornell University(康奈尔大学)
;
Weill Cornell Medicine(韦尔·康奈尔医学)
;
Technical University of Munich(慕尼黑技术大学)
;
New York Eye and Ear Infirmary of Mount Sinai(圣文森特医院)
;
Cornell Tech(康奈尔科技)
Comments6 pages, 3 figures. Code and data: https://github.com/Amiton7/Tri-Bench. Accepted to the AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
MindDrive: 一个整合世界模型和视觉-语言模型的全功能框架,用于端到端自动驾驶
Bin Sun, Yaoguang Cao, Yan Wang, Rui Wang, Jiachen Shang, Xiejie Feng, Jiayi Lu, Jia Shi, Shichun Yang, Xiaoyu Yan, Ziying Song
机构
*
School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
;
State Key Laboratory of Intelligent Transportation System, Beihang University(北京航空航天大学智能交通系统国家重点实验室)
;
Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新院)
;
Contemporary Amperex Technology Co., Limited (CATL)(当代电动车技术有限公司(CATL))
;
Research Institute of Aero-Engine, Beihang University(北京航空航天大学航空发动机研究院)
;
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
;
China Automotive Engineering Research Institute Co., Ltd.(中国汽车工程研究院股份有限公司)
机构
*
Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)
;
School of Biomedical Engineering, Southern Medical University(南方医科大学生物医学工程学院)
;
Singapore University of Technology and Design(新加坡科技与设计大学)
;
University of Auckland(奥克兰大学)
;
University of Science and Technology of China(中国科学技术大学)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
College of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
机构
*
Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院)
;
Shanghai Collaborative Innovation Center on Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
MiniMax
Text to Robotic Assembly of Multi Component Objects using 3D Generative AI and Vision Language Models
通过3D生成AI和视觉语言模型实现多组件物体的文本到机器人组装
Alexander Htet Kyaw, Richa Gupta, Dhruv Shah, Anoop Sinha, Kory Mathewson, Stefanie Pender, Sachin Chitta, Yotto Koga, Faez Ahmed, Lawrence Sass, Randall Davis
机构
*
Massachusetts Institute of Technology (MIT)(麻省理工学院)
;
MIT(麻省理工学院)
;
Google DeepMind(谷歌DeepMind)
;
Google, Paradigms of Intelligence(谷歌、范式智能)
;
Autodesk Research(Autodesk研究)
;
MIT Mechanical Engineering(麻省理工学院机械工程系)
;
MIT Architecture(麻省理工学院建筑系)
;
MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中
视觉推理
:vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI