arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-18 至 2026-03-18 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2603.16777 2026-03-18 cs.AI 79%

Anticipatory Planning for Multimodal AI Agents

多模态AI代理的前瞻性规划

Yongyuan Liang, Shijie Zhou, Yu Gu, Hao Tan, Gang Wu, Franck Dernoncourt, Jihyung Kil, Ryan A. Rossi, Ruiyi Zhang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) The Ohio State University(俄亥俄州立大学) Adobe Research(Adobe研究) State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出TraceR1框架,通过前瞻性轨迹推理提升多模态代理的规划能力,实现更稳定的执行和泛化性能。

Comments Published at CVPR 2026 Findings Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10376 2026-03-18 cs.CV cs.RO 79%

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

MSGNav: 释放多模态3D场景图在零样本具身导航中的潜力

Xun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu, Wanfa Zhang, Rongsheng Qu, Weixin Li, Yunhong Wang, Chenglu Wen

机构 * Fujian Key Laboratory of Urban Intelligent Sensing and Computing(福建智能感知与计算重点实验室) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(多媒体可信感知与高效计算重点实验室) Beihang University(北航) Nanyang Technological University(南洋理工大学) University of Chinese Academy of Sciences(中国科学院大学) Zhongguancun Academy(中关村学院)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

AI总结 MSGNav通过多模态3D场景图实现零样本具身导航,引入关键子图选择模块、自适应词汇更新模块和闭环推理模块,解决开放词汇和视觉证据保留问题,实验表明其在GOAT-Bench和HM3D-ObjNav基准上表现优异。

Comments 18 pages, Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14440 2026-03-18 cs.AI cs.CL cs.LG 76%

VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integration

VisTIRA: 通过结构化工具集成缩小图像-文本模态差距以提升视觉数学推理

Saeed Khaki, Ashudeep Singh, Nima Safaei, Kamal Ginotra

机构 * Microsoft AI(微软人工智能)

专题命中 多模态Agent :image-text(title);分类 cs.CL、cs.AI

AI总结 VisTIRA通过结构化工具集成框架,解决图像形式数学问题的推理难题,改进视觉数学推理能力,实验表明工具监督和OCR定位能有效缩小模态差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16664 2026-03-18 cs.CV cs.AI 62%

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Kestrel: 为降低LVLM幻觉而引入自反思

Jiawei Mao, Hardy Chen, Haoqin Tu, Yuhan Wang, Letian Zhang, Zeyu Zheng, Huaxiu Yao, Zirui Wang, Cihang Xie, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) UC Berkeley(加州大学伯克利分校) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Apple(苹果公司)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Kestrel提出一种无需训练的框架,通过显式视觉 grounding 与证据验证自反思机制减少LVLM幻觉,实验显示在POPE和MME-Hallucination基准上性能提升,同时提供透明的验证轨迹。

Comments 16 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21676 2026-03-18 cs.RO cs.NI 50%

Real-World Deployment of Cloud-based Autonomous Mobility Systems for Outdoor and Indoor Environments

云原生自主移动系统在户外和室内环境中的实际部署

Yufeng Yang, Minghao Ning, Keqi Shu, Aladdin Saleh, Ehsan Hashemi, Amir Khajepour

机构 * Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系) Technology Partnerships and Innovations, Rogers Communications, Canada Inc.(罗杰斯通讯加拿大有限公司技术伙伴关系与创新部) Mechanical Engineering Department, University of Alberta(阿尔伯塔大学机械工程系)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出云原生自主移动框架,通过基础设施智能传感与云计算协调提升自主操作能力,实验证明在城市环岛和医院类室内环境中的感知鲁棒性和安全性提升。

Comments This paper has been submitted to IEEE Robotics and Automation Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18373 2026-03-18 cs.RO cs.HC 50%

UGotMe: An Embodied System for Affective Human-Robot Interaction

UGotMe: 一种用于情感人机交互的具身系统

Peizhen Li, Longbing Cao, Xiao-Ming Wu, Xiaohan Yu, Runze Yang

机构 * School of Computing, Macquarie University(麦考瑞大学计算学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Department of Automation, Shanghai Jiao Tong University(上海交通大学自动化学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出UGotMe系统,解决多对话场景中视觉噪声和实时响应问题,通过去噪策略和高效数据传输提升情感识别能力。

Comments Accepted to the 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏