arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-05 至 2026-03-05 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 7 篇

2508.03099 2026-03-05 cs.RO 82%

Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping

Point2Act: 多模态大语言模型的高效3D蒸馏用于零样本情境感知抓取

Sang Min Kim, Hyeongjun Heo, Junho Kim, Yonghyeon Lee, Young Min Kim

机构 * Seoul National University(首尔国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 Point2Act通过多模态大语言模型高效蒸馏实现零样本情境感知抓取,生成空间定位响应以支持实际操作任务。

Comments Accepted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04349 2026-03-05 cs.CV 70%

FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

FocusGraph: 用于具身长视频问答的图结构帧选择

Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin

机构 * AXXX MIRAI Yandex FusionBrain Lab(FusionBrain实验室)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 FocusGraph通过图结构帧选择和轻量级模型实现高效长视频问答,提升性能并减少推理时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04222 2026-03-05 cs.RO cs.AI 57%

PRAM-R: A Perception-Reasoning-Action-Memory Framework with LLM-Guided Modality Routing for Adaptive Autonomous Driving

PRAM-R:一种具有LLM引导模态路由的感知-推理-行动-记忆框架,用于自适应自动驾驶

Yi Zhang, Xian Zhang, Saisi Zhao, Yinglei Song, Chengdong Wu, Nenad Petrovic, Alois Knoll

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 PRAM-R通过LLM引导的模态路由和分层记忆模块,实现高效自适应的多模态感知,提升自动驾驶的稳定性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04057 2026-03-05 cs.RO cs.AI 57%

Sim2Sea: Sim-to-Real Policy Transfer for Maritime Vessel Navigation in Congested Waters

Sim2Sea: 用于拥挤水域船舶导航的仿真到现实策略迁移

Xinyu Cui, Xuanfa Jin, Xue Yan, Yongcheng Zeng, Luoyang Sun, Siying Wei, Ruizhi Zhang, Jian Zhao, Haifeng Zhang, Jun Wang

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Zhongguancun Academy(中关村学院) University College London(伦敦大学学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 Sim2Sea通过GPU加速模拟器、双流时空策略和领域随机化方案,实现了拥挤水域中船舶自主导航的仿真到现实策略迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03985 2026-03-05 cs.CV 57%

RIVER: A Real-Time Interaction Benchmark for Video LLMs

RIVER:面向视频大语言模型的实时交互基准

Yansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng, Yi Wang, Limin Wang

机构 * School of Information Science And Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学) State Key Lab of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 RIVER基准通过引入实时交互任务框架,改进视频大语言模型的实时交互能力,提升长期记忆与未来感知表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16677 2026-03-05 cs.CV cs.LG cs.RO eess.IV 57%

Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence

基于动作的视频分割:面向具身智能的标签噪声鲁棒动作引导分割

Wenxin Li, Kunyu Peng, Di Wen, Ruiping Liu, Mengfei Duan, Kai Luo, Kailun Yang

机构 * School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(人工智能与机器人学院及机器人视觉感知与控制技术国家工程研究中心,湖南大学,中国) Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology, Germany(人机学与机器人研究所,卡尔斯鲁厄理工学院,德国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本研究首次探索基于动作的视频对象分割在标签噪声下的鲁棒性,引入两种噪声类型并提出并行掩码头机制以提升分割性能。

Comments Accepted to ICRA 2026. The established benchmark and source code will be made publicly available at https://github.com/mylwx/ActiSeg-NL

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03957 2026-03-05 cs.RO 50%

ArthroCut: Autonomous Policy Learning for Robotic Bone Resection in Knee Arthroplasty

ArthroCut:膝关节置换术中机器人骨切除的自主策略学习

Xu Lu, Yiling Zhang, Wenquan Cheng, Longfei Ma, Fang Chen, Hongen Liao

机构 * School of Biomedical Engineering, Tsinghua University(清华大学生物医学工程学院) School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院) Longwood Valley MedTech

专题命中 视频多模态 :multimodal(abstract)

AI总结 ArthroCut通过结合术前和术中数据,实现膝关节置换术中骨切除的自主策略学习,提升机器人手术的自主性和可解释性。

Comments Accepted for publication at the 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏