arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 159 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 159 篇

2604.27792 2026-07-16 cs.RO 版本更新 50%

Motubrain: An Advanced World Action Model for Robot Control

MotuBrain: 一种先进的世界动作模型用于机器人控制

Motubrain Team, Chendong Xiang, Fan Bao, Haitian Liu, Hengkai Tan, Hongzhe Bi, James Li, Jiabao Liu, Jingrui Pang, Kiro Jing, Louis Liu, Mengchen Cai, Rongxu Cui, Ruowen Zhao, Runqing Wang, Shuhe Huang, Yao Feng, Yinze Rong, Zeyuan Wang, Jun Zhu

机构 * MotuBrain Team(MotuBrain团队)

专题命中 视频多模态 :multimodal(abstract)

AI总结 MotuBrain提出一种统一的世界动作模型,结合视频和动作,支持多项任务,实现高效部署和高精度控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23087 2026-07-07 cs.RO 版本更新 50%

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation

CoLA-Flow Policy: 通过连续潜在动作流匹配实现机器人操作的时序一致模仿学习

Wu Songwei, Jiang Zhiduo, Sun Wandong, Xie Guanghu, Zhao Rui, Liu Hong, Liu Yang

机构 * State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学) The University of Sydney(悉尼大学) Honor Device Co., Ltd.(荣耀设备有限公司)

专题命中 视频多模态 :multimodal(abstract)

AI总结 本文提出CoLA-Flow Policy,一种基于连续潜在动作空间的轨迹级模仿学习框架,通过学习显式的潜在空间流,解耦全局运动结构与低层控制噪声,从而实现平滑可靠的长时程执行,并结合几何感知点云条件和执行时多模态调节,提升现实环境的鲁棒性。

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27039 2026-07-07 stat.ME 版本更新 50%

Measuring Human Behavior Through Controlled Perturbations: A Framework for Behavioral System Identification

通过受控扰动测量人类行为:行为系统识别的框架

Pietro Cipresso

专题命中 视频多模态 :multimodal(abstract)

AI总结 本文提出基于受控扰动的行为测量框架,通过动态系统视角重构测量问题,整合多模态数据与沉浸技术,推动从描述性模型向生成行为机制识别的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11475 2026-07-07 cs.LG cs.NI 版本更新 50%

Deep Learning Network-Temporal Models For Traffic Prediction

用于流量预测的深度学习网络-时间模型

Yufeng Xin, Ethan Fan

专题命中 视频多模态 :multi-modal(abstract)

AI总结 针对多元时间序列预测问题,提出拓扑感知学习框架,用图注意力模型捕捉相关性,评估微调大语言模型,引入聚类预处理,提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06480 2026-06-26 cs.RO 版本更新 50%

History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation

基于历史条件的时空视觉令牌剪枝用于高效视觉-语言导航

Qitong Wang, Yijun Liang, Ming Li, Tianyi Zhou, Christopher Rasmussen

机构 * Department of Computer and Information Sciences at the University of Delaware(德克萨斯大学达勒姆分校计算机与信息科学系) University of Maryland’s Department of Computer Science(马里兰大学计算机科学系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出一种无需训练的时空视觉令牌剪枝框架,通过空间令牌选择和时空压缩减少冗余计算,在保持导航精度的同时显著提升推理效率,并在真实机器人上验证了低延迟指令跟随导航。

Comments International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22126 2026-06-23 cs.RO 版本更新 50%

EasyUUV: An LLM-Enhanced Universal and Lightweight Sim-to-Real Reinforcement Learning Framework for UUV Attitude Control

EasyUUV:一种用于UUV姿态控制的LLM增强通用轻量级仿真到现实强化学习框架

Guanwen Xie, Jingzehua Xu, Jiwei Tang, Yubo Huang, Zixi Wang, Shuai Zhang, Dongfang Ma, Juntian Qu, Xiaofan Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Department of Mechanical Engineering, The University of Hong Kong(香港大学机械工程系) Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系) School of Civil Engineering, Southwest Jiaotong University(西南交通大学土木工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) Department of Data Science, New Jersey Institute of Technology(新泽西理工学院数据科学系) Ocean College, Zhejiang University(浙江大学海洋学院)

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出EasyUUV框架,结合并行强化学习与混合控制架构,并集成多模态大语言模型自适应调整参数,实现UUV姿态控制的鲁棒性和泛化性,经仿真与实船验证。

Comments 12 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23340 2026-06-23 cs.SI cs.DC cs.LG 版本更新 50%

CrediBench: Building Web-Scale Network Datasets for Information Integrity

CrediBench: 构建用于信息完整性的网络规模数据集

Emma Kondrup, Sebastian Sabry, Hussein Abdallah, Zachary Yang, Jiaqi Xiong, Kellin Pelrine, James Zhou, Zhijin Guo, Michael M. Bronstein, Jean-François Godbout, Reihaneh Rabbany, Shenyang Huang

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) University of Oxford(牛津大学) University of California, Berkeley(加州大学伯克利分校) AITHYRA Research Institute(AITHYRA研究院) Université de Montréal(蒙特利尔大学)

专题命中 视频多模态 :multi-modal(abstract)

AI总结 针对现有数据集忽略网络拓扑、时序和文本内容等关键模态的问题,提出包含八个月网络图数据的CrediBench数据集,支持回归和分类任务,多模态模型显著提升性能。

Comments 16 pages,4 figures

Journal ref KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27583 2026-06-17 q-bio.NC cs.RO 版本更新 50%

Simulating Infant First-Person Sensorimotor Experience via Motion Retargeting from Babies to Humanoids

通过从婴儿到类人机器人的运动重定向模拟婴儿第一人称感觉运动经验

Francisco M. López, Hoshinori Kanazawa, Ondrej Fiala, Yakov Balashov, Valentin Marcel, Lukas Rustler, Miles Lenz, Dongmin Kim, Yasuo Kuniyoshi, Jochen Triesch, Matej Hoffmann

机构 * German Research Foundation(德国研究基金会) Johanna Quandt foundation(约翰娜·克文特基金会) Czech Science Foundation(捷克科学基金)

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出一种从单视频重建婴儿3D姿态并映射到物理/虚拟类人平台的方法,实现亚厘米级精度的多感觉流模拟,为发育研究和神经发育障碍早期检测提供新工具。

Comments Accepted at IEEE ICDL 2026. 8 pages, 6 figures. Cite as: F. M. López, H. Kanazawa, O. Fiala, Y. Balashov, V. Marcel, L. Rustler, M. Lenz, D. Kim, Y. Kuniyoshi, J. Triesch, and M. Hoffmann, "Simulating infant first-person sensorimotor experience via motion retargeting from babies to humanoids'', in 2026 IEEE International Conference on Development and Learning (ICDL). IEEE, 2026, pp. 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09580 2026-06-08 cs.RO cs.LG 版本更新 50%

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows

SERNF: 通过动作块评论家和归一化流实现样本高效的真实世界灵巧策略微调

Chenyu Yang, Denis Tarasov, Davide Liconti, Romain Guntz, Hehui Zheng, Robert K. Katzschmann

机构 * Soft Robotics Lab, D-MAVT(软机器人实验室,D-MAVT) ETH Zurich(苏黎世联邦理工学院)

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出SERNF框架,结合归一化流策略和动作块评论家,实现真实世界灵巧操作策略的样本高效微调,解决多模态动作分布和信用分配问题。

Comments https://srl-ethz.github.io/SERNF/

详情

展开后加载摘要…

URL PDF HTML 收藏