arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2509.22295 2026-03-10 cs.LG 78%

Aurora: Towards Universal Generative Multimodal Time Series Forecasting

Aurora:迈向通用生成多模态时间序列预测

Xingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen, Yang Shu, Bin Yang, Chenjuan Guo

机构 * East China Normal University(东华师范大学)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 Aurora通过多模态输入和零样本推理,提升跨领域时间序列预测的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09095 2026-03-09 cs.LG 78%

Temporal Misalignment Attacks against Multimodal Perception in Autonomous Driving

针对自动驾驶多模态感知的时序错位攻击

Md Hasan Shahriar, Md Mohaimin Al Barat, Harshavardhan Sundar, Ning Zhang, Naren Ramakrishnan, Y. Thomas Hou, Wenjing Lou

机构 * Amazon.com, Inc.(亚马逊公司) Washington University in St. Louis(华盛顿大学)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出DejaVu攻击,通过制造时序错位破坏自动驾驶多模态感知,导致目标检测和跟踪性能显著下降。

Comments 19 pages, 18 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02569 2026-03-04 cs.HC 78%

An LLM-Assisted Toolkit for Inspectable Multimodal Emotion Data Annotation

一种辅助多模态情绪数据标注的工具包

Zheyuan Kuang, Weiwei Jiang, Nicholas Koemel, Matthew Ahmadi, Emmanuel Stamatakis, Benjamin Tag, Anusha Withana, Zhanna Sarsenbayeva

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出一种LLM辅助的多模态情绪数据标注工具包,通过可检查的工作流实现细粒度标注,提升跨模态一致性检查效率。

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00110 2026-03-03 cs.RO 78%

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

从预训练视频模型中学习物理:一种多模态连续和序列世界交互模型用于机器人操作

Zijian Song, Qichang Li, Sihan Qin, Yuhao Chen, Tianshui Chen, Liang Lin, Guangrun Wang

机构 * Sun Yat-sen University(中山大学) Guangdong Key Laboratory of Big Data Analysis(广东大数据分析与处理重点实验室) X-Era AI Lab(X-Era人工智能实验室) Guangdong University of Technology(广东工业大学)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 PhysGen通过预训练视频模型学习物理知识,实现机器人操作的连续和序列世界交互,优于现有基线方法。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17068 2026-02-20 cs.LG cs.SY eess.SY 78%

Spatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control

时空双阶段超图多智能体强化学习用于以人为中心的多模式走廊交通信号控制

Xiaocai Zhang, Neema Nassir, Milad Haghani

机构 * Department of Infrastructure Engineering, Faculty of Engineering and Information Technology, The University of Melbourne, VIC 3010, Australia(工程与信息技术学院基础设施工程系,墨尔本大学)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出STDSH-MARL方法,通过时空双阶段超图注意力机制和混合动作空间,提升走廊交通信号对多模式旅行者和公共交通的适应性与优先级。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18891 2026-02-18 cs.LG 78%

CAAT-EHR: Cross-Attentional Autoregressive Transformer for Multimodal Electronic Health Record Embeddings

CAAT-EHR: 多模态电子健康记录嵌入的交叉注意力自回归变压器

Mohammad Al Olaimat, Shaika Chowdhury, Serdar Bozdag

机构 * Department of Computer Science and Engineering, University of North Texas(计算机科学与工程系,北卡罗来纳大学达顿分校) BioDiscovery Institute, University of North Texas(生物发现研究所,北卡罗来纳大学达顿分校) Center for Computational Life Sciences, University of North Texas(计算生命科学中心,北卡罗来纳大学达顿分校)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 CAAT-EHR通过多模态嵌入和自回归机制提升EHR分析的通用性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.12433 2026-02-17 cs.RO cs.SY eess.SY 78%

Model Predictive Control with Gaussian Processes for Flexible Multi-Modal Physical Human Robot Interaction

基于高斯过程的模型预测控制用于柔性多模态人机协作交互

Kevin Haninger, Christian Hegeler, Luka Peternel

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 本文提出基于高斯过程的模型预测控制方法,用于多模态人机协作交互,通过贝叶斯推断和在线控制提升任务灵活性和效率。

Comments Submitted, ICRA 2022. Video: https://youtu.be/0GT1pPpXvt8 Data and code: https://owncloud.fraunhofer.de/index.php/s/kmCZvlKOghclHy9

Journal ref 2022 IEEE International Conference on Robotics and Automation (ICRA), May 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13459 2026-02-17 eess.SP 78%

Towards Causality-Aware Modeling for Multimodal Brain-Muscle Interactions

迈向因果意识的多模态脑-肌相互作用建模

Farwa Abbas, Wei Dai, Zoran Cvetkovic, Verity McClelland

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出一种结合几何流形重建与概率时间建模的DBN启发CCM框架,用于多模态脑-肌交互的因果建模,揭示了肌张力障碍中特定频率的通路重组织,并展示了其在生物标志物开发和神经调节干预中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10897 2026-02-12 hep-ex 78%

Multi-Modal Track Reconstruction using Graph Neural Networks at Belle II

利用图神经网络进行多模态轨迹重建:Belle II实验

Lea Reuter, Tristan Brandes, Giacomo De Pietro, Torben Ferber

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 Belle II实验利用图神经网络和多模态输入提升轨迹重建效率和纯度,通过模拟验证效果。

Comments proceedings for ACAT25

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06303 2026-02-11 cs.LG 78%

Multimodal Graph Neural Networks for Prognostic Modeling of Brain Network Reorganization

多模态图神经网络用于脑网络重组的预后建模

Preksha Girish, Rachana Mysore, Kiran K. N., Hiranmayee R., Shipra Prashanth, Shrey Kumar

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出多模态图神经网络用于脑网络重组的预后建模,通过整合多种影像数据,生成可解释的生物标志物以预测认知下降风险。

Comments Fundamental methodological error invalidating results

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07388 2026-02-10 cs.RO 78%

Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation

面向长时程机器人操作的多模态动作消歧的轨迹聚焦扩散策略

Yuxuan Hu, Xiangyu Chen, Chuhao Zhou, Yuxi Liu, Gen Li, Jindou Jia, Jianfei Yang

机构 * MARS Lab, Nanyang Technological University(MARS实验室,南洋理工大学)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 TF-DP通过轨迹聚焦解决长时程机器人操作中的多模态动作模糊问题,提升时间一致性和鲁棒性,实验结果优于普通扩散策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14550 2026-01-22 cs.RO 78%

TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks

TacUMI: 一种多模态通用操控接口用于接触密集型任务

Tailai Cheng, Kejia Chen, Lingyun Chen, Liding Zhang, Yue Zhang, Yao Ling, Mahdi Hamad, Zhenshan Bing, Fan Wu, Karan Sharma, Alois Knoll

机构 * School of Computation, Information and Technology, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院) Agile Robots SE(敏捷机器人公司) State Key Laboratory for Novel Software Technology and the School of Science and Technology, Nanjing University (Suzhou Campus)(南京大学软件新技术国家重点实验室及科学与技术学院(苏州校区)) Shanghai University(上海大学)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 TacUMI通过整合多种传感器,提供了一种多模态数据采集系统,用于提高接触密集型任务中多模态演示的分割和收集效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04299 2026-01-09 cs.LG q-bio.QM 78%

Transformer-Based Multi-Modal Temporal Embeddings for Explainable Metabolic Phenotyping in Type 1 Diabetes

基于Transformer的多模态时间嵌入用于1型糖尿病的可解释代谢表型分析

Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba, Sule Yildrim Yayilgan, Sarang Shaikh

专题命中 视频多模态 :multi-modal(title);multimodal(abstract)

AI总结 本研究提出基于Transformer的多模态时间嵌入框架,用于1型糖尿病的可解释代谢表型分析,识别出5种代谢亚组并揭示其与心血管风险的关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18785 2026-01-01 cs.LG 78%

ALF: Advertiser Large Foundation Model for Multi-Modal Advertiser Understanding

ALF:多模态广告商基础模型用于多模态广告商理解

Santosh Rajagopalan, Jonathan Vronsky, Songbai Yan, S. Alireza Golestaneh, Shubhra Chandra, Min Zhou

机构 * Google(谷歌)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 ALF通过多模态Transformer架构和多任务优化,实现了广告商行为理解的高精度和高召回率,显著提升了欺诈检测和政策识别的性能。

Comments KDD 2026 ADS Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20538 2025-12-24 cs.RO 78%

Uni-Mapper: Unified Mapping Framework for Multi-modal LiDARs in Complex and Dynamic Environments

Uni-Mapper:多模态激光雷达在复杂动态环境中的统一映射框架

Gilhwan Kang, Hogyun Kim, Byunghee Choi, Seokhwan Jeong, Young-Sik Shin, Younggun Cho

机构 * Hyundai Motor Company(现代汽车公司) Inha University(inha大学) Korea Institute of Machinery and Materials(韩国机械材料研究院)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 Uni-Mapper通过动态感知和多模态激光雷达融合技术,实现复杂动态环境下的统一地图构建与回环检测。

Comments 18 pages, 14 figures

Journal ref 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01882 2025-12-02 cs.LG 78%

New Spiking Architecture for Multi-Modal Decision-Making in Autonomous Vehicles

多模态决策中自动驾驶车辆的新脉冲架构

Aref Ghoreishee, Abhishek Mishra, Lifeng Zhou, John Walsh, Nagarajan Kandasamy

机构 * Electrical and Computer Engineering Department(电子与计算机工程系)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 本文提出了一种基于脉冲神经元的高效多模态决策架构,用于提升自动驾驶车辆的实时决策能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22362 2025-12-01 cs.LG 78%

Efficient-Husformer: Efficient Multimodal Transformer Hyperparameter Optimization for Stress and Cognitive Loads

Efficient-Husformer: 高效多模态Transformer超参数优化用于压力和认知负荷检测

Merey Orazaly, Fariza Temirkhanova, Jurn-Gyu Park

机构 * The School of Engineering and Digital Sciences, Nazarbayev University(工程与数字科学学院,纳扎尔拜耶夫大学)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 Efficient-Husformer通过超参数优化提升多模态生理信号分析的效率和准确性,实现压力和认知负荷检测的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16169 2025-11-21 eess.SP 78%

UT-OSANet: A Multimodal Deep Learning model for Evaluating and Classifying Obstructive Sleep Apnea

UT-OSANet: 一种用于评估和分类阻塞性睡眠呼吸暂停的多模态深度学习模型

Zijian Wang, Xiaoyu Bao, Chenhao Zhao, Jihui Zhang, Sizhi Ai, Yuanqing Li

专题命中 视频多模态 :multimodal(title);cross-modal(abstract)

AI总结 UT-OSANet是一种多模态深度学习模型,用于高精度评估和分类阻塞性睡眠呼吸暂停,通过多种输入模态和灵活训练策略实现事件层面的诊断。

Comments 12 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09773 2025-11-14 cs.LG eess.SP 78%

NeuroLingua: A Language-Inspired Hierarchical Framework for Multimodal Sleep Stage Classification Using EEG and EOG

Mahdi Samaee, Mehran Yazdi, Daniel Massicotte

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09342 2025-11-13 eess.SP 78%

A cross-modal pre-training framework with video data for improving performance and generalization of distributed acoustic sensing

Junyi Duan, Jiageng Chen, Zuyuan He

专题命中 视频多模态 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09137 2025-11-13 eess.SP 78%

xHAP: Cross-Modal Attention for Haptic Feedback Estimation in the Tactile Internet

Georgios Kokkinis, Alexandros Iosifidis, Qi Zhang

专题命中 视频多模态 :cross-modal(title,abstract)

Comments 12 pages, 13 figures, 3 tables, 2 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11321 2025-11-07 cs.RO 78%

HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data

Ruizhe Liu, Pei Zhou, Qian Luo, Li Sun, Jun Cen, Yibing Song, Yanchao Yang

机构 * HKU Musketeers Foundation Institute of Data Science(香港大学数据科学学院) The University of Hong Kong(香港大学) DAMO Academy(达摩院) Alibaba Group(阿里巴巴集团)

专题命中 视频多模态 :multi-modal(title);cross-modal(abstract)

Comments Accepted at 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01405 2025-11-05 eess.SP cs.ET 78%

MM-2FSK: Multimodal Frequency Shift Keying for Ultra-Efficient and Robust High-Resolution MIMO Radar Imaging

Vanessa Wirth, Johanna Bräunig, Martin Vossiek, Tim Weyrich, Marc Stamminger

专题命中 视频多模态 :multimodal(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00716 2025-11-04 cs.LG 78%

Enhancing Heavy Rain Nowcasting with Multimodal Data: Integrating Radar and Satellite Observations

Rama Kassoumeh, David Rügamer, Henning Oppel

机构 * Bochum Institute of Technology(波恩技术学院) LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Okeanos Smart Data Solutions GmbH(Okeanos智能数据解决方案有限公司)

专题命中 视频多模态 :multimodal(title,abstract)

Comments accepted to ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13522 2025-11-04 cs.LG cs.NE q-fin.CP 78%

Cross-Modal Temporal Fusion for Financial Market Forecasting

Yunhua Pei, John Cartlidge, Anandadeep Mandal, Daniel Gold, Enrique Marcilio, Riccardo Mazzon

机构 * School of Computer Science, University of Bristol, Bristol, UK(布里斯托大学计算机科学学院) School of Engineering Mathematics and Technology, University of Bristol, Bristol, UK(布里斯托大学工程数学与技术学院) School of Engineering Mathematics(工程数学学院) Technology, University of Bristol, Bristol, UK(技术学院) Business School, University of Birmingham, Birmingham, UK(伯明翰大学商学院) Stratiphy Limited, London, UK(Stratiphy公司)

专题命中 视频多模态 :cross-modal(title,abstract)

Comments 10 pages, 4 figures, manuscript accepted to PAIS at ECAI-2025 European Conference on Artificial Intelligence, October 25-30, 2025, Bologna, Italy

Journal ref Frontiers in Artificial Intelligence and Applications, vol. 413, ECAI 2025, pp. 5360 - 5367

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24238 2025-10-29 cond-mat.mtrl-sci 78%

Unlocking Dynamic Luminescent Mapping of pH with Sustainable Lignin-Derived Carbon Dots with Multimodal Readout Capacity

Maja Szymczak, Jan Hočevar, Jernej Iskra, Darja Lisjak, Jelena Papan Djaniš, Lukasz Marciniak, Karolina Elzbieciak-Piecka

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20952 2025-10-27 cs.LG 78%

LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting

Sungjun Cho, Changho Shin, Suenggwan Jo, Xinya Yan, Shourjo Aditya Chaudhuri, Frederic Sala

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01116 2025-10-14 cs.RO cs.LG cs.SY eess.SY 78%

Scalable Multi-modal Model Predictive Control via Duality-based Interaction Predictions

Hansung Kim, Siddharth H. Nair, Francesco Borrelli

机构 * Model Predictive Control Laboratory, UC Berkeley(模型预测控制实验室,伯克利大学)

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted at IEEE Intelligent Vehicles Symposium 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26636 2025-10-01 cs.LG 78%

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Ming Jin, Boyi Li, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学) UCL(伦敦大学学院) Virginia Tech(弗吉尼亚理工学院) Nvidia(英伟达公司)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21346 2025-09-29 cs.NE cs.LG q-bio.BM 78%

Spiking Neural Networks for Mental Workload Classification with a Multimodal Approach

Jiahui An, Sara Irina Fabrikant, Giacomo Indiveri, Elisa Donati

机构 * Institute of Neuroinformatics, University of Zurich(神经信息学研究所,苏黎世大学) ETH Zurich(苏黎世联邦理工学院) Digital Society Initiative, University of Zurich(数字社会倡议,苏黎世大学) Department of Geography(地理系)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏