arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4726 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4726 篇

2604.02093 2026-04-03 cs.CV 74%

GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding

GroundVTS:多模态大语言模型中的视频时间定位视觉标记采样

Rong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang, Kai Dai, Zhao Yang

机构 * Newcapec AI Research(新开普人工智能研究院) Fudan University(复旦大学) Tongji University(同济大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出GroundVTS,通过细粒度查询引导机制筛选视觉标记,提升视频时间定位的时空信息保留与时间连贯性,实验表明其在三个标准基准上优于现有方法。

Comments Published as a conference paper at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01712 2026-04-03 cs.LG cs.AI eess.SP physics.comp-ph 74%

Transformer self-attention encoder-decoder with multimodal deep learning for response time series forecasting and digital twin support in wind structural health monitoring

基于多模态深度学习的Transformer自注意力编码器-解码器:用于响应时间序列预测和风力结构健康监测的数字孪生支持

Feiyu Zhou, Marios Impraimakis

机构 * Department of Mechanical Engineering, University of Bath(巴斯大学机械工程系) Department of Civil Engineering, Zhejiang University(浙江大学土木工程系)

专题命中 视频多模态 :multimodal(title);分类 cs.AI

AI总结 本文提出一种基于Transformer的多模态深度学习方法,用于风力结构响应时间序列预测和数字孪生支持,通过捕捉系统时间特性提升结构健康监测的准确性与鲁棒性。

Comments 21 pages, 22 figures, 9 tables. This version corresponds to the published article in Computers & Structures. https://doi.org/10.1016/j.compstruc.2026.108216

Journal ref Computers and Structures 326 (2026) 108216

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16997 2026-03-31 cs.CV cs.LG cs.RO 74%

Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention

通过多模态注意力解决视频预测中的时空纠缠

Shreyam Gupta, P. Agrawal, Priyam Gupta

机构 * Indian Institute of Technology (BHU), Varanasi(印度理工学院(巴纳拉斯印度教大学),瓦拉纳西) University of Colorado, Boulder(科罗拉多大学博尔德分校) Erasmus+, Intelligent Field Robotic Systems (IFRoS), University of Girona(伊拉斯谟+,智能现场机器人系统(IFRoS),赫罗纳大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 本文提出MAUCell架构,结合GAN与层级处理策略及三种注意力机制,解决RNN在长时间序列中的局限,提升视频预测的时空一致性与实时性。

Comments 11 pages, 3 figures, 5 tables, and 3 Algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18990 2026-03-13 cs.CV 74%

IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition

IDSelect: 一种基于强化学习的成本感知选择代理用于基于视频的多模态人识别

Yuyang Ji, Yixuan Shen, Kien Nguyen, Lifeng Zhou, Feng Liu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 IDSelect 通过强化学习选择最优模型提升多模态人识别的效率与准确度,实验显示其在计算资源上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01284 2026-03-03 cs.CV 74%

FoSS: Modeling Long Range Dependencies and Multimodal Uncertainty in Trajectory Prediction via Fourier State Space Integration

FoSS:通过傅里叶状态空间整合建模长距离依赖性和多模态不确定性在轨迹预测中

Yizhou Huang, Gengze Jiang, Yihua Cheng, Kezhi Wang

机构 * Brunel University of London(伦敦布鲁内尔大学) University of Birmingham(伯明翰大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 FoSS通过结合频域和时域建模,有效提升轨迹预测的准确性和效率,减少计算与参数消耗,实现多模态不确定性建模。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22923 2026-02-27 cs.CV cs.RO 74%

WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents

WaterVideoQA: 以ASV为中心的感知与符合规则的推理 via 多模态智能体

Runwei Guan, Shaofeng Liang, Ningwei Ouyang, Weichen Fei, Shanliang Yao, Wei Dai, Chenhao Ge, Penglei Sun, Xiaohui Zhu, Tao Huang, Ryan Wen Liu, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)人工智能研究所) Hubei Key Laboratory of Inland Shipping Technology (Wuhan University of Technology)(湖北内河航运技术重点实验室(武汉理工大学)) School of Navigation, Wuhan University of Technology(武汉理工大学航海学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院) School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) School of Information Engineering, Yancheng Institute of Technology(盐城职业技术学院信息工程学院) School of Engineering, Stanford University(斯坦福大学工程学院) Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 WaterVideoQA通过多模态智能体系统,实现ASV在复杂水域环境中的感知与规则合规推理,提升自主航行的安全性和精确性。

Comments 11 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08683 2026-02-27 cs.CV 74%

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

OneVision-Encoder: 编码对齐的稀疏性作为多模态智能的基础原则

Feilong Tang, Xiang An, Yunyao Yan, Yin Xie, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Chunyuan Li, Shikun Feng, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng

机构 * Glint Lab(Glint实验室) AIM for Health Lab(健康人工智能实验室) MVP Lab(MVP实验室)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 OneVision-Encoder通过编码对齐的稀疏性原则,实现高效的多模态视觉理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17785 2026-02-23 cs.CV 74%

Multi-Modal Monocular Endoscopic Depth and Pose Estimation with Edge-Guided Self-Supervision

多模态单目内窥镜深度与姿态估计与边缘引导自监督学习

Xinwei Ju, Rema Daher, Danail Stoyanov, Sophia Bano, Francisco Vasconcelos

机构 * UCL Hawkes Institute, Department of Computer Science(伦敦大学学院UCL霍克斯研究所、计算机科学系)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 PRISM通过边缘引导自监督学习,结合解剖学和照明先验知识,提升内窥镜单目深度与姿态估计性能。

Comments 14 pages, 6 figures; early accepted by IPCAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00132 2026-02-03 cs.CV 74%

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

去除伪装,连接领域:通过测试时适应检测转移多模态仇恨视频

Jiao Li, Jian Lang, Xikai Tang, Wenzheng Shu, Ting Zhong, Qiang Gao, Yong Wang, Leiting Chen, Fan Zhou

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 SCANNER通过测试时适应框架,利用仇恨内容中稳定的内核连接源与目标领域,有效应对多模态仇恨视频检测中的语义漂移问题。

Comments Accepted by AAAI2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22661 2026-01-22 cs.IR cs.AI 74%

Next Point-of-interest (POI) Recommendation Model Based on Multi-modal Spatio-temporal Context Feature Embedding

基于多模态时空上下文特征嵌入的下一个兴趣点(POI)推荐模型

Lingyu Zhang, Pengfei Xu, Rui Ban, Zhenchao Zhang, Songtao Liu, Yan Wang, Yunhai Wang

机构 * Institute of Trustworthy Autonomous Systems, Southern University of Science and Technology(SUSTech)(可信自主系统研究院,南方科技大学) School of Information Sciences and Technology, Northwest University(信息科学与技术学院,西北大学) China Information Technology Designing & Consulting Institute Co., Ltd.(中国信息科技设计与咨询研究院有限公司) Renmin University of China(中国人民大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

AI总结 本文提出基于多模态时空上下文特征嵌入的POI推荐模型,通过双流时空注意力机制有效区分长期习惯与短期意图,提升个性化出行预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19850 2026-01-09 cs.CV 74%

FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs

FALCONEye: 在一小时内视频中寻找答案并定位内容的多模态大语言模型

Carlos Plou, Cesar Borja, Ruben Martinez-Cantin, Ana C. Murillo

机构 * DIIS-I3A, University of Zaragoza(DIIS-I3A,西班牙阿利坎特大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 FALCONEye通过多模态大语言模型在小时长视频中高效定位内容并回答问题,超越现有模型性能并显著降低推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23361 2026-01-01 cs.CV 74%

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

OmniVCus: 基于多模态控制条件的前馈主体驱动视频定制

Yuanhao Cai, He Zhang, Xi Chen, Jinbo Xing, Yiwei Hu, Yuqian Zhou, Kai Zhang, Zhifei Zhang, Soo Ye Kim, Tianyu Wang, Yulun Zhang, Xiaokang Yang, Zhe Lin, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) Adobe Research(Adobe研究) The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 OmniVCus通过多模态控制条件和改进的嵌入机制实现高效的多主体视频定制。

Comments NeurIPS 2025; A data construction pipeline and a diffusion Transformer framework for controllable subject-driven video customization

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21693 2025-12-01 cs.MM 74%

Designing a Multimodal Viewer for Piano Performance Analysis -- a Pedagogy-First Approach

为钢琴表演分析设计多模态查看器——一种以教学为导向的方法

Joonhyung Bae, Hyeyoon Cho, Kirak Kim, Dawon Park, Taegyun Kwon, Yoon-Seok Choi, Hyeon Hur, Shigeru Kai, Yohei Wada, Satoshi Obata, Akira Maezawa, Jaebum Park, Jonghwa Park, Juhan Nam

专题命中 视频多模态 :multimodal(title);分类 cs.MM

AI总结 本研究提出了一种以教学为导向的多模态查看器,通过整合视频、动作捕捉和乐谱,为钢琴教学提供具体的视觉反馈,以提高教学指导的清晰度和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16524 2025-11-21 cs.CV 74%

BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization

BoxingVI:一种多模态基准,用于拳击动作识别与定位

Rahul Kumar, Vipul Baghel, Sudhanshu Singh, Bikash Kumar Badatya, Shivam Yadav, Babji Srinivasan, Ravi Hegde

机构 * Indian Institute of Technology Gandhinagar(印度理工学院甘地纳加尔) Indian Institute of Technology Madras(印度理工学院马德拉斯) Dr. A. P. J. Abdul Kalam Technical University(阿卜杜勒·卡拉姆技术大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 BoxingVI提供了一个多模态数据集,用于拳击动作识别与定位,旨在促进低资源环境下的实时视觉动作识别研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15342 2025-11-20 cs.HC cs.AI 74%

Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions

Shan Shan

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01562 2025-11-17 cs.CV 74%

Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion

Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01481 2025-10-28 cs.CV cs.LG 74%

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du, Fuxiao Liu, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06509 2025-10-13 cs.CV 74%

From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding

Shih-Yao Lin, Sibendu Paul, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07963 2025-10-06 cs.RO cs.AI cs.SY eess.SY 74%

Hierarchical Contact-Rich Trajectory Optimization for Multi-Modal Manipulation using Tight Convex Relaxations

Yuki Shirai, Arvind Raghunathan, Devesh K. Jha

机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 2025 IEEE International Conference on Robotics and Automation (2025 ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00139 2025-10-01 cs.CV 74%

SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection for SLAM

Yannick Burkhardt, Simon Schaefer, Stefan Leutenegger

机构 * Technical University of Munich(慕尼黑技术大学) ETH Zürich(苏黎世联邦理工学院) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13838 2025-09-30 eess.SP cs.CV cs.IT eess.IV math.IT 74%

Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model

Hang Yin, Li Qiao, Yu Ma, Shuo Sun, Kan Li, Zhen Gao, Dusit Niyato

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments IEEE Transactions on Vehicular Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12341 2025-08-14 cs.MM 74%

Multimodal LLM-based Query Paraphrasing for Video Search

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Sheng-Hua Zhong, Xiong-Yong Wei, Qing Li

专题命中 视频多模态 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23524 2025-06-06 cs.CV 74%

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

Rui Xia, Dan Jiang, Quan Zhang, Ke Zhang, Chun Yuan

专题命中 视频多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02252 2025-06-03 cs.CV 74%

MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning

Jianghui Wang, Yuxuan Wang, Dongyan Zhao, Zilong Zheng

机构 * Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院) Wangxuan Institute of Computer Technology(王轩计算机技术研究所)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05541 2025-04-10 cs.CV 74%

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Yunlong Tang, Jing Bi, Chao Huang, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Yizhi Song, Pinxin Liu, Mingqian Feng, Junjia Guo, Zhuo Liu, Luchuan Song, Ali Vosoughi, Jinxi He, Liu He, Zeliang Zhang, Jiebo Luo, Chenliang Xu

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05673 2025-04-09 cs.CV 74%

VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs

Dongjun Qian, Kai Su, Yiming Tan, Qishuai Diao, Xian Wu, Chang Liu, Bingyue Peng, Zehuan Yuan

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03112 2025-03-06 cs.SI cs.AI cs.NE 74%

A Multimodal Framework for Topic Propagation Classification in Social Networks

Yuchuan Jiang, Chaolong Jia, Yunyi Qin, Wei Cai, Yongsen Qian

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17805 2024-12-24 cs.CV 74%

Large Motion Video Autoencoding with Cross-modal Video VAE

Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Project Website: https://yzxing87.github.io/vae/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00517 2024-12-03 cs.AI cs.ET cs.RO 74%

LAMBDA: Covering the Multimodal Critical Scenarios for Automated Driving Systems by Search Space Quantization

Xinzheng Wu, Junyi Chen, Xingyu Xing, Jian Sun, Ye Tian, Lihao Liu, Yong Shen

专题命中 视频多模态 :multimodal(title);分类 cs.AI

Comments 17pages, 21figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19825 2024-10-29 cs.CV 74%

Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis

Elia Fantini

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 150 pages, 60 figures

详情

展开后加载摘要…

URL PDF HTML 收藏