arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4728 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4728 篇

2512.18318 2025-12-23 cs.MM cs.AI cs.CV cs.DC cs.NI 67%

Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems

异步流水线并行用于实时多语言唇同步视频通信系统

Eren Caglar, Amirkia Rafiei Oskooei, Mehmet Kutanoglu, Mustafa Keles, Mehmet S. Aktas

机构 * Department of Data Science Big Data Yildiz Technical University Istanbul, Turkey Department of Computer Engineering Yildiz Technical University Istanbul, Turkey R\&D Center Aktif Investment Bank Inc. Istanbul, Turkey 3.5cm R\&D Center 3.5cm Aktif Investment Bank Inc. 3.5cm Istanbul, Turkey 3.5cm 0.25cm Department of Computer Engineering 0.25cm Yildiz Technical University 0.25cm Istanbul, Turkey 0.25cm

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出一种异步流水线并行的Transformer框架,用于实时多语言唇同步,通过模块化设计和优化技术提升处理速度与资源利用率。

Comments Accepted to IEEE Big Data 2025, AIDE4IoT Workshop. Copyright \c{opyright} 2025 IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08051 2025-12-12 cs.LG 67%

ST-GraphNet: A Spatio-Temporal Graph Neural Network for Understanding and Predicting Automated Vehicle Crash Severity

ST-GraphNet:一种用于理解和预测自动驾驶车辆碰撞严重程度的时空图神经网络

Mahmuda Sultana Mimi, Md Monzurul Islam, Anannya Ghosh Tusti, Shriyank Somvanshi, Subasish Das

机构 * Texas State University(德克萨斯州立大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract)

AI总结 ST-GraphNet通过时空图神经网络整合多模态数据,以高准确率预测自动驾驶车辆碰撞严重程度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09841 2025-12-11 cs.CL cs.CV cs.MM 67%

ChronusOmni: Improving Time Awareness of Omni Large Language Models

ChronusOmni: 提升 Omni 大语言模型的时间感知能力

Yijing Chen, Yihan Wu, Kaisi Guan, Yuchen Ren, Yuyue Wang, Ruihua Song, Liyun Ru

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学光荣人工智能学院) Baichuan Inc.(百川科技公司)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 ChronusOmni 通过增强时间感知能力,提升音频视觉时间定位任务的性能,实现跨模态统一建模和细粒度推理。

Comments Code available at https://github.com/YJCX330/Chronus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23400 2025-12-05 astro-ph.IM astro-ph.SR 67%

Solar flare forecasting with foundational transformer models across image, video, and time-series modalities

基于基础Transformer模型的太阳耀斑预测:跨图像、视频和时间序列模态

S. Riggi, P. Romano, A. Pilzer, U. Becciani

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 本文通过比较三种基础Transformer模型在太阳耀斑预测中的表现,发现Moirai2在时间序列预测中表现最佳,展示了预训练模型在多模态空间天气预测中的潜力。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16595 2025-11-27 cs.CV cs.AI cs.CL 67%

TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

TimeViper: 一种用于高效长视频理解的混合Mamba-Transformer视觉语言模型

Boshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, Qin Jin

机构 * AIM3 Lab, Renmin University of China(中国人民大学人工智能实验室) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 TimeViper是一种混合Mamba-Transformer模型,通过TransV模块实现高效长视频理解,提升多模态处理能力。

Comments Project page: https://xuboshen.github.io/TimeViper; Code: https://github.com/xiaomi-research/timeviper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19475 2025-11-26 cs.CV cs.AI cs.MM 67%

Tracking and Segmenting Anything in Any Modality

任何模态下任何事物的跟踪与分割

Tianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong Han

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 SATA提出了一种通用框架,通过解耦的专家混合机制和任务感知多目标跟踪流程,统一了多种跟踪和分割任务,提升了模型的泛化能力。

Comments Accpetd by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19002 2025-11-18 cs.CV cs.AI cs.CL cs.LG 67%

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Ziqi Yin, Wentao Hu, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara

机构 * Waseda University(早稻田大学) CyberAgent, Inc.(CyberAgent公司) AI Shift, Inc.(AI Shift公司) Nara Institute of Science and Technology(奈良研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10008 2025-11-18 cs.MM cs.AI cs.CV 67%

Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Updated with the ICIDS 2025 camera-ready version. This revision includes the final title, updated abstract, improved explanations of the narrative coherence framework, and minor editorial changes. Figures and examples have been refined for clarity. No new experiments were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16781 2025-11-14 cs.CV cs.AI cs.CL 67%

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21786 2025-10-28 cs.CV cs.AI cs.MM 67%

EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction

Qile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang, Chao Tong

机构 * Beihang University(北京航空航天大学) University of Science and Technology Beijing(北京科技大学) School of Computer Science and Engineering(计算机科学与工程学院) State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 15 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12299 2025-10-15 cs.IR 67%

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Zhi Li, Yanan Wang, Hao Niu, Julio Vizcarra, Masato Taya

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25745 2025-10-01 cs.CV cs.CL cs.MM 67%

FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos

Siddhant Sukhani, Yash Bhardwaj, Riya Bhadani, Veer Kejriwal, Michael Galarnyk, Sudheer Chava

机构 * Stanford University(斯坦福大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments ICCV Short Video Understanding Workshop Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07966 2025-10-01 cs.CV cs.AI cs.CL 67%

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu, Hongxu Yin, Yao Lu, Song Han

机构 * NVIDIA MIT(麻省理工学院) HKU(香港大学) UC Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS 2025. Code at https://github.com/NVlabs/Long-RL and model at https://huggingface.co/Efficient-Large-Model/LongVILA-R1-7B

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11915 2025-09-18 cs.SD cs.CV cs.LG cs.MM eess.AS 67%

Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

机构 * Graduate School of AI, KAIST(韩国国立庆熙大学人工智能研究生院) Graduate School of CT, KAIST(韩国国立庆熙大学CT研究生院)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12231 2025-09-17 cs.DC 67%

Research on fault diagnosis and root cause analysis based on full stack observability

Jian Hou

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04957 2025-09-08 cs.CV cs.MM cs.SD eess.AS 67%

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

Gehui Chen, Guan'an Wang, Xiaowen Huang, Jitao Sang

机构 * School of Computer Science Technology, Beijing Jiaotong University Beijing China Beijing Key Laboratory of Traffic Data Mining Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education Beijing China Technology, Beijing Jiaotong University Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 67%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21586 2025-08-08 cs.CL cs.AI cs.CV 67%

Can Vision Language Models Understand Mimed Actions?

Hyundong Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) Institute for Creative Technologies(创意技术研究所) Department of Computer Science(计算机科学系) University of Southern California(南加州大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Aristotle University of Thessaloniki(希腊雅典纳大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09068 2025-07-24 cs.CV cs.AI cs.IR cs.LG cs.MM 67%

Infinite Video Understanding

Dell Zhang, Xiangyu Chen, Jixiang Luo, Mengxi Jia, Changzhi Sun, Ruilong Ren, Jingren Liu, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) Peking University(北京大学) Tianjin University(天津大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06260 2025-07-10 cs.CR cs.CY 67%

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework

Satyapriya Krishna, Ninareh Mehrabi, Abhinav Mohanty, Matteo Memelli, Vincent Ponzo, Payal Motwani, Rahul Gupta

专题命中 视频多模态 :multimodal(abstract);multimodal foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03365 2025-07-08 cs.RO 67%

Label-Free Long-Horizon 3D UAV Trajectory Prediction via Motion-Aligned RGB and Event Cues

Hanfang Liang, Shenghai Yuan, Fen Liu, Yizhuo Yang, Bing Wang, Zhuyu Huang, Chenyang Shi, Jing Jin

机构 * Jianghan University(江汉大学) Nanyang Technological University(南洋理工大学) Beihang University (BUAA)(北京航空航天大学)

专题命中 视频多模态 :multimodal(abstract);audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21862 2025-06-30 cs.CV cs.AI cs.HC cs.MM 67%

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Boyuan Sun, Jiaxing Zhao, Xihan Wei, Qibin Hou

机构 * VCIP, School of Computer Science, Nankai University(南开大学计算机学院) Tongyi Lab, Alibaba Group(阿里巴巴集团 Tongyi 实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 21 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08260 2025-06-25 cs.SE 67%

FixDrive: Automatically Repairing Autonomous Vehicle Driving Behaviour for $0.08 per Violation

Yang Sun, Christopher M. Poskitt, Kun Wang, Jun Sun

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract)

Comments Accepted by the 47th IEEE/ACM International Conference on Software Engineering (ICSE 2025)

Journal ref Proc. ICSE'25, pages 1921-1933. IEEE, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16267 2025-06-11 cs.CV cs.AI cs.CL cs.LG 67%

xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Michael S. Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Jongwoo Park, Kanchana Ranasinghe, Silvio Savarese, Ran Xu, Caiming Xiong, Juan Carlos Niebles

机构 * Salesforce AI Research(Salesforce AI研究院) Stony Brook University(石溪大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12559 2025-06-10 cs.CV cs.CL cs.MM 67%

AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu, Li Cao, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Huawei Technologies Co., Ltd.(华为技术有限公司) Shandong University(山东大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22019 2025-06-04 cs.CL cs.AI cs.CV 67%

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Qiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen, Lin Chen, Shihang Wang, Pengjun Xie, Fei Huang, Feng Zhao

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23268 2025-05-30 cs.CV cs.AI cs.MM 67%

Unsupervised Transcript-assisted Video Summarization and Highlight Detection

Spyros Barbakos, Charalampos Antoniadis, Gerasimos Potamianos, Gianluca Setti

机构 * University of Thessaly(塞萨洛尼基大学) King Abdullah University of Science and Technology(卡布斯大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05874 2025-05-30 cs.CV cs.AI cs.CL cs.IR cs.LG 67%

VideoRAG: Retrieval-Augmented Generation over Video Corpus

Soyeong Jeong, Kangsan Kim, Jinheon Baek, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22457 2025-05-29 cs.CV cs.AI cs.CL 67%

Fostering Video Reasoning via Next-Event Prediction

Haonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du, Kenji Kawaguchi, Ye Wang, Tianyu Pang

机构 * National University of Singapore(新加坡国立大学) Sea AI Lab(Sea AI实验室)

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20987 2025-05-28 cs.IR 67%

LifeIR at the NTCIR-18 Lifelog-6 Task

Jiahan Chen, Da Li, Keping Bi

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏