arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2408.07867 2025-12-02 cs.CV 79%

Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models

持续感知至关重要:多模态模型中时间整合失败的诊断

Zeyu Wang, Zhenzhen Weng, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CP-Bench,通过简单任务揭示多模态模型在持续感知中的时间整合缺陷,指出现有模型无法有效跨时间积累证据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22774 2025-12-01 cs.CV 79%

Alzheimer's Disease Prediction Using EffNetViTLoRA and BiLSTM with Multimodal Longitudinal MRI Data

使用EffNetViTLoRA和BiLSTM的多模态纵向MRI数据进行阿尔茨海默病预测

Mahdieh Behjat Khatooni, Mohsen Soryani

机构 * School of Computer Engineering, Iran University of Science and Technology(伊朗科学技术大学计算机工程学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出结合EffNetViTLoRA和BiLSTM的多模态模型,利用纵向MRI数据实现高精度的阿尔茨海默病预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18773 2025-12-01 cs.CV 79%

Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains

Spacewalk-18:多模态和长形式程序视频理解的基准测试,用于新领域

Zitian Tang, Rohan Myer Krishnan, Zhiqiu Yu, Chen Sun

机构 * Brown University(布朗大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 Spacewalk-18是一个用于多模态和长形式程序视频理解的新领域基准测试,通过步骤识别和视频问答任务评估模型在新领域和长时序上下文中的泛化能力。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05034 2025-11-26 cs.CV 79%

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

视频大模型训练后:深入探讨大型多模态模型在视频推理中的应用

Yolo Y. Tang, Jing Bi, Pinxin Liu, Zhenyu Pan, Zhangyun Tan, Qianxiang Shen, Jiani Liu, Hang Hua, Junjia Guo, Yunzhong Xiao, Chao Huang, Zhiyuan Wang, Susan Liang, Xinyi Liu, Yizhi Song, Junhua Huang, Jia-Xing Zhong, Bozheng Li, Daiqing Qi, Ziyun Zeng, Ali Vosoughi, Luchuan Song, Zeliang Zhang, Daiki Shimada, Han Liu, Jiebo Luo, Chenliang Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文深入探讨了Video-LMMs训练后处理方法,分析了监督微调、强化学习和测试时扩展等关键技术,总结了设计原则与挑战,为视频推理研究提供了统一框架。

Comments Version v1.1

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18823 2025-11-25 cs.CV 79%

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

VideoPerceiver: 提升视频多模态大语言模型的细粒度时序感知能力

Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao, Chen Ye, Yang Cai, Jian Gao, Danfeng Yan

机构 * State Key Laboratory of Networking and Switching Technology, BUPT(网络与交换技术国家重点实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoPerceiver通过两阶段训练框架提升视频多模态大语言模型对细粒度动作和短暂事件的感知能力,显著优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18104 2025-11-25 cs.CV 79%

Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning

通过统一多模态伪造学习巩固扩散生成视频检测

Xiaohong Liu, Xiufeng Song, Huayu Zheng, Lei Bai, Xiaoming Liu, Guangtao Zhai

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学信息科学与电子工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Computer Science and Engineering, Michigan State University(密歇根州立大学计算机科学与工程系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MM-Det++算法,通过统一多模态学习检测扩散生成视频,结合时空分支和多模态分支,提升视频伪造检测的准确性和泛化能力。

Comments Code and dataset are available at https://github.com/SparkleXFantasy/MM-Det-Plus

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17945 2025-11-25 cs.CV 79%

Test-Time Temporal Sampling for Efficient MLLM Video Understanding

测试时时间采样用于高效多模态大语言模型视频理解

Kaibin Wang, Mingbao Lin

机构 * SenseTime, China(深睿时代,中国) Rakuten, Singapore(拉结恩,新加坡)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 T3S通过测试时时间采样提高多模态大语言模型处理长视频的效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18883 2025-11-24 cs.CV 79%

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

通用视频时间定位与生成多模态大语言模型

Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学SAI实验室) ByteDance Seed(字节跳动种子)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出UniTime模型,利用生成多模态大语言模型实现通用视频时间定位,有效处理多类型视频并提升VideoQA任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16227 2025-11-21 cs.CV 79%

SwiTrack: Tri-State Switch for Cross-Modal Object Tracking

SwiTrack:跨模态目标跟踪的三态开关

Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Mathematics, Southeast University(东南大学数学学院) Purple Mountain Laboratories(紫金山实验室)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 SwiTrack通过三态开关框架提升跨模态目标跟踪的鲁棒性和精度,实现7.2%和4.3%的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15722 2025-11-21 cs.AI 79%

Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods

多模态大语言模型中的空间推理:任务、基准和方法的调查

Weichen Liu, Qiyao Xue, Haoming Wang, Xiangyu Yin, Boyuan Yang, Wei Gao

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文调查多模态大语言模型中的空间推理问题,从认知角度分类任务和基准,分析评估方法与改进策略,揭示模型与人类推理间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14770 2025-11-20 cs.IR cs.AI 79%

ExplainRec: Towards Explainable Multi-Modal Zero-Shot Recommendation with Preference Attribution and Large Language Models

Bo Ma, LuYao Liu, ZeHua Hu, Simon Lau

机构 * Department of Software \& Microelectronics Peking University Beijing, China Department of Software \& Microelectronics Peking University Beijing, China hangli\ Department of Software \& Microelectronics Peking University Beijing, China zehua\ Department of Software \& Microelectronics Peking University Beijing, China xiaofan\ Economic Law School China University of Political Science School of Computer Science Peking University Beijing, China

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10210 2025-11-20 cs.CV 79%

MK-SGN: A Spiking Graph Convolutional Network with Multimodal Fusion and Knowledge Distillation for Skeleton-based Action Recognition

Naichuan Zheng, Hailun Xia, Zeyu Liang, Yuchen Du

机构 * Beijing Laboratory of Advanced Information Networks, Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications(北京先进信息网络实验室、网络系统架构与收敛重点实验室、信息与通信工程学院、北京邮电大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14057 2025-11-19 cs.LG cs.AI 79%

A Machine Learning-Based Multimodal Framework for Wearable Sensor-Based Archery Action Recognition and Stress Estimation

Xianghe Liu, Jiajia Liu, Chuxian Xu, Minghan Wang, Hongbo Peng, Tao Sun, Jiaqi Xu

机构 * Beijing PsychTech Technology Co., Ltd.(北京心理科技技术有限公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13655 2025-11-18 cs.CV cs.LG 79%

OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

Henry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng, Joseph Redmon, Hadrien Sablon, Ryan Park, Jacob Morrison, Alexandra Buraczynski, Karen Farley, Joshua Hansen, Andrew Howe, Patrick Alan Johnson, Mark Otterlee, Ted Schmitt, Hunter Pitelka, Stephen Daspit, Rachel Ratner, Christopher Wilhelm, Sebastian Wood, Mike Jacobi, Hannah Kerner, Evan Shelhamer, Ali Farhadi, Ranjay Krishna, Patrick Beukema

机构 * Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学) Arizona State University(亚利桑那州立大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12196 2025-11-18 cs.CV cs.HC 79%

Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System

Aditi Bhalla, Christian Hellert, Enkelejda Kasneci

机构 * School of Social Sciences and Technology(社会科学与技术学院)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11002 2025-11-17 cs.CV 79%

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 15 pages, 12 figures. Accepted as an Oral presentation at AAAI 2026. For code and dataset, see https://zane-zyqiu.github.io/EmoVid

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10134 2025-11-14 cs.CV 79%

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

Mingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li, Qi Zeng, Yifan Zhang, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23990 2025-11-13 cs.AI 79%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02182 2025-11-05 cs.CV 79%

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

Jinhwan Seo, Yoonki Cho, Junhyug Noh, Sung-eui Yoon

机构 * KAIST(韩国科学技术院) Ewha Womans University(成均馆大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 1st place winner of Grounded Videoqa track at the ICCV2025 Perception Test

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17897 2025-10-30 q-bio.NC cs.CV cs.LG 79%

Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)

Semih Eren, Deniz Kucukahmetler, Nico Scherf

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) TU Dresden(德累斯顿技术大学) School for Embedded and Composite AI (SECAI)(嵌入式与复合人工智能学院) Center for Scalable Data Analytics & AI (ScaDS.AI)(可扩展数据与人工智能中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 2 figures, 1 table. Invited report, CCN 2025 Algonauts Project session (3rd-place team). Code: https://github.com/erensemih/Algonauts2025_ModalityRNN v3: Added equal contribution footnote to author list. Corrected reference list

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25332 2025-10-30 cs.CV 79%

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu

机构 * Henan Institute of Advanced Technology, Zhengzhou University(河南高级技术研究所,郑州大学) Institute of Automation, CAS(自动化研究所,中国科学院) UCAS(中国科学院大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23934 2025-10-29 cs.CY cs.AI cs.ET 79%

MFiSP: A Multimodal Fire Spread Prediction Framework

Alec Sathiyamoorthy, Wenhao Zhou, Xiangmin Zhou, Xiaodong Li, Iqbal Gondal

机构 * School of Computing Technologies, RMIT University(计算技术学院,皇家墨尔本理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03727 2025-10-29 eess.AS cs.LG 79%

Detecting Neurocognitive Disorders through Analyses of Topic Evolution and Cross-modal Consistency in Visual-Stimulated Narratives

Jinchao Li, Yuejiao Wang, Junan Li, Jiawen Kang, Bo Zheng, Ka Ho Wong, Brian Mak, Helene H. Fung, Jean Woo, Man-Wai Mak, Timothy Kwok, Vincent Mok, Xianmin Gong, Xixin Wu, Xunying Liu, Patrick C. M. Wong, Helen Meng

机构 * The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 eess.AS

Comments 16 pages, 5 figures, accepted by "IEEE Journal of Selected Topics in Signal Processing"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06499 2025-10-28 q-bio.NC cs.AI 79%

The ISLab Solution to the Algonauts Challenge 2025: A Multimodal Deep Learning Approach to Brain Response Prediction

Andrea Corsico, Giorgia Rigamonti, Simone Zini, Luigi Celona, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication, University of Milano-Bicocca(信息学、系统与通信系,米兰-比科卡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21406 2025-10-27 cs.CV 79%

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

Yue Feng, Jinwei Hu, Qijia Lu, Jiawei Niu, Li Tan, Shuo Yuan, Ziyi Yan, Yizhen Jia, Qingzhi He, Shiping Ge, Ethan Q. Chen, Wentong Li, Limin Wang, Jie Qin

机构 * MoE Key Laboratory of Brain-Machine Intelligence Technology, College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(脑机智能技术MoE实验室,人工智能学院,南京航空航天大学) Nanjing University(南京大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22854 2025-10-27 cs.CV 79%

CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting

Kornel Howil, Joanna Waczyńska, Piotr Borycki, Tadeusz Dziarmaga, Marcin Mazur, Przemysław Spurek

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03340 2025-10-27 cs.CV 79%

Seeing the Arrow of Time in Large Multimodal Models

Zihui Xue, Mi Luo, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025, Project website: https://vision.cs.utexas.edu/projects/SeeAoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18411 2025-10-22 cs.CL cs.LG 79%

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

Yue Jiang, Jichu Li, Yang Liu, Dingkang Yang, Feng Zhou, Quyu Kong

机构 * Fudan University(复旦大学) Center for Applied Statistics and School of Statistics, Renmin University of China(应用统计中心和中国人民大学统计学院) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高级创新中心)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07447 2025-10-15 cs.CV 79%

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, School of CSE, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学计算机科学与工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11110 2025-10-14 cs.LG cs.AI 79%

PhysioME: A Robust Multimodal Self-Supervised Framework for Physiological Signals with Missing Modalities

Cheol-Hui Lee, Hwa-Yeon Lee, Min-Kyung Jung, Dong-Joo Kim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏