arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4699 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4699 篇

2603.16495 2026-03-18 cs.AI 87%

ExpressMind: A Multimodal Pretrained Large Language Model for Expressway Operation

ExpressMind:一种用于高速公路运营的多模态预训练大语言模型

Zihe Wang, Yihuan Wang, Haiyang Yu. Zhiyong Cui, Xiaojian Liao, Chengcheng Wang, Yonglin Tian, Yongxin Tong

机构 * Beihang University(北航) Shandong Hi-speed Group Co., Ltd(山东高速集团有限公司) Institute of automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);cross-modal(abstract)

AI总结 本文提出ExpressMind,一种针对高速公路运营的多模态预训练大语言模型,通过构建全栈数据集和双层预训练方法,提升事件检测、安全响应生成和复杂交通分析能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02963 2026-07-07 cs.CV cs.AI cs.MM 新提交 87%

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

用于全模态密集视频字幕的并行自回归解码

Wenzheng Zeng, Siyi Jiao, Chen Gao, Hwee Tou Ng, Mike Zheng Shou

机构 * National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 研究密集视频字幕生成,提出并行自回归框架,利用事件间弱局部依赖重组因果依赖图,引入全局规划和事件分解并行解码机制,提升效率与字幕性能。

Comments ECCV 2026. Project website: https://github.com/showlab/PadCaptioner

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07541 2026-06-09 cs.HC cs.AI cs.CV cs.CY cs.MM 新提交 87%

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

多模态大语言模型作为视频研究中的合成参与者:一项评估

Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI、cs.MM

AI总结 本研究评估多模态大语言模型在视频感知任务中模拟人类主观评分的表现,发现模型存在偏差且与人类一致性有限。

Comments Accepted to SocialLLM @ ICWSM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19178 2025-09-30 cs.CV cs.AI cs.CL cs.IR cs.LG 87%

Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval

Yang Du, Yuqi Liu, Qin Jin

机构 * Renmin University of China(中国人民大学)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACMMM 2024 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19772 2025-03-21 cs.CV cs.CL cs.LG cs.MM 87%

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Tiantian Geng, Jinrui Zhang, Qingni Wang, Teng Wang, Jinming Duan, Feng Zheng

专题命中 视频多模态 :omni-modal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06020 2025-02-11 cs.CV cs.MM cs.SD eess.AS 87%

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, Jiang Gui

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14200 2024-06-03 cs.CV cs.AI 87%

Awesome Multi-modal Object Tracking

Chunhui Zhang, Li Liu, Hao Wen, Xi Zhou, Yanfeng Wang

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);vision-language-audio(abstract);分类 cs.CV、cs.AI

Comments A continuously updated project to track the latest progress in multi-modal object tracking

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13024 2026-06-12 cs.LG cs.AI 新提交 86%

CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

CausalMoE:基于模式路由异构专家的十亿规模多模态基础模型用于格兰杰因果发现

Bo Liu, Di Dai, Jingwei Liu, Jiarui Jin, Xiaocheng Fang, Guangkun Nie, Hongyan Li, Shenda Hong

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) National Institute of Health Data Science, and Institute for Artificial Intelligence, Peking University(北京大学健康医疗大数据国家研究院、人工智能研究院)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

AI总结 提出CausalMoE,一种十亿规模多模态格兰杰因果基础模型,通过模式路由混合异构专家解耦动态机制,结合因果自注意力与LLM/VLM先验,实现稀疏因果图恢复,在监督和少样本场景中达到最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07642 2026-05-11 cs.CV 86%

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

EggHand:一种用于第一人称手姿态预测的多模态基础模型

Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park

机构 * DGIST, Republic of Korea(韩国忠南科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

AI总结 本文提出EggHand模型,通过结合视觉-语言-动作模型的动作解码器和第一人称视频-文本编码器,实现对第一人称视频中手姿态序列的预测,提升了在剧烈视角变化下的鲁棒性和可控性。

Comments CVPR Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07497 2025-06-23 cs.CV 86%

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Xiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu, Lijun Zhou, Gangwei Xu, Shaoqing Xu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17088 2024-06-24 cs.CV 86%

Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies

Mulin Tian, Mahyar Khayatkhoei, Joe Mathai, Wael AbdAlmageed

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(title);分类 cs.CV

Comments 11 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02927 2026-08-11 cs.CV cs.AI 86%

PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding

PrismVAU: 用于多模态视频异常理解的提示优化推理系统

Iñaki Erregue, Kamal Nasrollahi, Sergio Escalera

机构 * Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心) Aalborg University(奥胡斯大学) Milestone Systems(Milestone系统)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 PrismVAU通过轻量级系统和自动提示工程实现高效的多模态视频异常理解,无需复杂标注和外部模块。

Comments This paper has been accepted to the 6th Workshop on Real-World Surveillance: Applications and Challenges (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05833 2026-06-19 cs.CV cs.AI 版本更新 86%

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

从视频中学习几何表示以实现空间智能多模态大语言模型

Haibo Wang, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出GeoVR框架,通过从2D视频序列中蒸馏3D几何知识(包括相机姿态、深度图、尺度因子和多尺度3D特征),重塑多模态大语言模型的内部表示以赋予其空间智能,在空间推理基准上达到最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17422 2026-04-21 cs.CV cs.MM 86%

Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

聚焦何处:用于长视频理解的查询调节多模态关键帧选择

Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong

机构 * Department of CSE, HKUST(香港科技大学计算机科学与工程系)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.MM

AI总结 本文提出Q-Gate框架,通过动态模态路由解决长视频理解中关键帧选择问题,有效抑制模态噪声,提升多模态大语言模型的推理能力。

Comments 9 pages, 7 figures, 9 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28696 2026-03-31 cs.CV cs.AI 86%

AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding

AdaptToken: 基于熵的自适应令牌选择用于多模态大语言模型长视频理解

Haozhe Qi, Kevin Qu, Mahdi Rad, Rui Wang, Alexander Mathis, Marc Pollefeys

机构 * Microsoft Spatial AI Lab(微软空间人工智能实验室) EPFL(瑞士联邦理工学院洛桑) ETH Zurich(苏黎世联邦理工学院)

专题命中 视频多模态 :MLLM(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 AdaptToken通过利用模型自不确定性实现全局控制信号,提升多模态大语言模型对长视频的理解能力,通过熵信号分配令牌预算并支持提前停止,提升准确率并减少推理时间。

Comments Project page: https://haozheqi.github.io/adapt-token

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03722 2025-08-07 cs.CV cs.AI 86%

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi, Feng Chen, Xinming Wang, Jun Xie

机构 * Lenovo Research(联想研究院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Automation, CAS(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学理工学院) School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07138 2025-03-03 cs.CV cs.CL cs.LG 86%

Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content

Girish A. Koushik, Diptesh Kanojia, Helen Treharne

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to the MM4SG Workshop at the WebConf 2025

Journal ref Companion Proceedings of the ACM Web Conference 2025 (WWW Companion '25), April 28-May 2, 2025, Sydney, NSW, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16552 2024-07-25 cs.CV cs.MM 86%

MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues

Liyun Zhang

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12059 2021-02-02 cs.CV cs.CL 86%

VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs

Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, Lorenzo Torresani

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10212 2025-11-14 cs.CV 86%

Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization

Ashutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong Chng

机构 * College of Computing and Data Science(计算与数据科学学院) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Under Review, Multimodal Deepfake detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18314 2026-02-13 q-bio.QM cs.LG q-bio.NC 86%

BrainSymphony: A parameter-efficient multimodal foundation model for brain dynamics with limited data

BrainSymphony: 一种参数高效、多模态的基础模型,用于在有限数据下的脑动态

Moein Khajehnejad, Forough Habibollahi, Devon Stoliker, Adeel Razi

机构 * Turner Institute for Brain and Mental Health(大脑与心理健康Turner研究所) School of Psychological Sciences, Monash University(墨尔本大学心理学科学学院) Cortical Labs(皮层实验室) CIFAR Azrieli Global Scholars Program(CIFAR阿兹里埃利全球学者计划)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title)

AI总结 BrainSymphony是一种参数高效、多模态的基础模型,通过整合fMRI和扩散MRI数据,实现有限数据下的脑动态分析,优于更大模型并提升神经科学应用。

Comments 32 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04824 2026-08-07 cs.CV 85%

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

SOVABench:多模态大语言模型的车辆监控动作检索基准

Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera

机构 * Milestone Systems A/S(Milestone Systems公司) Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心) Aalborg Universitet(奥胡斯大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 SOVABench为多模态大语言模型提供车辆监控动作检索基准,通过定义两种评估协议评估跨动作区分和时间方向理解,展示了模型在复杂监控任务中的性能。

Comments This work has been accepted at Real World Surveillance: Applications and Challenges, 6th (in WACV Workshops)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16978 2026-06-17 cs.CV 版本更新 85%

A Benchmark for Omni-Modal Reasoning in Long Videos

长视频全模态推理基准

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou, Sahal Shaji Mullappilly, Mohammad Almansoori, Noor Ahsan, Beknur Kalmakhanbet, Sambal Shikhar, Rishabh Lalla, Jean Lahoud, Mariette Awad, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) American University of Beirut(贝鲁特美国大学) Linköping University(利尔贝里大学)

专题命中 视频多模态 :omni-modal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 提出LongShOTBench基准,用于评估长视频中视觉、语音和环境音频的全模态推理,并引入无训练的全模态证据搜索代理LongShOTAgent,在105个模型上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03920 2026-06-03 cs.CV 85%

Benchmarking Visual State Tracking in Multimodal Video Understanding

多模态视频理解中的视觉状态追踪基准测试

Sihyun Yu, Nanye Ma, Pinzhi Huang, Hyunseok Lee, Shusheng Yang, June Suk Choi, Ellis Brown, Oscar Michel, Boyang Zheng, Jinwoo Shin, Saining Xie

机构 * New York University(纽约大学) KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出VSTAT基准,通过需要连续感知和整合整个视频流的问题评估多模态大语言模型的视觉状态追踪能力,发现当前模型远低于人类表现,失败主要源于视觉感知而非文本推理。

Comments Website: https://vision-x-nyu.github.io/vstat-site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11283 2026-06-02 cs.CV 85%

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

多模态大语言模型驱动的视频翻译:面向角色的综述

Bingzheng Qu, Kehai Chen, Xuefeng Bai, Min Zhang

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文通过面向角色的分类法,系统综述了多模态大语言模型在视频翻译中的应用,将其分为语义推理器、表达执行器和视觉合成器三个功能角色,并总结了数据集、基准和评估指标,指出了端到端视频翻译的挑战与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22185 2026-05-22 cs.CV cs.LG 85%

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

增强多模态大语言模型以用于安全关键驾驶视频分析

Tomaso Trinci, Henrique Piñeiro Monteagudo, Leonardo Taccari

机构 * Verizon Connect

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本研究通过融合降采样视频帧与同步高频 telemetry 数据及专用计算机视觉模型的语义信息,提升多模态大语言模型在安全关键驾驶场景中的感知与推理能力,从而更准确地识别和描述现实驾驶中的安全关键事件。

Comments Accepted at the 2026 IEEE International Conference on Intelligent Transportation Systems (ITSC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18431 2026-05-20 cs.CV 85%

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

协同视见:基于多模态大语言模型的多机器人协作自体空间推理

Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Hao Shi, Yi Zhou, M. Saquib Sarfraz, Danda Pani Paudel, Luc Van Gool

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Hunan University(湖南大学) University of Oxford(牛津大学) Zhejiang University(浙江大学) ETH Zurich(苏黎世联邦理工学院) Ant Group(蚂蚁集团)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文研究了多机器人协作动态空间推理问题,提出了首个针对该任务的基准CoopSR以及多机器人自体问答数据集EgoTeam,通过引入SP-CoR框架实现了细粒度的协作空间推理,显著提升了多机器人协作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22492 2026-04-27 eess.IV cs.CV 85%

MTT-Bench: Predicting Social Dominance in Mice via Multimodal Large Language Models

MTT-Bench:通过多模态大语言模型预测小鼠的社会支配

Yunquan Chen, Haoyu Chen

机构 * Department of Communication Systems, KTH Royal Institute of Technology(通信系统系,皇家理工学院) CMVS, University of Oulu(奥卢大学CMVS)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MTT-Bench基准,利用多模态大语言模型分析小鼠行为视频,预测社会支配等级,无需显式标签进行零样本推理,展示了在乙学和社会行为分析中的应用潜力。

Comments 8 pages, 2 figures. Submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12961 2025-02-28 cs.CV 85%

Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Zuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu, Jiwen Lu, Yongming Rao

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00832 2024-12-03 cs.CV 85%

EventGPT: Event Stream Understanding with Multimodal Large Language Models

Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Xin Meng, Fei Richard Yu, Xiangyang Ji, Ming Li

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏