arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4699 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4699 篇

2410.12828 2024-10-18 cs.CV cs.LG 85%

GCM-Net: Graph-enhanced Cross-Modal Infusion with a Metaheuristic-Driven Network for Video Sentiment and Emotion Analysis

Prasad Chaudhari, Aman Kumar, Chandravardhan Singh Raghaw, Mohammad Zia Ur Rehman, Nagendra Kumar

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17880 2024-06-27 cs.CV 85%

MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval

Weitong Cai, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu

专题命中 视频多模态 :MLLM(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04154 2024-01-10 cs.CV cs.AI cs.LG cs.MM cs.SD eess.AS 85%

Efficient Selective Audio Masked Multimodal Bottleneck Transformer for Audio-Video Classification

Wentao Zhu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by WACV 2024; well-formatted PDF is in https://drive.google.com/file/d/1qvW52lamsvNGMCqPS7q8g8L4NaR_LlbR/view?usp=sharing. arXiv admin note: text overlap with arXiv:2401.04023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17642 2023-10-27 cs.RO cs.CV cs.LG 85%

Drive Anywhere: Generalizable End-to-end Autonomous Driving with Multi-modal Foundation Models

Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, Daniela Rus

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Project webpage: https://drive-anywhere.github.io Explainer video: https://www.youtube.com/watch?v=4n-DJf8vXxo&feature=youtu.be

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07646 2022-07-18 cs.CV cs.LG 85%

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Rui Qian, Yeqing Li, Zheng Xu, Ming-Hsuan Yang, Serge Belongie, Yin Cui

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12485 2022-03-30 cs.CV 85%

CroMo: Cross-Modal Learning for Monocular Depth Estimation

Yannick Verdié, Jifei Song, Barnabé Mas, Benjamin Busam, Aleš Leonardis, Steven McDonagh

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted for publication at CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07515 2021-12-15 cs.CV cs.AI cs.CL cs.MM 85%

CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising

Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao, Hongyang Chao, Tao Mei

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACM Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10949 2021-10-22 cs.CV 85%

Multimodal Learning using Optimal Transport for Sarcasm and Humor Detection

Shraman Pramanick, Aniket Roy, Vishal M. Patel

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to WACV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04124 2021-09-23 cs.CV 85%

Parameter Efficient Multimodal Transformers for Video Representation Learning

Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, Yale Song

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted to ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.03152 2017-04-12 cs.CV 85%

Deep Multimodal Representation Learning from Temporal Data

Xitong Yang, Palghat Ramesh, Radha Chitta, Sriganesh Madhvanath, Edgar A. Bernal, Jiebo Luo

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments To appear in CVPR 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10024 2026-07-20 cs.CV cs.AI cs.LG 版本更新 85%

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

LVSum:一个用于时间感知长视频摘要的基准测试

Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan

机构 * Apple(苹果公司)

专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出LVSum基准测试,用于评估长视频摘要中时间对齐的性能,通过引入新的评估指标揭示现有MLLM在时间理解上的系统性差距。

Comments 25 pages, 5 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16560 2026-07-21 cs.AI cs.CV cs.LG cs.MM 新提交 85%

From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence

从模态到命题:多模态智能的语言中心框架

Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez

机构 * NVIDIA(英伟达) University of Ottawa(渥太华大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 该研究提出多模态数据语言表示框架,将观察结果表示为原子命题,通过全局语义码本统一为共享词汇表,置于可解释空间,实现跨模态理解等,还在自动驾驶等数据上进行了展示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02945 2026-04-06 cs.DC 85%

MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

MSAO:基于边缘-云协作的自适应模态稀疏性感知卸载框架用于高效多模态大语言模型推理

Zheming Yang, Qi Guo, Jun Wan, Jiarui Ruan, Yunqing Hu, Chang Zhao, Xiangyang Li

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract)

AI总结 本文提出MSAO框架,通过边缘-云协作和模态稀疏性分析,降低多模态大语言模型推理的延迟和资源消耗,提升吞吐量。

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14267 2026-04-06 cs.CV cs.AI cs.MM cs.SD 85%

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

DiFlowDubber:通过跨模态对齐与同步实现的离散流匹配自动化视频配音

Ngoc-Son Nguyen, Thanh V. T. Tran, Jeongsoo Choi, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen

机构 * FPT Software AI Center, Vietnam(FPT软件人工智能中心,越南) KAIST, South Korea(韩国科学技术院) University of Alabama at Birmingham, USA(阿拉巴马大学伯明翰分校)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出DiFlowDubber框架,通过离散流匹配和两阶段训练策略,解决视频配音中内容准确性、表达语气、高质量音频和精确唇同步的问题。

Comments Accepted at CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28610 2026-04-01 cs.CV cs.AI cs.CL 85%

ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning

ResAdapt:面向高效多模态推理的自适应分辨率

Huanxuan Liao, Zhongtao Jiang, Yupu Hao, Yuqiao Tan, Shizhu He, Ben Wang, Jun Zhao, Kun Xu, Kang Liu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 ResAdapt通过自适应输入分辨率框架,在保持高空间分辨率的同时提升多模态推理效率,尤其在压缩条件下显著提升性能。

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04356 2025-12-05 cs.CV cs.AI cs.CL cs.LG 85%

Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment

通过自增强对比对齐缓解多模态大语言模型中的对象和动作幻觉

Kai-Po Chang, Wei-Yuan Cheng, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究所) NVIDIA

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 SANTA框架通过自增强对比对齐方法,有效缓解多模态大语言模型中的对象和动作幻觉问题。

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. Project page: https://kpc0810.github.io/santa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15870 2025-10-29 cs.CV cs.AI cs.CL 85%

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, Zhijian Liu, Yukang Chen, Ambrish Dantrey, Ehsan Jahangiri, Sreyan Ghosh, Daguang Xu, Ehsan Hosseini-Asl, Danial Mohseni Taheri, Vidya Murali, Sifei Liu, Yao Lu, Oluwatobi Olabiyi, Yu-Chiang Frank Wang, Rafael Valle, Bryan Catanzaro, Andrew Tao, Song Han, Jan Kautz, Hongxu Yin, Pavlo Molchanov

机构 * NVIDIA

专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report. Code: https://github.com/NVlabs/OmniVinci

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04651 2025-08-29 cs.IR 85%

FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation

Maolin Wang, Yutian Xiao, Binhao Wang, Sheng Zhang, Shanshan Ye, Wanyu Wang, Hongzhi Yin, Ruocheng Guo, Zenglin Xu

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)

Comments Accepted by KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17204 2025-07-24 cs.LG 85%

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation

Zixuan Wang, Jinghao Shi, Hanzhong Liang, Xiang Shen, Vera Wen, Zhiqian Chen, Yifan Wu, Zhixin Zhang, Hongyu Xiong

机构 * TikTok(字节跳动)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract)

Comments Camera Ready for ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10604 2025-06-18 cs.CV cs.AI cs.CL 85%

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

Ruixuan Zhang, Beichen Wang, Juexiao Zhang, Zilin Bian, Chen Feng, Kaan Ozbay

机构 * New York University(纽约大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10011 2025-06-13 cs.MM cs.AI cs.CV eess.SP 85%

WDMIR: Wavelet-Driven Multimodal Intent Recognition

Weiyin Gong, Kai Zhang, Yanghai Zhang, Qi Liu, Xinjie Sun, Junyu Lu, Linbo Zhu

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能国家重点实验室) School of Computer Science, Liupanshui Normal University(黎平师范学院计算机科学学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted at IJCAI 2025, 9pages, 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06814 2025-05-13 cs.CV cs.AI cs.CL 85%

Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge

Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, Shoujun Zhou

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Engineering, Westlake University(工程学院,西湖大学)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12623 2025-05-05 cs.LG cs.AI cs.CV cs.MM 85%

MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network

Vrushank Ahire, Kunal Shah, Mudasir Nazir Khan, Nikhil Pakhale, Lownish Rai Sookha, M. A. Ganaie, Abhinav Dhall

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) Monash University(墨尔本大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12231 2025-01-22 cs.CV cs.AI cs.CL 85%

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Pha Nguyen, Sailik Sengupta, Girik Malik, Arshit Gupta, Bonan Min

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19493 2024-12-30 cs.CV cs.AI cs.MM 85%

Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection

Yihao Wang, Lizhi Chen, Zhong Qian, Peifeng Li

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18286 2024-09-30 cs.CV cs.AI cs.CL 85%

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

Huthaifa I. Ashqar, Ahmed Jaber, Taqwa I. Alhadidi, Mohammed Elhenawy

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00552 2024-09-04 eess.AS cs.CV cs.MM cs.SD 85%

Digit Recognition using Multimodal Spiking Neural Networks

William Bjorndahl, Jack Easton, Austin Modoff, Eric C. Larson, Joseph Camp, Prasanna Rangarajan

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00022 2024-09-04 cs.MM cs.AI cs.CV 85%

Detecting Misinformation in Multimedia Content through Cross-Modal Entity Consistency: A Dual Learning Approach

Zhe Fu, Kanlun Wang, Wangjiaxuan Xin, Lina Zhou, Shi Chen, Yaorong Ge, Daniel Janies, Dongsong Zhang

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to PACIS 2024. 15 pages, 3 figures

Journal ref https://aisel.aisnet.org/pacis2024/track07_secprivacy/track07_secprivacy/2

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15164 2024-01-30 cs.SD cs.CV cs.LG cs.MM eess.AS 85%

AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations

Naresh Kumar Devulapally, Sidharth Anand, Sreyasee Das Bhattacharjee, Junsong Yuan, Yu-Ping Chang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00220 2023-12-04 cs.MM cs.CL cs.CV 85%

Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation

Linzi Xing, Quan Tran, Fabian Caba, Franck Dernoncourt, Seunghyun Yoon, Zhaowen Wang, Trung Bui, Giuseppe Carenini

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at the 30th International Conference on Multimedia Modeling (MMM 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏