arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2411.08466 2025-06-10 cs.CV 83%

Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

Quan Zhang, Jinwei Fang, Rui Yuan, Xi Tang, Yuxin Qi, Ke Zhang, Chun Yuan

机构 * Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00304 2025-06-04 cs.CV 83%

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa, Shao-Yuan Lo

机构 * Johns Hopkins University(约翰霍普金斯大学) Honda Research Institute USA(本田研究院美国)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Paper is accepted by IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08282 2025-06-03 cs.CV 83%

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang, Tianrui Hui, Jialin Gao, Xiaoming Wei, Si Liu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Meituan(美团)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24476 2025-06-02 cs.CV 83%

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model

Yuting Zhang, Hao Lu, Qingyong Hu, Yin Wang, Kaishen Yuan, Xin Liu, Kaishun Wu

机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science & Technology(香港科技大学) Zhejiang University(浙江大学) Lappeenranta-Lahti University of Technology(拉佩兰塔-拉赫蒂技术大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19400 2025-05-27 cs.AI cs.CL cs.CV cs.MM 83%

TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding

Max Ku, Thomas Chong, Jonathan Leung, Krish Shah, Alvin Yu, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Votee AI Vector Institute(向量研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to ACL 2025 main, camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11865 2025-05-19 cs.CV 83%

From Image to Video, what do we need in multimodal LLMs?

Suyuan Huang, Haoxin Zhang, Linqing Zhong, Honggu Chen, Yan Gao, Yao Hu, Zengchang Qin

机构 * Intelligent Computing and Machine Learning Lab, School of ASEE, Beihang University(北京航空航天大学自动化学院智能计算与机器学习实验室) Xiaohongshu(小红书) School of Sino-French Engineer, Beihang University(北京航空航天大学中法工程师学院) College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05714 2025-05-12 cs.CL 83%

TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries

Jinze Lv, Jian Chen, Zi Long, Xianghua Fu, Yin Chen

机构 * College of Application and Technology, Shenzhen University, China(应用技术学院,深圳大学,中国) College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02096 2025-05-06 cs.MM 83%

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing

Yaru Chen, Peiliang Zhang, Fei Li, Faegheh Sardari, Ruohao Guo, Zhenbo Li, Wenwu Wang

专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM

Comments Accepted by ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13983 2025-04-14 cs.CV 83%

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li, Jiannan Ge, Hongtao Xie, Yongdong Zhang

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03735 2025-04-02 cs.CV 83%

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li, Eun Woo Im, Pooyan Fazli

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23660 2025-04-01 cs.CV 83%

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

Junjie Zheng, Zihao Chen, Chaofan Ding, Xinhan Di

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19406 2025-03-26 cs.CV 83%

M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation

Ziyuan Liu, Jiawei Zhang, Wenyu Wang, Yuantao Gu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19134 2025-03-26 cs.CL cs.CR 83%

MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks

Wenhao You, Bryan Hooi, Yiwei Wang, Youke Wang, Zong Ke, Ming-Hsuan Yang, Zi Huang, Yujun Cai

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13281 2025-03-25 cs.CV cs.AI cs.CL cs.MM 83%

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, Junnan Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2025, Project Page: https://videoautoarena.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10523 2025-03-14 cs.CV 83%

Interactive Multimodal Fusion with Temporal Modeling

Jun Yu, Yongqi Wang, Lei Wang, Yang Zheng, Shengfan Xu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17599 2025-03-14 cs.CL 83%

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Hui Shen, Xin Wang, Che Liu, Zheda Mai, Mi Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20105 2024-12-31 cs.CV 83%

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

Jiedong Zhuang, Lu Lu, Ming Dai, Rui Hu, Jian Chen, Qiang Liu, Haoji Hu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18060 2024-12-25 cs.CV 83%

An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM

Wen Wen, Yilin Wang, Neil Birkbeck, Balu Adsumilli

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14006 2024-12-19 cs.CV 83%

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng, Yong Liu, Zheng Zhao, Yujiu Yang

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15829 2024-12-03 cs.CV 83%

SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization

Sicheng Liu, Lintao Wang, Xiaogang Zhu, Xuequan Lu, Zhiyong Wang, Kun Hu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures, submitted to ACM Multimedia Asia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17773 2024-12-02 cs.CV 83%

XTrack: Multimodal Training Boosts RGB-X Video Object Trackers

Yuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou, Guolei Sun, Eduard Zamfi, Chao Ma, Danda Pani Paudel, Luc Van Gool, Radu Timofte

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09875 2024-10-15 cs.CV cs.IR 83%

ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification

Chen Mao, Chong Tan, Jingqi Hu, Min Zheng

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04955 2024-10-01 cs.CV 83%

Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations

Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang, Peng Zhai, Song Wang, Lihua Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by TCSVT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05930 2024-09-18 cs.CV 83%

MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

Rezaul Karim, He Zhao, Richard P. Wildes, Mennatullah Siam

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments Extension of CVPR'23 paper for journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10213 2024-09-17 cs.CV 83%

Neuromorphic Facial Analysis with Cross-Modal Supervision

Federico Becattini, Luca Cultrera, Lorenzo Berlincioni, Claudio Ferrari, Andrea Leonardo, Alberto Del Bimbo

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted for publication at the ECCV 2024 workshop on Neuromorphic Vision: Advantages and Applications of Event Cameras (NEVI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09362 2024-09-17 cs.CL 83%

Generating Event-oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM

Yuanjie Lyu, Tong Xu, Zihan Niu, Bo Peng, Jing Ke, Enhong Chen

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08150 2024-09-06 cs.CV 83%

Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding

Minghui Wu, Chenxu Zhao, Anyang Su, Donglin Di, Tianyu Fu, Da An, Min He, Ya Gao, Meng Ma, Kun Yan, Ping Wang

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ACM MULTIMEDIA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01766 2024-08-20 cs.CV 83%

MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition

Ruoyu Wang, Wenqian Wang, Jianjun Gao, Dan Lin, Kim-Hui Yap, Bingbing Li

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05523 2024-08-15 cs.HC cs.CV 83%

DeepFace-Attention: Multimodal Face Biometrics for Attention Estimation with Application to e-Learning

Roberto Daza, Luis F. Gomez, Julian Fierrez, Aythami Morales, Ruben Tolosana, Javier Ortega-Garcia

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Article accepted in the IEEE Access journal. Accessible at https://ieeexplore.ieee.org/document/10633208

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11340 2024-08-06 cs.CV cs.LG 83%

CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition

Ruoyu Wang, Chen Cai, Wenqian Wang, Jianjun Gao, Dan Lin, Wenyang Liu, Kim-Hui Yap

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏