arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2504.01591 2025-09-03 cs.CV 79%

Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval

Adriano Fragomeni, Dima Damen, Michael Wray

机构 * School of Computer Science University of Bristol(计算机科学学院英国布里斯托尔大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18284 2025-08-27 cs.LG cs.AI cs.SY eess.SY 79%

Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models

Rahmat K. Adesunkanmi, Alexander W. Brandt, Masoud Deylami, Gustavo A. Giraldo Echeverri, Hamidreza Karbasian, Adel Alaeddini

机构 * Departments of Mechanical and Electrical Engineering, Southern Methodist University(机械与电气工程系,南方 Methodist 大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17890 2025-08-26 cs.CV 79%

UniAPO: Unified Multimodal Automated Prompt Optimization

Qipeng Zhu, Yanzhe Chen, Huasong Zhong, Yan Li, Jie Chen, Zhixin Zhang, Junping Zhang, Zhenheng Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17817 2025-08-26 cs.CV 79%

TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration

Meiqi Gong, Hao Zhang, Xunpeng Yi, Linfeng Tang, Jiayi Ma

机构 * Electronic Information School, Wuhan University, Wuhan 430072, China(武汉大学电子信息学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14395 2025-08-21 cs.HC cs.AI 79%

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding

Running Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang, Handi Chen, Weipeng Deng, Luyao Jin, Xiaojuan Qi, Xun Qian, Edith C. H. Ngai

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Google(谷歌)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to UIST 2025. Project website: https://zhaorunning.github.io/NoteIt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13692 2025-08-20 cs.CV 79%

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

Keliang Li, Hongze Shen, Hao Shi, Ruibing Hou, Hong Chang, Jie Huang, Chenghao Jia, Wen Wang, Yiling Wu, Dongmei Jiang, Shiguang Shan, Xilin Chen

机构 * Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, China(中国科学院智能信息处理重点实验室(中国科学院)、计算技术研究所、中国)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 79%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10828 2025-08-15 cs.RO cs.AI 79%

A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots

Henry Powell, Guy Laban, Emily S. Cross

机构 * Amazon(亚马逊) School of Psychology and Neuroscience, University of Glasgow(心理学与神经科学学院,格拉斯哥大学) Ben-Gurion University of the Negev(内盖夫本·古里安大学) University of Cambridge(剑桥大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02935 2025-08-13 cs.CL 79%

Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in Conversation

Yuntao Shou, Tao Meng, Wei Ai, Keqin Li

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学教育部长江网络与网络安全重点实验室) College of Computer and Mathematics, Central South University of Forestry and Technology(中南林业科技大学计算机与数学学院) Department of Computer Science, State University of New York(纽约州立大学新帕尔茨分校计算机科学系)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06939 2025-08-12 cs.AI cs.LG 79%

Intrinsic Explainability of Multimodal Learning for Crop Yield Prediction

Hiba Najjar, Deepak Pathak, Marlon Nuske, Andreas Dengel

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03694 2025-08-06 cs.CV 79%

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Jianfeng Feng, Chenyang Si, Yanwei Fu, Yu Qiao, Ziwei Liu

机构 * Nanjing University(南京大学) Fudan University(复旦大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) NVIDIA(NVIDIA公司) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Project page: https://vchitect.github.io/LongVie-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02179 2025-08-05 cs.CV 79%

Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning

Wenbo Xu, Wei Lu, Xiangyang Luo

机构 * School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University(计算机科学与工程学院、信息科技教育部重点实验室、广东省信息安全技术重点实验室、中山大学) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages,4 figures. arXiv admin note: text overlap with arXiv:2507.16596

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01701 2025-08-05 cs.LG cs.AI 79%

MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model

Asmit Bandyopadhyay, Rohit Basu, Tanmay Sen, Swagatam Das

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01168 2025-08-05 cs.MM 79%

Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis

Hu Zhangfeng, Shi mengxin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06603 2025-08-04 cs.CV 79%

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition

Xu Shaowu, Jia Xibin, Gao Junyu, Sun Qianmei, Chang Jing, Fan Chao

机构 * Beijing University of Technology(北京理工大学) Chinese Academy of Sciences(中国科学院) Capital Medical University(首都医科大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22878 2025-07-31 cs.IR cs.CL cs.CY 79%

GeoOutageKG: A Multimodal Geospatiotemporal Knowledge Graph for Multiresolution Power Outage Analysis

Ethan Frakes, Yinghui Wu, Roger H. French, Mengjie Li

机构 * University of Central Florida(佛罗里达中央大学) Case Western Reserve University(凯斯西储大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to the 24th International Semantic Web Conference Resource Track (ISWC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 79%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10541 2025-07-30 cs.IR cs.AI 79%

Multi-Modal Hypergraph Enhanced LLM Learning for Recommendation

Xu Guo, Tong Zhang, Yuanzhi Wang, Chenxu Wang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, Zhen Cui

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Shituoyun (Nanjing) Technology Co., Ltd(石图云(南京)科技有限公司) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 12 pages, 4 figures, submitted to IEEE Transactions on Knowledge and Data Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20451 2025-07-29 cs.AI 79%

STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction

Pritom Ray Nobin, Imran Ahammad Rifat

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20163 2025-07-29 cs.CV 79%

Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning

Zeyu Xi, Haoying Sun, Yaofei Wu, Junchi Yan, Haoran Zhang, Lifang Wu, Liang Wang, Changwen Chen

机构 * Beijing University of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) Chinese Academy of Sciences(中国科学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025 (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18104 2025-07-28 cs.CV q-bio.NC 79%

A Multimodal Seq2Seq Transformer for Predicting Brain Responses to Naturalistic Stimuli

Qianyi He, Yuan Chang Leong

机构 * Data Science Institute University of Chicago(芝加哥大学数据科学研究所) Department of Psychology, Neuroscience Institute University of Chicago(芝加哥大学心理学系、神经科学研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17050 2025-07-24 cs.CV 79%

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas, Subarna Tripathi

机构 * Intel(英特尔)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to CVAM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13482 2025-07-21 cs.LG cs.CV 79%

Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning

Seyyed Saeid Cheshmi, Buyao Lyu, Thomas Lisko, Rajesh Rajamani, Robert A. McGovern, Yogatheesan Varatharajah

机构 * Department of Computer Science & Engineering University of Minnesota(计算机科学与工程系明尼苏达大学) Department of Mechanical Engineering University of Minnesota(机械工程系明尼苏达大学) Department of Neurosurgery University of Minnesota(神经外科系明尼苏达大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13403 2025-07-21 cs.CV cs.LG 79%

UL-DD: A Multimodal Drowsiness Dataset Using Video, Biometric Signals, and Behavioral Data

Morteza Bodaghi, Majid Hosseini, Raju Gottumukkala, Ravi Teja Bhupatiraju, Iftikhar Ahmad, Moncef Gabbouj

机构 * University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校) Tietoevry(蒂奥维瑞) Tampere University(塔尔库大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12666 2025-07-18 cs.AI cs.LG 79%

Fly, Fail, Fix: Iterative Game Repair with Reinforcement Learning and Large Multimodal Models

Alex Zook, Josef Spjut, Jonathan Tremblay

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published at Reinforcement Learning and Video Games workshop https://sites.google.com/view/rlvg-workshop-2025/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15681 2025-07-18 cs.CV 79%

Vidi: Large Multimodal Models for Video Understanding and Editing

Vidi Team, Celong Liu, Chia-Wen Kuo, Dawei Du, Fan Chen, Guang Chen, Jiamin Yuan, Lingxi Zhang, Lu Guo, Lusha Li, Longyin Wen, Qingyu Chen, Rachel Deng, Sijie Zhu, Stuart Siew, Tong Jin, Wei Lu, Wen Zhong, Xiaohui Shen, Xin Gu, Xing Mei, Xueqiong Qu, Zhenfang Chen

机构 * Intelligent Editing Team(智能编辑团队) Intelligent Creation, ByteDance Inc.(智能创作,字节跳动公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09617 2025-07-15 cs.AI cs.RO 79%

Bridging Bots: from Perception to Action via Multimodal-LMs and Knowledge Graphs

Margherita Martorana, Francesca Urgese, Mark Adamik, Ilaria Tiddi

机构 * Vrije Universiteit Amsterdam(范·艾克大学阿姆斯特丹)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07938 2025-07-11 cs.MM 79%

Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency

Abolfazl Zarghani, Amirhossein Ebrahimi, Amir Malekesfandiari

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06072 2025-07-09 cs.CV 79%

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

Tongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang, Jing Wang, Kai Liu

机构 * Department of Computer Science, Chongqing University, China(重庆大学计算机科学系) National Elite Institute of Engineering, Chongqing University, China(重庆大学工程精英研究院) College of Computer Science and Technology, National University of Deffense Technology, China(国防科技大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05899 2025-07-09 cs.CV 79%

What You Have is What You Track: Adaptive and Robust Multimodal Tracking

Yuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li, Zhaochong An, Chao Ma, Danda Paudel, Luc Van Gool, Radu Timofte, Zongwei Wu

机构 * TeleAI, China Telecom(TeleAI,中国电信) Computer Vision Lab, CAIDAS & IFI, University of Wurzburg(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学) INSAIT, Sofia University(INSAIT,索菲亚大学) ShanghaiTech University(上海科技大学) University of Copenhagen(哥本哈根大学) AI Institute, Shanghai Jiao Tong University(人工智能研究院,上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICCV2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏