arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2205.03569 2025-10-06 cs.CV cs.AI 81%

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

Bing Li, Jiaxin Chen, Dongming Zhang, Xiuguo Bao, Di Huang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25393 2025-10-02 cs.CV cs.AI 81%

Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction

Wendong Yao, Binhua Huang, Soumyabrata Dev

机构 * ADAPT SFI Research Centre, School of Computer Science, University College Dublin(ADAPT SFI研究所以及计算机科学学院,都柏林大学学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This paper is submitted to IEEE Transactions on Geoscience and Remote Sensing for reviewing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07032 2025-10-01 cs.CL cs.CV 81%

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫德赫·本·扎耶德人工智能大学) University of Central Florida(中央佛罗里达大学) Islamic University of Technology(伊斯兰技术大学) Air University(空军大学) ETH Zurich(苏黎世联邦理工学院) Technische Universität München(慕尼黑技术大学) National Institute of Informatics(国家信息研究所) Australian National University(澳大利亚国立大学) Linköping University(利尔贝里大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23044 2025-09-30 cs.CV cs.AI 81%

MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition

Ye-eun Kim, Suhyeon Lim, Andrew J. Choi

机构 * National Rehabilitation Center, Ministry of Health and Welfare, Korea(韩国卫生福利部国家康复中心) Gachon University(高丽大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18436 2025-09-30 cs.AI cs.CL cs.DB 81%

Memory-QA: Answering Recall Questions Based on Multimodal Memories

Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora, Shiun-Zu Kuo, Jiayang Xu, Ankur Bansal, Christopher Brossman, Yue Liu, Aaron Colak, Ahmed Aly, Anuj Kumar, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22014 2025-09-29 cs.CV cs.AI cs.HC cs.RO 81%

Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics

Saurav Jha, Stefan K. Ehrlich

机构 * SETLabs Resarch GmbH(SETL实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13763 2025-09-16 cs.CV cs.AI 81%

Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

Zhawnen Chen, Tianchun Wang, Yizhou Wang, Michal Kosinski, Xiang Zhang, Yun Fu, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) The Pennsylvania State University(宾夕法尼亚州立大学) Northeastern University(东北大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01177 2025-09-03 cs.CV cs.AI cs.HC eess.SP 81%

DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

Junxiang Liu, Junming Lin, Jiangtong Li, Jie Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00357 2025-09-03 cs.CV cs.AI cs.LG 81%

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院) CAMP, Technische Universität München(CAMP,慕尼黑技术大学) Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09362 2025-08-14 cs.CV cs.AI cs.LG 81%

FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition

Md. Milon Islam, Md Rezwanul Haque, S M Taslim Uddin Raju, Fakhri Karray

机构 * University of Waterloo(滑铁卢大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for the IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, Hawaii, USA. 1st MSLR Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04900 2025-08-08 cs.CV cs.AI 81%

Revealing Temporal Label Noise in Multimodal Hateful Video Classification

Shuonan Yang, Tailin Chen, Rahul Singh, Jiangbei Yue, Jianbo Jiao, Zeyu Fu

机构 * Multimodal Intelligence Lab(多模态智能实验室) Department of Computer Science(计算机科学系) University of Exeter(埃克塞特大学) University of Birmingham(伯明翰大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05185 2025-08-05 q-fin.CP cs.AI cs.MM 81%

Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance

Fengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Peking University(北京大学) University of Science and Technology of China(中国科学技术大学) Estates Pte Ltd(6Estates私人有限公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted by MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21161 2025-07-30 cs.CV cs.AI cs.LG 81%

Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues

Pallavi Zambare, Venkata Nikhil Thanikella, Ying Liu

机构 * Departmrnt of computer science(计算机科学系) Texas Tech University(得克萨斯科技大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in IEEE 3rd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18252 2025-07-25 cs.HC cs.AI cs.CL cs.LG 81%

Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning

Dongyang Guo, Yasmeen Abdrabou, Enkeleda Thaqi, Enkelejda Kasneci

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14766 2025-07-22 cs.LG cs.AI cs.CV 81%

CXR-TFT: Multi-Modal Temporal Fusion Transformer for Predicting Chest X-ray Trajectories

Mehak Arora, Ayman Ali, Kaiyuan Wu, Carolyn Davis, Takashi Shimazui, Mahmoud Alwakeel, Victor Moas, Philip Yang, Annette Esper, Rishikesan Kamaleswaran

机构 * Duke University(杜克大学) Emory University(埃默里大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments In Review for MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05939 2025-07-09 cs.CL cs.MM 81%

Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors

Bing Wang, Ximing Li, Mengzhe Ye, Changchun Li, Bo Fu, Jianfeng Qu, Lin Yuanbo Wu

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) College of Software, Jilin University(吉林大学软件学院) School of Computer and Artificial Intelligence, Liaoning Normal University(辽宁师范大学计算机与人工智能学院) School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Department of Computer Science, Swansea University(斯旺西大学计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted by ACM MM 2025. 10 pages, 6 figures. Code: https://github.com/wangbing1416/DAEDCMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02904 2025-07-08 cs.CV cs.AI 81%

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis

Charlton Teo

机构 * Department of Computer Science(计算机科学系) School of Computing(计算学院) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments B.Comp. dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15220 2025-06-18 cs.CV cs.AI 81%

Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Kun Yuan, Vinkle Srivastav, Tong Yu, Joel L. Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, Nicolas Padoy

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心,法国国家卫生研究院,ICube,UMR7357,法国斯特拉斯堡) University Digestive Health Care Center – Clarunis, 4002 Basel, Switzerland(消化健康研究中心–Clarunis,瑞士巴塞尔)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by Medical Image Analysis (MedIA), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13322 2025-06-17 cs.CV cs.AI 81%

Active Multimodal Distillation for Few-shot Action Recognition

Weijia Feng, Yichen Zhu, Ruojia Zhang, Chenyang Wang, Fei Ma, Xiaobao Wang, Xiaobai Li

机构 * College of Computer and Information Engineering, Tianjin Normal University(天津师范大学计算机与信息工程学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) College of Intelligence and Computing, Tianjin University(天津大学智能科学与计算学院) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江省区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments IJCAI 2025, the 34th International Joint Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10415 2025-06-13 cs.CL cs.CV 81%

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?

Yingjin Song, Yupei Du, Denis Paperno, Albert Gatt

机构 * Utrecht University(乌特雷赫大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 27 pages, 14 figures. Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09081 2025-06-11 cs.CV cs.AI 81%

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

Xiaowei Bi, Zheyuan Xu

机构 * Northwestern University(西北大学) IEEE Member(IEEE会员)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01757 2025-06-03 cs.CV cs.AI 81%

Efficient Egocentric Action Recognition with Multimodal Data

Marco Calzavara, Ard Kastrati, Matteo Macchini, Dushan Vasilevski, Roger Wattenhofer

机构 * ETH Zurich(苏黎世联邦理工学院) Magic Leap(Magic Leap公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as an extended abstract at the Second Joint Egocentric Vision (EgoVis) Workshop, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02406 2025-05-29 cs.CV cs.AI cs.DC cs.LG 81%

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models

Tzu-Tao Chang, Shivaram Venkataraman

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16630 2025-05-23 cs.CV cs.AI 81%

SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding

Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Mubarak Shah

机构 * SimulaMet and OsloMet, Norway(SimulaMet和OsloMet,挪威) Forzasys, Norway(Forzasys,挪威) SimulaMet, Norway(SimulaMet,挪威) SimulaMet, OsloMet, and Forzasys, Norway(SimulaMet、OsloMet和Forzasys,挪威) University of Central Florida, USA(佛罗里达中央大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16279 2025-05-23 cs.MM cs.CV 81%

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

Junjie Zheng, Zihao Chen, Chaofan Ding, Yunming Liang, Yihan Fan, Huan Yang, Lei Xie, Xinhan Di

机构 * Giant NetworkChina(巨网中国) The East China University of Science and TechnologyChina(东华大学) Northwestern Polytechnical UniversityChina(西北工业大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments 5 pages, 4 figures, accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12368 2025-05-21 cs.CV cs.CL 81%

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, Kai Chen, Dahua Lin, Jiaqi Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Nanjing University(南京大学) Fudan University(复旦大学) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11083 2025-05-21 cs.CV cs.CL 81%

Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning

Xiaohao Xu, Yunkang Cao, Huaxin Zhang, Nong Sang, Xiaonan Huang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Best Student Paper Award at IEEE International Conference on Computer Supported Cooperative Work in Design, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11926 2025-05-20 cs.CV cs.AI 81%

SafeVid: Toward Safety Aligned Video Large Multimodal Models

Yixu Wang, Jiaxin Song, Yifeng Gao, Xin Wang, Yang Yao, Yan Teng, Xingjun Ma, Yingchun Wang, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10836 2025-05-19 cs.CL cs.CV 81%

Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs

Abhishek Dey, Aabha Bothera, Samhita Sarikonda, Rishav Aryan, Sanjay Kumar Podishetty, Akshay Havalgi, Gaurav Singh, Saurabh Srivastava

机构 * George Mason University(乔治·马歇尔大学) Amazon(亚马逊公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01680 2025-05-06 cs.CV cs.AI cs.HC math.PR 81%

Automated ARAT Scoring Using Multimodal Video Analysis, Multi-View Fusion, and Hierarchical Bayesian Models: A Clinician Study

Tamim Ahmed, Thanassis Rikakis

机构 * University of Southern California(南加州大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏