arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2404.01932 2025-05-29 cs.RO cs.LG 82%

Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks

Gabriela Sejnova, Michal Vavrecka, Karla Stepanova

机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) Czech Technical University in Prague(布拉格捷克技术大学)

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract)

Comments 7 pages, 5 figures, 2 tables, conference

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14535 2025-05-21 cs.LG cs.HC 82%

Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning

Jiangrong Shen, Yulin Xie, Qi Xu, Gang Pan, Huajin Tang, Badong Chen

机构 * Faculty of Electronic and Information Engineering, Xi’an Jiaotong University(电子与信息工程学院,西安交通大学) Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究院,西安交通大学) State Key Lab of Brain-Machine Intelligence, Zhejiang University(脑机智能国家重点实验室,浙江大学) School of Computer Science, Dalian University of Technology(计算机科学学院,大连理工大学) National Key Lab of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(人机混合增强智能国家实验室,西安交通大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12051 2025-05-20 cs.MM cs.AI cs.CV 82%

Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion

Yinghui Zhang, Tailin Chen, Yuchen Zhang, Zeyu Fu

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICDMW 2024, Github: https://github.com/EvelynZ10/cmfusion

Journal ref 2024 IEEE International Conference on Data Mining Workshops (ICDMW), Abu Dhabi, United Arab Emirates, 2024, pp. 183-190

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16036 2025-03-21 cs.CV cs.AI cs.CL 82%

Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

Zhihang Liu, Chen-Wei Xie, Pandeng Li, Liming Zhao, Longxiang Tang, Yun Zheng, Chuanbin Liu, Hongtao Xie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05889 2025-03-21 cs.CV cs.AI cs.CL 82%

CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion

Shoubin Yu, Jaehong Yoon, Mohit Bansal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025; first two authors contributed equally. Project page: https://CREMA-VideoLLM.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05092 2025-03-19 cs.CV cs.AI cs.CL 82%

Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs

Rohit Saxena, Aryo Pradipta Gema, Pasquale Minervini

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at the ICLR 2025 Workshop on Reasoning and Planning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16755 2025-02-25 cs.CY 82%

Watch Out E-scooter Coming Through: Multimodal Sensing of Mixed Traffic Use and Conflicts Through Riders Ego-centric Views

Hiruni Nuwanthika Kegalle, Danula Hettiachchi, Jeffrey Chan, Mark Sanderson, Flora D. Salim

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract)

Comments Accepted in Proc. ACM Interactive, Mobile, Wearable and Ubiquitous Technologies,(March 2025), 23 pages. https://doi.org/10.1145/3712284

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05474 2025-01-13 cs.CL cs.AI cs.LG cs.SD eess.AS 82%

Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis

Xincheng Wang, Liejun Wang, Yinfeng Yu, Xinxin Jiao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted for publication by 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18748 2025-01-03 cs.MM cs.CL cs.SD eess.AS 82%

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

Yuan Zhao, Rui Liu, Gaoxiang Cong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16407 2024-10-23 cs.CL cs.AI cs.MM 82%

Enhancing Multimodal Affective Analysis with Learned Live Comment Features

Zhaoyuan Deng, Amith Ananthram, Kathleen McKeown

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10818 2024-10-16 cs.CV cs.AI cs.CL cs.LG 82%

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao, Yong Jae Lee, Jianwei Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://temporalbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19467 2024-10-11 cs.CL cs.AI cs.CV 82%

TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning

Kate Sanders, Nathaniel Weir, Benjamin Van Durme

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 9 pages, EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15766 2024-10-04 cs.AI cs.CL cs.CV 82%

Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development

Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Aman Chadha, Samrat Mondal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11593 2024-09-05 cs.MM cs.CV cs.SD eess.AS 82%

MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Yuan Zhao, Zhenqi Jia, Rui Liu, De Hu, Feilong Bao, Guanglai Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by NCMMSC2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14895 2024-08-29 cs.AI cs.CL cs.CV 82%

VHAKG: A Multi-modal Knowledge Graph Based on Synchronized Multi-view Videos of Daily Activities

Shusaku Egami, Takahiro Ugai, Swe Nwe Nwe Htun, Ken Fukuda

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 5 pages, 4 figures, accepted by CIKM2024 Resource Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07694 2024-08-15 cs.CV cs.AI cs.LG cs.MM 82%

End-to-end Semantic-centric Video-based Multimodal Affective Computing

Ronghao Lin, Ying Zeng, Sijie Mai, Haifeng Hu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10711 2024-07-24 cs.CV cs.AI cs.CL 82%

Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering

Haibo Wang, Chenghang Lai, Yixuan Sun, Weifeng Ge

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted by ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08743 2024-06-18 cs.AI cs.CL cs.CV cs.LG 82%

MMToM-QA: Multimodal Theory of Mind Question Answering

Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua B. Tenenbaum, Tianmin Shu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2024. 26 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06964 2024-06-12 cs.CL cs.MM cs.SD eess.AS 82%

Missingness-resilient Video-enhanced Multimodal Disfluency Detection

Payal Mohapatra, Shamika Likhite, Subrata Biswas, Bashima Islam, Qi Zhu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted to Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16728 2024-05-28 cs.CV cs.AI cs.LG cs.MM 82%

Towards Multi-Task Multi-Modal Models: A Video Generative Perspective

Lijun Yu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05291 2024-04-01 cs.CV cs.AI cs.CL 82%

GlitchBench: Can large multimodal models detect video game glitches?

Mohammad Reza Taesiri, Tianjun Feng, Anh Nguyen, Cor-Paul Bezemer

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02051 2024-03-29 cs.CV cs.AI cs.CL 82%

TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, Lu Hou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2024 camera-ready version, code is available at https://github.com/RenShuhuai-Andy/TimeChat

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17172 2023-12-29 cs.CV cs.AI cs.CL 82%

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, Aniruddha Kembhavi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 38 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00347 2023-12-19 cs.CV cs.CL cs.MM 82%

RTQ: Rethinking Video-language Understanding Based on Image-text Model

Xiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang, Jingjing Lv, Liqiang Nie

专题命中 视频多模态 :image-text(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by ACM MM 2023 as Oral representation

Journal ref In International Conference on Multimedia. ACM, 557--566 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03741 2023-08-08 cs.CV cs.AI cs.LG cs.MM 82%

MAiVAR-T: Multimodal Audio-image and Video Action Recognizer using Transformers

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, Naveed Akhtar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 6 pages, 7 figures, 4 tables, Peer reviewed, Accepted @ The 11th European Workshop on Visual Information Processing (EUVIP) will be held on 11th-14th September 2023, in Gjøvik, Norway. arXiv admin note: text overlap with arXiv:2103.15691 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03625 2023-05-11 cs.CL cs.CV cs.MM 82%

C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval

Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogerio Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, James Glass

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at ICASSP 2023. The code, models, and dataset are available at https://github.com/roudimit/c2kd

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03915 2023-05-09 cs.CV cs.CL cs.MM 82%

HateMM: A Multi-Modal Dataset for Hate Video Classification

Mithun Das, Rohit Raj, Punyajoy Saha, Binny Mathew, Manish Gupta, Animesh Mukherjee

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at ICWSM 2023(dataset track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08875 2023-04-12 physics.optics cs.GR 82%

Multi-modal imaging using a cascaded microscope design

Xi Yang, Mark Harfouche, Kevin C. Zhou, Lucas Kreiss, Shiqi Xu, Kanghyun Kim, Roarke Horstmeyer

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01814 2023-01-23 cs.LG 82%

Multimodal Frame-Scoring Transformer for Video Summarization

Jeiyoon Park, Kiho Kwoun, Chanhee Lee, Heuiseok Lim

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract)

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08481 2022-10-18 cs.CV cs.CL cs.MM 82%

TLDW: Extreme Multimodal Summarisation of News Videos

Peggy Tang, Kun Hu, Lei Zhang, Jiebo Luo, Zhiyong Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏