arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2502.15363 2025-03-17 cs.HC cs.CV 79%

M2LADS Demo: A System for Generating Multimodal Learning Analytics Dashboards

Alvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Julian Fierrez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Published in the Workshop on Innovation and Responsibility in AI-Supported Education (iRAISE25) at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08370 2025-03-12 cs.GR cs.CV 79%

Ev-Layout: A Large-scale Event-based Multi-modal Dataset for Indoor Layout Estimation and Tracking

Xucheng Guo, Yiran Shen, Xiaofang Xiao, Yuanfeng Zhou, Lin Wang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05936 2025-03-11 cs.CV 79%

CASP: Compression of Large Multimodal Models Based on Attention Sparsity

Mohsen Gholami, Mohammad Akbari, Kevin Cannons, Yong Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15393 2025-03-04 cs.CV 79%

LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models

Hongchen Wei, Zhihong Tan, Yaosi Hu, Chang Wen Chen, Zhenzhong Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18371 2025-02-26 cs.AI 79%

MindMem: Multimodal for Predicting Advertisement Memorability Using LLMs and Deep Learning

Sepehr Asgarian, Qayam Jetha, Jouhyun Jeon

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, 5 figures, 4 Tables, AAAI 2025 Economics of Modern ML: Markets, Incentives, and Generative AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17038 2025-02-25 cs.MM 79%

Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction

Jiacheng Lu, Mingyuan Xiao, Weijian Wang, Yuxin Du, Zhengze Wu, Cheng Hua

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17359 2025-02-24 cs.CL 79%

Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

Zhenyu Pan, Haozheng Luo, Manling Li, Han Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments International Conference on Learning Representations (ICLR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14227 2025-02-21 cs.LG cs.AI 79%

SleepGMUformer: A gated multimodal temporal neural network for sleep staging

Chenjun Zhao, Xuesen Niu, Xinglin Yu, Long Chen, Na Lv, Huiyu Zhou, Aite Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13716 2025-02-20 cs.CV 79%

Event-Based Video Frame Interpolation With Cross-Modal Asymmetric Bidirectional Motion Fields

Taewoo Kim, Yujeong Chae, Hyun-Kurl Jang, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted in CVPR2023(Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12604 2025-02-19 cs.CV 79%

S2C: Learning Noise-Resistant Differences for Unsupervised Change Detection in Multimodal Remote Sensing Images

Lei Ding, Xibing Zuo, Danfeng Hong, Haitao Guo, Jun Lu, Zhihui Gong, Lorenzo Bruzzone

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06332 2025-02-18 cs.CV 79%

X-VARS: Introducing Explainability in Football Refereeing with Multi-Modal Large Language Model

Jan Held, Hani Itani, Anthony Cioppa, Silvio Giancola, Bernard Ghanem, Marc Van Droogenbroeck

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09895 2025-02-11 cs.CV 79%

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

Yating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv, Lingtong Min, Yanning Zhang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14269 2025-01-31 cs.IR cs.AI 79%

Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation

Shengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang, Hui Xiong

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15508 2025-01-28 cs.MM 79%

Learning Complex Heterogeneous Multimodal Fake News via Social Latent Network Inference

Mingxin Li, Yuchen Zhang, Haowei Xu, Xianghua Li, Chao Gao, Zhen Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10692 2025-01-22 cs.CV 79%

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Zien Xie, Youyao Jia, Sidan Du

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09499 2025-01-17 cs.CV 79%

VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

Zixun Fang, Zhiheng Liu, Kai Zhu, Yu Liu, Ka Leong Cheng, Wei Zhai, Yang Cao, Zheng-Jun Zha

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06475 2025-01-14 cs.CV cs.LG 79%

Enhancing Multi-Modal Video Sentiment Classification Through Semi-Supervised Clustering

Mehrshad Saadatinia, Minoo Ahmadi, Armin Abdollahi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05884 2025-01-13 cs.CV 79%

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

Dabing Cheng, Haosen Zhan, Xingchen Zhao, Guisheng Liu, Zemin Li, Jinghui Xie, Zhao Song, Weiguo Feng, Bingyue Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 16pages conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05733 2025-01-13 cs.CV 79%

TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos

Korawat Charoenpitaks, Van-Quang Nguyen, Masanori Suganuma, Kentaro Arai, Seiji Totsuka, Hiroshi Ino, Takayuki Okatani

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Main Paper: 8 pages, Supplementary Materials: 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11280 2025-01-09 cs.CV 79%

ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO

Daechul Ahn, Yura Choi, San Kim, Youngjae Yu, Dongyeop Kang, Jonghyun Choi

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00546 2025-01-08 cs.AI cs.LG 79%

AllSpark: A Multimodal Spatio-Temporal General Intelligence Model with Ten Modalities via Language as a Reference Framework

Run Shao, Cheng Yang, Qiujun Li, Qing Zhu, Yongjun Zhang, YanSheng Li, Yu Liu, Yong Tang, Dapeng Liu, Shizhong Yang, Haifeng Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 19 pages, 19 tables, 3 figures

Journal ref IEEE Transactions on Geoscience and Remote Sensing. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03103 2025-01-07 cs.CV 79%

MVP: Multimodal Emotion Recognition based on Video and Physiological Signals

Valeriya Strizhkova, Hadi Kachmar, Hava Chaptoukaev, Raphael Kalandadze, Natia Kukhilava, Tatia Tsmindashvili, Nibras Abo-Alzahab, Maria A. Zuluaga, Michal Balazia, Antitza Dantcheva, François Brémond, Laura Ferrari

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Preprint. Final paper accepted at Affective Behavior Analysis in-the-Wild (ABAW) at IEEE/CVF European Conference on Computer Vision (ECCV), Milan, September, 2024. 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00795 2025-01-03 cs.CV 79%

Multimodal Large Models Are Effective Action Anticipators

Binglu Wang, Yao Tian, Shunzhou Wang, Le Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19842 2024-12-31 cs.CV 79%

Multimodal joint prediction of traffic spatial-temporal data with graph sparse attention mechanism and bidirectional temporal convolutional network

Dongran Zhang, Jiangnan Yan, Kemal Polat, Adi Alhudhaif, Jun Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15798 2024-12-30 eess.IV cs.CV 79%

M3-CVC: Controllable Video Compression with Multimodal Generative Models

Rui Wan, Qi Zheng, Yibo Fan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17832 2024-12-25 eess.SP cs.AI cs.LG 79%

MANGO: Multimodal Acuity traNsformer for intelliGent ICU Outcomes

Jiaqing Zhang, Miguel Contreras, Sabyasachi Bandyopadhyay, Andrea Davidson, Jessica Sena, Yuanfang Ren, Ziyuan Guan, Tezcan Ozrazgat-Baslanti, Tyler J. Loftus, Subhash Nerella, Azra Bihorac, Parisa Rashidi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19944 2024-12-25 cs.CV 79%

A Multimodal Approach For Endoscopic VCE Image Classification Using BiomedCLIP-PubMedBERT

Nagarajan Ganapathy, Podakanti Satyajith Chary, Teja Venkata Ramana Kumar Pithani, Pavan Kavati, Arun Kumar S

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 11 Pages, 2 Figures, Capsule Vision 2024 Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15691 2024-12-23 cs.CV 79%

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

Xiantao Hu, Ying Tai, Xu Zhao, Chen Zhao, Zhenyu Zhang, Jun Li, Bineng Zhong, Jian Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10410 2024-12-17 cs.AI cs.LG cs.RO 79%

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents

Shaofei Cai, Bowei Zhang, Zihao Wang, Haowei Lin, Xiaojian Ma, Anji Liu, Yitao Liang

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10219 2024-12-16 cs.CV 79%

Learning Complex Non-Rigid Image Edits from Multimodal Conditioning

Nikolai Warner, Jack Kolb, Meera Hahn, Vighnesh Birodkar, Jonathan Huang, Irfan Essa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏