arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2504.19918 2025-04-29 cs.CV cs.AI 81%

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI

Hugo Georgenthum, Cristian Cosentino, Fabrizio Marozzo, Pietro Liò

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17213 2025-04-29 cs.CV cs.AI 81%

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding

Shiwen Cao, Zhaoxing Zhang, Junming Jiao, Juyi Qiao, Guowen Song, Rong Shen, Xiangbing Meng

机构 * Li Auto Inc.(利亚汽车公司) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14432 2025-04-22 cs.CV cs.AI 81%

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task

Ahmad Khalil, Mahmoud Khalil, Alioune Ngom

机构 * University of Windsor(温莎大学)

专题命中 视频多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14429 2025-04-22 cs.CV cs.AI 81%

ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations

Ahmad Khalil, Mahmoud Khalil, Alioune Ngom

机构 * University of Windsor(温莎大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11493 2025-04-17 cs.RO cs.AI cs.CV 81%

Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning

Azizul Zahid, Jie Fan, Farong Wang, Ashton Dy, Sai Swaminathan, Fei Liu

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments ICRA'25 Workshop: Human-Centered Robot Learning in the Era of Big Data and Large Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11232 2025-04-16 cs.CV cs.MM 81%

Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset

Elisa Ancarani, Julie Tores, Lucile Sassatelli, Rémy Sun, Hui-Yin Wu, Frédéric Precioso

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 6 pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05030 2025-04-08 cs.CV cs.MM 81%

AsyReC: A Multimodal Graph-based Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification

Wang Tang, Fethiye Irmak Dogan, Linbo Qing, Hatice Gunes

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04550 2025-04-08 cs.CV cs.AI cs.LG cs.RO 81%

Advancing Egocentric Video Question Answering with Multimodal Large Language Models

Alkesh Patel, Vibhav Chitalia, Yinfei Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22233 2025-04-01 cs.CV cs.AI cs.IR 81%

ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising

Ashutosh Chaubey, Anoubhav Agarwaal, Sartaki Sinha Roy, Aayush Agrawal, Susmita Ghose

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22610 2025-03-31 cs.HC cs.AI cs.CL cs.CY cs.LG 81%

Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users

Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05421 2025-03-21 cs.CV cs.AI 81%

EPAM-Net: An Efficient Pose-driven Attention-guided Multimodal Network for Video Action Recognition

Ahmed Abdelkawy, Asem Ali, Aly Farag

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Neurocomputing, Volume 633, 7 June 2025, 129781

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07667 2025-03-21 cs.LG cs.AI cs.CV eess.SP 81%

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

Wei Dai, Peilin Chen, Malinda Lu, Daniel Li, Haowen Wei, Hejie Cui, Paul Pu Liang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13531 2025-03-19 cs.CV cs.AI cs.CY cs.LG 81%

Context-aware Multimodal AI Reveals Hidden Pathways in Five Centuries of Art Evolution

Jin Kim, Byunghwee Lee, Taekho You, Jinhyuk Yun

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 30 pages, 4 figures. Some example paintings are blurred to avoid potential copyright violations

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12077 2025-03-18 cs.CV cs.AI 81%

V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents

Zhengrong Yue, Shaobin Zhuang, Kunchang Li, Yanbo Ding, Yali Wang

专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00064 2025-03-10 cs.LG cs.AI cs.CV cs.RO 81%

M2Distill: Multi-Modal Distillation for Lifelong Imitation Learning

Kaushik Roy, Akila Dissanayake, Brendan Tidd, Peyman Moghadam

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments IEEE ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13130 2025-02-19 cs.CV cs.AI cs.HC cs.LG cs.RO 81%

Magma: A Foundation Model for Multimodal AI Agents

Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, Lars Liden, Jianfeng Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 29 pages, 16 figures, technical report from MSR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10674 2025-02-19 cs.CV cs.CL 81%

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Mohamed Fazli Imam, Chenyang Lyu, Alham Fikri Aji

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Our dataset can be found at \url{https://huggingface.co/datasets/fazliimam/temporal-vqa}

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19100 2025-02-18 cs.CV cs.AI 81%

VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

Lawrence Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, Kazuhito Koishida

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10145 2025-02-17 cs.CV cs.MM 81%

Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling

Xinyu Li, Marwa Mahmoud

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03857 2025-02-07 cs.LG cs.CL cs.CV 81%

MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition

Stefan Gerd Fritsch, Cennet Oguz, Vitor Fortes Rey, Lala Ray, Maximilian Kiefer-Emmanouilidis, Paul Lukowicz

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16786 2025-01-29 cs.CV cs.CL 81%

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding

Yun Li, Zhe Liu, Yajing Kong, Guangrui Li, Jiyuan Zhang, Chao Bian, Feng Liu, Lina Yao, Zhenbang Sun

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15438 2025-01-28 cs.CV cs.MM 81%

Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection

Han Wang, Rui Yang Tan, Roy Ka-Wei Lee

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments 10 pages, 4 figures, THE WEB CONFERENCE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09355 2025-01-17 cs.AI cs.CV cs.ET cs.MA 81%

YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks

Saptarashmi Bandyopadhyay, Vikas Bahirwani, Lavisha Aggarwal, Bhanu Guda, Lin Li, Andrea Colaco

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01960 2025-01-07 cs.CV cs.AI cs.GR cs.LG 81%

GAF-FusionNet: Multimodal ECG Analysis via Gramian Angular Fields and Split Attention

Jiahao Qin, Feng Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 1 figure, accepted by ICONIP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11621 2024-12-17 cs.CV cs.MM 81%

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar, Mike Zheng Shou, Qianli Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted for The 39th Annual AAAI Conference on Artificial Intelligence 2025 in Main Track, 19 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10360 2024-12-16 cs.CV cs.AI 81%

Apollo: An Exploration of Video Understanding in Large Multimodal Models

Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, Xide Xia

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments https://apollo-lmms.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18938 2024-12-04 cs.CV cs.AI 81%

From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Heqing Zou, Tianze Luo, Guiyang Xie, Victor, Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13053 2024-11-21 cs.CV cs.AI cs.LG 81%

MEGL: Multimodal Explanation-Guided Learning

Yifei Zhang, Tianxu Jiang, Bo Pan, Jingyu Wang, Guangji Bai, Liang Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16620 2024-11-13 cs.CV cs.CL 81%

OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer

Lu Zhang, Tiancheng Zhao, Heting Ying, Yibo Ma, Kyusong Lee

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05767 2024-10-28 cs.AI cs.CL cs.IR 81%

A Survey of Knowledge Graph Reasoning on Graph Types: Static, Dynamic, and Multimodal

Ke Liang, Lingyuan Meng, Meng Liu, Yue Liu, Wenxuan Tu, Siwei Wang, Sihang Zhou, Xinwang Liu, Fuchun Sun

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏