arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2103.12541 2022-09-30 cs.MM cs.AI cs.CL cs.CR cs.CY cs.LG cs.SI 82%

A Survey on Multimodal Disinformation Detection

Firoj Alam, Stefano Cresci, Tanmoy Chakraborty, Fabrizio Silvestri, Dimiter Dimitrov, Giovanni Da San Martino, Shaden Shaar, Hamed Firooz, Preslav Nakov

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted at COLING-2022, disinformation, misinformation, factuality, harmfulness, fake news, propaganda, multimodality, text, images, videos, network structure, temporality

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13537 2022-09-28 cs.NI cs.DB 82%

Sensing Multi-modal Mobility Patterns: A Case Study of Helsinki using Bluetooth Beacons and a Mobile Application

Zhiren Huang, Alonso Espinosa Mireles de Villafranca, Charalampos Sipetas

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04446 2022-08-19 cs.CV cs.CL cs.SD eess.AS 82%

Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, Hilde Kuehne

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、eess.AS

Comments CVPR2022. The final published version of the proceedings will be available on IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05814 2022-08-12 cs.CV cs.AI cs.MM 82%

Seeing your sleep stage: cross-modal distillation from EEG to infrared video

Jianan Han, Shaoxing Zhang, Aidong Men, Yang Liu, Ziming Yao, Yan Yan, Qingchao Chen

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments We have submitted this paper to an academic journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07898 2022-06-17 cs.AI cs.CL cs.CV cs.LG 82%

Multimodal Dialogue State Tracking

Hung Le, Nancy F. Chen, Steven C. H. Hoi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at NAACL 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08264 2022-05-11 cs.CV cs.AI cs.CL cs.HC 82%

End-to-end Generative Pretraining for Multimodal Video Captioning

Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia Schmid

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03734 2022-04-11 cs.CV cs.CL cs.MM 82%

MHMS: Multimodal Hierarchical Multimedia Summarization

Jielin Qiu, Jiacheng Zhu, Mengdi Xu, Franck Dernoncourt, Trung Bui, Zhaowen Wang, Bo Li, Ding Zhao, Hailin Jin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03763 2022-02-03 cs.LG 82%

Creating Multimodal Interactive Agents with Imitation and Self-Supervised Learning

DeepMind Interactive Agents Team, Josh Abramson, Arun Ahuja, Arthur Brussee, Federico Carnevale, Mary Cassin, Felix Fischer, Petko Georgiev, Alex Goldin, Mansi Gupta, Tim Harley, Felix Hill, Peter C Humphreys, Alden Hung, Jessica Landon, Timothy Lillicrap, Hamza Merzic, Alistair Muldal, Adam Santoro, Guy Scully, Tamara von Glehn, Greg Wayne, Nathaniel Wong, Chen Yan, Rui Zhu

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11178 2021-12-08 cs.CV cs.AI cs.LG cs.MM eess.IV 82%

VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text

Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, Boqing Gong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Published in the 35th Conference on Neural Information Processing Systems (NeurIPS 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11985 2021-04-30 cs.CL cs.CV cs.LG cs.MM 82%

MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences

Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, Louis-Philippe Morency

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00848 2021-01-13 cs.CR 82%

Towards Cross-Modal Forgery Detection and Localization on Live Surveillance Videos

Yong Huang, Xiang Li, Wei Wang, Tao Jiang, Qian Zhang

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract)

Comments Accepted by IEEE International Conference on Computer Communications 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06353 2020-09-16 cs.CV cs.CL cs.LG eess.AS eess.IV 82%

UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, Ming Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.03546 2020-08-11 cs.CV cs.AI cs.LG cs.MM eess.IV 82%

Online Multi-modal Person Search in Videos

Jiangyue Xia, Anyi Rao, Qingqiu Huang, Linning Xu, Jiangtao Wen, Dahua Lin

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments ECCV2020. Project page: http://movienet.site/projects/eccv20onlineperson.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.07758 2020-05-07 cs.CV cs.CL cs.LG cs.SD eess.AS eess.IV 82%

Multi-modal Dense Video Captioning

Vladimir Iashin, Esa Rahtu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、eess.AS

Comments To appear in the proceedings of CVPR Workshops 2020; Code: https://github.com/v-iashin/MDVC Project Page: https://v-iashin.github.io/mdvc

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04917 2020-04-13 cs.LG cs.AI cs.CL cs.CV 82%

Multimodal Categorization of Crisis Events in Social Media

Mahdi Abavisani, Liwei Wu, Shengli Hu, Joel Tetreault, Alejandro Jaimes

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Conference on Computer Vision and Pattern Recognition (CVPR 2020)

Journal ref Conference on Computer Vision and Pattern Recognition (CVPR 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.01060 2018-05-04 cs.AI cs.CL cs.CV 82%

Multimodal Emotion Recognition for One-Minute-Gradual Emotion Challenge

Ziqi Zheng, Chenjie Cao, Xingwei Chen, Guoqiang Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.01455 2017-12-06 cs.AI cs.CL cs.CV 82%

Multimodal Storytelling via Generative Adversarial Imitation Learning

Zhiqian Chen, Xuchao Zhang, Arnold P. Boedihardjo, Jing Dai, Chang-Tien Lu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments IJCAI 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.03126 2017-03-13 cs.CV cs.CL cs.MM 82%

Attention-Based Multimodal Fusion for Video Description

Chiori Hori, Takaaki Hori, Teng-Yok Lee, Kazuhiro Sumi, John R. Hershey, Tim K. Marks

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Resubmitted to the rebuttal for CVPR 2017 for review, 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11147 2026-04-03 cs.MM cs.CV cs.LG 82%

Catalogue Grounded Multimodal Attribution for Museum Video under Resource and Regulatory Constraints

博物馆视频下的资源与监管约束下的目录引导多模态归因

Minsak Nanang, Adrian Hilton, Armin Mustafa

机构 * University of Surrey(萨里大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出一种目录引导的多模态归因方法,用于博物馆视频内容的自动化元数据生成,以提高档案馆发现性并满足资源和监管要求。

Comments Demo video url: https://jn00767.pages.surrey.ac.uk/catalogue-grounded-multimodal-attribution-for-museum-video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21815 2026-02-02 cs.CY cs.AI cs.CL cs.SI 82%

Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US

道德愤怒影响承诺:韩国和美国YouTube上的多模态道德情感

Seongchan Park, Jaehong Kim, Hyeonseung Kim, Heejin Bin, Sue Moon, Wonjae Lee

机构 * KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究通过多模态道德情感分类器分析YouTube上道德愤怒对用户参与度的影响,发现其在不同文化中均能提升观看和评论等互动行为。

Comments Accepted at The Web Conference 2026. We release Korean and English multimodal moral emotion classifiers

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12454 2025-08-13 cs.CV cs.AI cs.HC cs.LG 82%

Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches

He Zhang, Xinyi Fu

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, accepted to MRAC'25: 3rd International Workshop on Multimodal and Responsible Affective Computing (ACM-MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08657 2023-06-16 cs.CV cs.LG cs.MM 82%

EMERSK -- Explainable Multimodal Emotion Recognition with Situational Knowledge

Mijanur Palash, Bharat Bhargava

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM;multi-modal(comments)

Comments Emotion Recognition, Deep Learning, Multi-modal, Convolutional neural network (CNN), LSTM, Situational-Knowledge, Novelty

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.13362 2020-05-29 cs.CL cs.CV 82%

A Multi-modal Approach to Fine-grained Opinion Mining on Video Reviews

Edison Marrese-Taylor, Cristian Rodriguez-Opazo, Jorge A. Balazs, Stephen Gould, Yutaka Matsuo

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL;multimodal(comments)

Comments Second Grand Challenge and Workshop on Multimodal Language ACL 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05780 2026-08-07 cs.CV 新提交 81%

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

用于高效长视频理解的证据驱动型动态视觉选择器

Bo Zhang, Wenxin Wang, Feng Chen, Zhihao Zhang, Zixuan Wang, Changsheng Li, Yinjie Lei

专题命中 视频多模态 :MLLM(summary_cn,abstract);分类 cs.CV

AI总结 本文提出基于目标MLLM内部注意力证据的动态视觉选择框架EviSelect,通过GRPO优化的随机策略实现高效长视频理解,在三个基准上性能更优,视觉token减少约50%、端到端加速3.9倍。

Comments Project Page: https://zhangbo135.github.io/EviSelect/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03890 2026-06-05 cs.CV 81%

4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

4DPC$^2$hat: 面向动态点云理解的失败感知自举学习

Xindan Zhang, Weilong Yan, Yufei Shi, Xuerui Qiu, Tao He, Ying Li, Ming Li, Hehe Fan

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出首个针对动态点云理解的多模态大语言模型4DPC$^2$hat,通过构建大规模跨模态数据集4DPC$^2$hat-200K和引入Mamba增强的时间推理模块及失败感知自举学习策略,显著提升了动作理解与时间推理能力。

Comments Accept by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.13369 2020-10-27 cs.CV cs.HC cs.LG 81%

Introducing Representations of Facial Affect in Automated Multimodal Deception Detection

Leena Mathur, Maja J Matarić

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, Accepted at ACM International Conference on Multimodal Interaction (ICMI), October 2020

Journal ref Proceedings of the 2020 International Conference on Multimodal Interaction (ICMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13558 2026-08-14 cs.AI cs.CL 新提交 81%

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

OmniScientist:一种全模态全学科的AI科学家

Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) University of Oxford(牛津大学)

专题命中 视频多模态 :omni-modal(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究提出全模态AI科学家OmniScientist,通过端到端全模态流水线处理多学科异构原始证据,在36个真实案例中完成从原始数据到手稿的全流程,性能优于仅用预计算特征的系统,为通用AI科学家提供可行方案。

Comments 30 pages, 13 figures, 19 tables. Project page: https://omni-scientist.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03979 2026-08-05 cs.CV cs.AI 新提交 81%

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Video-DeepResearch:迈向新一代多模态深度研究智能体

Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 该研究提出Video-DR智能体框架,通过解耦感知-探索流水线与分阶段工具解锁解决现有多模态模型的模态偏差和知识泄漏问题,构建基准并在视频问答任务上取得SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01827 2026-08-04 cs.CV cs.AI 新提交 81%

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

DeepVoyager-VL:为长视野多模态智能体激励视觉在环搜索

Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang

机构 * PKU(北京大学) HKUST(GZ)(香港科技大学(广州)) NUDT(中国人民解放军国防科技大学) OUC(中国海洋大学) HITSZ(哈尔滨工业大学(深圳)) Huawei Cloud BU(华为云业务部)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对现有多模态深度搜索忽视视觉中间推理作用的局限,提出DeepVoyager-VL框架,构建多模态事件图合成数据,设计智能体框架实现主动视觉获取,经10个基准实验验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17696 2026-08-03 cs.CV cs.AI 81%

Hierarchical and Multimodal Data for Daily Activity Understanding

层级化和多模态数据用于日常活动理解

Ghazal Kaviani, Yavuz Yarici, Seulgi Kim, Mohit Prabhushankar, Ghassan AlRegib, Mashhour Solh, Ameya Patil

机构 * Georgia Institute of Technology(佐治亚理工学院) Amazon Lab126(亚马逊Lab126)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 DARai数据集通过多模态和层级注释,旨在理解现实环境中的人类活动,支持单模态和多模态传感器融合实验,揭示日常活动理解中的挑战。

Comments Accepted for publication in DMLR

详情

展开后加载摘要…

URL PDF HTML 收藏