arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4726 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4726 篇

2103.10699 2021-11-09 cs.CV 74%

MDMMT: Multidomain Multimodal Transformer for Video Retrieval

Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov, Aleksandr Petiushko

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Journal ref CVPR Workshops 2021: 3354-3363

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10211 2021-10-28 cs.CV 74%

Space-Time Crop & Attend: Improving Cross-modal Video Representation Learning

Mandela Patrick, Yuki M. Asano, Bernie Huang, Ishan Misra, Florian Metze, Joao Henriques, Andrea Vedaldi

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted to ICCV 2021. Code at https://github.com/facebookresearch/GDT

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13992 2021-10-28 cs.CV cs.LG 74%

Leveraging Local Temporal Information for Multimodal Scene Classification

Saurabh Sahu, Palash Goyal

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.14633 2021-05-03 physics.soc-ph cs.AI cs.LG cs.SI 74%

Modelling Urban Dynamics with Multi-Modal Graph Convolutional Networks

Krittika D'Silva, Jordan Cambe, Anastasios Noulas, Cecilia Mascolo, Adam Waksman

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 10 pages, 4 figures. arXiv admin note: substantial text overlap with arXiv:2104.13981

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06838 2021-02-18 cs.RO cs.CV 74%

Unified Multi-Modal Landmark Tracking for Tightly Coupled Lidar-Visual-Inertial Odometry

David Wisth, Marco Camurri, Sandipan Das, Maurice Fallon

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Video: https://youtu.be/MjXYAHurWe8

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14462 2020-12-18 cs.LG astro-ph.IM cs.CV eess.IV eess.SP 74%

Deep Probabilistic Imaging: Uncertainty Quantification and Multi-modal Solution Characterization for Computational Imaging

He Sun, Katherine L. Bouman

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments This paper has been accepted to AAAI 2021. Keywords: Computational Imaging, Normalizing Flow, Uncertainty Quantification, Interferometry, MRI

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.02568 2020-09-08 cs.CV 74%

Multimodal Memorability: Modeling Effects of Semantics and Decay on Video Memorability

Anelise Newman, Camilo Fosco, Vincent Casser, Allen Lee, Barry McNamara, Aude Oliva

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments European Conference on Computer Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.03715 2020-08-21 eess.SP cs.DC cs.MM 74%

A Modular Approach for Synchronized Wireless Multimodal Multisensor Data Acquisition in Highly Dynamic Social Settings

Chirag Raman, Stephanie Tan, Hayley Hung

专题命中 视频多模态 :multimodal(title);分类 cs.MM

Comments 9 pages, 8 figures, Proceedings of the 28th ACM International Conference on Multimedia (MM '20), October 12--16, 2020, Seattle, WA, USA. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07544 2020-04-17 cs.CV eess.IV 74%

Multimodal and multiview distillation for real-time player detection on a football field

Anthony Cioppa, Adrien Deliège, Noor Ul Huda, Rikke Gade, Marc Van Droogenbroeck, Thomas B. Moeslund

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments Accepted for the CVSports workshop of CVPR 2020 ; 8 pages + references

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13594 2020-03-31 cs.CV 74%

Speech2Action: Cross-modal Supervision for Action Recognition

Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted to CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08830 2020-01-27 cs.SD cs.LG eess.AS eess.SP 74%

Scattering Features for Multimodal Gait Recognition

Srđan Kitić, Gilles Puy, Patrick Pérez, Philippe Gilberton

专题命中 视频多模态 :multimodal(title);分类 eess.AS

Comments Published at IEEE GlobalSIP 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01886 2019-02-07 cs.AI 74%

Situational Grounding within Multimodal Simulations

James Pustejovsky, Nikhil Krishnaswamy

专题命中 视频多模态 :multimodal(title);分类 cs.AI

Comments AAAI-19 Workshop on Games and Simulations for Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.00303 2018-12-04 cs.CV 74%

Multi-modal Capsule Routing for Actor and Action Video Segmentation Conditioned on Natural Language Queries

Bruce McIntosh, Kevin Duarte, Yogesh S Rawat, Mubarak Shah

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.07212 2018-10-18 cs.CV 74%

Cross-Modal and Hierarchical Modeling of Video and Text

Bowen Zhang, Hexiang Hu, Fei Sha

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by ECCV 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.00599 2018-10-02 cs.CV 74%

Unsupervised Trajectory Segmentation and Promoting of Multi-Modal Surgical Demonstrations

Zhenzhou Shao, Hongfa Zhao, Jiexin Xie, Ying Qu, Yong Guan, Jindong Tan

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06774 2018-04-19 cs.AI cs.NE cs.RO 74%

Encoding Longer-term Contextual Multi-modal Information in a Predictive Coding Model

Junpei Zhong, Tetsuya Ogata, Angelo Cangelosi

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments Submitted to ICDL/EpiRob 2018 (8th Joint IEEE International Conference on Development and Learning and on Epigenetic Robotics )

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.10330 2017-10-31 cs.CV 74%

Multi-modal Aggregation for Video Classification

Chen Chen, Xiaowei Zhao, Yang Liu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.07728 2017-10-30 cs.SI cs.CL cs.CY physics.soc-ph 74%

A Computational Framework for Multi-Modal Social Action Identification

Jason Anastasopoulos, Jake Ryland Williams

专题命中 视频多模态 :multi-modal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.01077 2017-09-06 cs.CV 74%

A Nonparametric Model for Multimodal Collaborative Activities Summarization

Guy Rosman, John W. Fisher, Daniela Rus

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.09430 2017-06-30 cs.CV 74%

Summarization of ICU Patient Motion from Multimodal Multiview Videos

Carlos Torres, Kenneth Rose, Jeffrey C. Fried, B. S. Manjunath

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments Total of 17 pages: 1-12 (body), 13-16 (result figures), and 17(references) Number of figures: 21

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.07475 2017-03-29 cs.CV 74%

PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding

Chunhui Liu, Yueyu Hu, Yanghao Li, Sijie Song, Jiaying Liu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.01376 2016-11-11 cs.CV 74%

Recognizing and Presenting the Storytelling Video Structure with Deep Multimodal Networks

Lorenzo Baraldi, Costantino Grana, Rita Cucchiara

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1503.01800 2015-03-31 cs.LG cs.CV 74%

EmoNets: Multimodal deep learning approaches for emotion recognition in video

Samira Ebrahimi Kahou, Xavier Bouthillier, Pascal Lamblin, Caglar Gulcehre, Vincent Michalski, Kishore Konda, Sébastien Jean, Pierre Froumenty, Yann Dauphin, Nicolas Boulanger-Lewandowski, Raul Chandias Ferrari, Mehdi Mirza, David Warde-Farley, Aaron Courville, Pascal Vincent, Roland Memisevic, Christopher Pal, Yoshua Bengio

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1008.5163 2010-09-01 cs.AI 74%

Learning Multi-modal Similarity

Brian McFee, Gert Lanckriet

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12826 2026-08-12 cs.CV cs.AI 版本更新 73%

DIMOS: Disentangling Instance-level Moving Object Segmentation

DIMOS: 解耦实例级运动目标分割

Hongxiang Huang, Hongwei Ren, Xiaopeng Lin, Yulong Huang, Zeke Xie, Bojun Cheng

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出双解耦特征提取框架分离图像与事件模态的外观和运动信息,并通过多粒度跨模态对齐实现有效融合,在运动实例分割任务中尤其对快速运动和低光下的小目标取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07417 2026-08-10 cs.CV cs.AI 新提交 73%

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

我在视频中寻找你:面向以人为中心的视频推理的身份条件查询

Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang

机构 * Beijing Jiaotong University(北京交通大学) HUJING Digital Media & Entertainment Group(汇晶数字媒体与娱乐集团) MAIS Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 该研究提出了身份条件查询(ICQ)任务,构建了ISYV基准、训练集与框架,实验显示主流多模态大语言模型在该任务上表现不佳,ISYV模型性能优于强基线且接近闭源模型。

Comments Accepted to ACM Multimedia 2026 (MM '26). 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28678 2026-08-03 cs.AI cs.CV 新提交 73%

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

ViSAGE:为长视频理解构建自修正记忆

Xinkui Zhao, Enbo Chen, Yifan Zhang, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu

机构 * Zhejiang University(浙江大学) Xidian University(西安电子科技大学)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 ViSAGE是一种多模态智能体记忆框架,通过跨模态绑定、双向记忆精调及多智能体交叉验证解决长视频理解中的实体混淆等问题,较最强基线准确率提升5.9%。

Comments Accept by ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11719 2026-07-30 cs.CV cs.AI 版本更新 73%

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning

Ouroboros-Spatial:闭环数据-模型循环的空间推理

Enhan Zhao, Wei Wu, Yuanrui Zhang, Xueliang Zhao, Di He

机构 * Peking University(北京大学) Ant International(蚂蚁国际) The University of Hong Kong(香港大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 提出Ouroboros-Spatial自演化框架,通过提议器与求解器闭环交互,动态生成与模型能力匹配的训练样本,在六个空间推理基准上以十分之一数据量显著提升Qwen3-VL性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23327 2026-07-07 cs.CV cs.AI 新提交 73%

VideoAgent: All-in-One Framework for Video Understanding and Editing

VideoAgent: 视频理解与编辑的全能框架

Hengji Zhou, Lingxuan Huang, Jian Wang, Bing Zhou, Si Wu, Lianghao Xia, Chao Huang

机构 * South China University of Technology(华南理工大学) The University of Hong Kong(香港大学) Snap Inc(Snap公司) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出VideoAgent,一种基于多智能体编排的全能框架,通过自动镜头创建和30多个专业编辑智能体,解决现有系统在长视频理解和多样化编辑操作上的局限,在VideoEdit基准上实现87-95%编排成功率并降低60% API成本。

Comments Preprint. Code available at https://github.com/HKUDS/VideoAgent

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10359 2026-06-30 cs.CV cs.AI 73%

Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task

增强工具的时空推理用于视频问答任务的优化

Sunqi Fan, Jiashuo Cui, Meng-Hao Guo, Shuojin Yang

机构 * BNRist, Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系,北京信息科学与技术国家研究中心)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出STAR框架和视频工具包,提升多模态大语言模型的时空推理能力,在VideoMME和LongVideoBench上分别提升8.2%和4.6%。

Comments Accepted by NeurIPS 2025 main track

详情

展开后加载摘要…

URL PDF HTML 收藏