arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

1805.11372 2018-05-30 cs.CV 79%

"How to rate a video game?" - A prediction system for video games based on multimodal information

Vishal Batchu, Varshit Battu, Murali Krishna Reddy, Radhika Mamidi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICPRAI-18

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06057 2018-04-18 cs.MM 79%

Multimodal Co-Training for Selecting Good Examples from Webly Labeled Video

Ryota Hinami, Junwei Liang, Shin'ichi Satoh, Alexander Hauptmann

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.05939 2018-04-04 cs.CV q-bio.NC 79%

AJILE Movement Prediction: Multimodal Deep Learning for Natural Human Neural Recordings and Video

Nancy Xin Ru Wang, Ali Farhadi, Rajesh Rao, Bingni Brunton

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Thirty-Second AAAI Conference On Artificial Intelligence (2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.08513 2018-02-08 cs.CL q-bio.NC 79%

Interactive Natural Language Acquisition in a Multi-modal Recurrent Neural Architecture

Stefan Heinrich, Stefan Wermter

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Received 25 June 2016; Accepted 1 February 2017

Journal ref Connection Science, vol 30, No 1, pp. 99-133, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.07721 2018-01-15 cs.CV 79%

An Order Preserving Bilinear Model for Person Detection in Multi-Modal Data

Oytun Ulutan, Benjamin S. Riggan, Nasser M. Nasrabadi, B. S. Manjunath

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.03321 2017-11-28 stat.ML cs.CV q-bio.NC 79%

A deep learning architecture for temporal sleep stage classification using multivariate and multimodal time series

Stanislas Chambon, Mathieu Galtier, Pierrick Arnal, Gilles Wainrib, Alexandre Gramfort

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.08569 2017-11-27 cs.CV 79%

Geometric Cross-Modal Comparison of Heterogeneous Sensor Data

Christopher J. Tralie, Abraham Smith, Nathan Borggren, Jay Hineman, Paul Bendich, Peter Zulch, John Harer

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 10 pages, 13 figures, Proceedings of IEEE Aeroconf 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.05219 2017-10-17 cs.LG cs.AI 79%

Mental Sampling in Multimodal Representations

Jian-Qiao Zhu, Adam N. Sanborn, Nick Chater

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.05461 2017-07-11 cs.CV 79%

Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text

Zhe Wang, Kingsley Kuan, Mathieu Ravaut, Gaurav Manek, Sibo Song, Yuan Fang, Seokhwan Kim, Nancy Chen, Luis Fernando D'Haro, Luu Anh Tuan, Hongyuan Zhu, Zeng Zeng, Ngai Man Cheung, Georgios Piliouras, Jie Lin, Vijay Chandrasekhar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, Accepted to CVPR'17 Workshop on YouTube-8M Large-Scale Video Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.05103 2017-05-16 cs.MM cs.IR 79%

Generative Adversarial Networks for Multimodal Representation Learning in Video Hyperlinking

Vedran Vukotic, Christian Raymond, Guillaume Gravier

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments 4 pages, 1 figure, 2 tables, published at ACM International Conference in Multimedia Retrieval (ICMR) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.05244 2017-04-14 cs.CL cs.IR 79%

Select-Additive Learning: Improving Generalization in Multimodal Sentiment Analysis

Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, Eric P. Xing

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Supplementary files at: http://www.cs.cmu.edu/~haohanw/document/sal_supp.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.01638 2017-02-07 cs.CV 79%

Concurrent Activity Recognition with Multimodal CNN-LSTM Structure

Xinyu Li, Yanyi Zhang, Jianyu Zhang, Shuhong Chen, Ivan Marsic, Richard A. Farneth, Randall S. Burd

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, 12 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.07120 2016-12-28 cs.CV 79%

Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos

Amir Shahroudy, Tian-Tsong Ng, Yihong Gong, Gang Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.7006 2016-11-15 cs.CV cs.LG cs.NE 79%

Multi-modal Sensor Registration for Vehicle Perception via Deep Neural Networks

Michael Giering, Vivek Venugopalan, Kishore Reddy

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, double column, IEEE format, accepted at IEEE HPEC 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.03112 2016-10-12 cs.CL 79%

Leveraging Recurrent Neural Networks for Multimodal Recognition of Social Norm Violation in Dialog

Tiancheng Zhao, Ran Zhao, Zhao Meng, Justine Cassell

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Submitted to NIPS Workshop. arXiv admin note: text overlap with arXiv:1608.02977 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0611138 2016-08-16 cs.AI 79%

Functional Brain Imaging with Multi-Objective Multi-Modal Evolutionary Optimization

Vojtech Krmicek, Michèle Sebag

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Journal ref Dans PPSN'06, 4193 (2006) 382-391

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.04780 2016-07-19 cs.CV cs.LG 79%

Exploiting Multi-modal Curriculum in Noisy Web Data for Large-scale Concept Learning

Junwei Liang, Lu Jiang, Deyu Meng, Alexander Hauptmann

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.02652 2016-07-12 cs.HC cs.CV 79%

Multimodal Affect Recognition using Kinect

Amol Patwardhan, Gerald Knapp

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages, 2 tables, 1 figure, Peer reviewed in ACM TIST

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.03237 2016-06-13 cs.CV 79%

Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-related Applications

Ciprian Corneanu, Marc Oliu, Jeffrey F. Cohn, Sergio Escalera

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.00022 2016-01-06 cs.CV 79%

Event Specific Multimodal Pattern Mining with Image-Caption Pairs

Hongzhi Li, Joseph G. Ellis, Shih-Fu Chang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.08761 2015-08-03 cs.CV 79%

Multimodal Multipart Learning for Action Recognition in Depth Videos

Amir Shahroudy, Gang Wang, Tian-Tsong Ng, Qingxiong Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.07892 2024-06-10 cs.LG math.DS nlin.CD 79%

Integrating Multimodal Data for Joint Generative Modeling of Complex Dynamics

Manuel Brenner, Florian Hess, Georgia Koppe, Daniel Durstewitz

专题命中 视频多模态 :multimodal(title,abstract)

Comments ICML 2024. Previously published as a workshop paper for the AAAI 2023 Workshop MLmDS as "Multimodal Teacher Forcing for Reconstructing Nonlinear Dynamical Systems"

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19528 2026-07-23 cs.CV cs.AI 新提交 79%

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景

Heesang Han, A. Lynn Abbott, Abhijit Sarkar

机构 * Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系) Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所) Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文针对自动驾驶中多模态大语言模型,提出D3VL框架,整合2D和3D时间序列数据,回答交通场景相关问题,在KITTI问答数据集上性能提升11%,并引入Waymo QA数据集扩展以评估模型在多样驾驶条件下处理3D和时间序列数据的能力。

Comments Accepted to IEEE IV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29136 2026-06-30 cs.CV cs.AI 79%

CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection

CMTFormer:融合Transformer与层次化信息交互的RGB-事件目标检测

Yu Li, Yuenan Hou, Yingmei Wei, Jiangming Chen, Yanming Guo

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出CMTFormer,通过浅层对齐、中层增强和深层可学习融合的层次化交互机制,有效融合RGB帧与事件流,在DSEC-Detection和PKU-DAVIS-SOD基准上超越现有方法。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22870 2026-06-23 cs.CV cs.AI 新提交 79%

VideoLatent: Video-Language Learning via Latent Self-Forcing

VideoLatent: 通过潜在自强制进行视频-语言学习

Zi-Yuan Hu, Zicong Tang, Shijia Huang, Yanyang Li, Michael R. Lyu, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学) Weitu AI(微图AI)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出VideoLatent,一种通过潜在自强制训练范式(包括潜在对齐和潜在多样性目标)进行视觉潜在推理的多模态大语言模型,仅依赖标准视频-问答三元组,在14个基准上优于现有模型,并大幅降低训练/推理开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12813 2026-06-16 cs.CV cs.MM 版本更新 79%

DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment

DPC-VQA: 解耦质量感知与残差校准用于视频质量评估

Xinyue Li, Shubo Xu, Zhichao Zhang, Zhaolin Cai, Yitong Chen, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Baidu Inc.(百度公司) Xinjiang University(新疆大学)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 提出DPC-VQA框架,通过冻结多模态大模型提供感知先验,轻量校准分支预测残差修正,实现低训练成本下视频质量评估,在UGC和AIGC基准上取得竞争性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09151 2026-06-09 cs.CV cs.AI cs.LG 版本更新 79%

Video Understanding by Design: How Datasets Shape Video Models

通过设计理解视频:数据集如何塑造视频模型

Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao

机构 * School of Engineering and Built Environment, Electrical and Electronic Engineering, Griffith University(工程与建筑环境学院,电气与电子工程学院,格里菲斯大学) School of Computer Science and Engineering, University of New South Wales(计算机科学与工程学院,新南威尔士大学)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

AI总结 本文从数据集视角出发,提出统一框架连接数据集结构、归纳偏差与架构设计,分析数据集特性如何驱动视频理解架构创新,并讨论不同数据体制下的表征偏差。

Comments Research report

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00101 2026-06-02 cs.CV cs.AI 79%

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

CoCoVideo: 基于商业模型的高质量对比基准用于AI生成视频检测

Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng

机构 * School of Informatics, Xiamen University(厦门大学信息学院) China Academy of Information and Communications Technology(中国信息通信技术研究院) AI Transcend Pte. Ltd.(AI Transcend有限公司)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 针对现有数据集依赖低质量开源模型且商业样本带水印的问题,提出包含13个商业生成器的CoCoVideo-26K对比数据集,并设计结合对比学习与置信门控多模态大语言模型的CoCoDetect检测框架,实现高保真AI生成视频的鲁棒检测。

Comments Accepected by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16953 2026-05-26 cs.AI cs.CL 79%

How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study

人类如何处理AI生成的幻觉内容:一项神经影像学研究

Shuqi Zhu, Yi Zhong, Ziyi Ye, Bangde Du, Yujia Zhou, Qingyao Ai, Yiqun Liu

机构 * Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系) Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China(复旦大学可信具身人工智能研究院)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 通过EEG实验,研究人类在处理多模态大语言模型生成的幻觉与非幻觉内容时的神经动力学差异,揭示误判的幻觉内容未能触发标准神经认知事实验证通路。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00891 2026-05-05 cs.CV cs.AI 79%

X2SAM: Any Segmentation in Images and Videos

X2SAM:图像和视频中的任意分割

Hao Wang, Limeng Qiao, Chi Zhang, Lin Ma, Guanglu Wan, Xiangyuan Lan, Xiaodan Liang

机构 * Sun Yat-Sen University(中山大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 X2SAM是一种统一的分割多模态大语言模型,能够将图像分割能力扩展到视频,结合LLM和Mask Memory模块,支持通用、开放词汇、指称、推理、基于视觉的对话生成及交互式分割。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏