"How to rate a video game?" - A prediction system for video games based on multimodal information
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments ICPRAI-18
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments ICPRAI-18
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Journal ref Thirty-Second AAAI Conference On Artificial Intelligence (2018)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL
Comments Received 25 June 2016; Accepted 1 February 2017
Journal ref Connection Science, vol 30, No 1, pp. 99-133, 2017
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV
Comments 10 pages, 13 figures, Proceedings of IEEE Aeroconf 2017
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments 8 pages, Accepted to CVPR'17 Workshop on YouTube-8M Large-Scale Video Understanding
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM
Comments 4 pages, 1 figure, 2 tables, published at ACM International Conference in Multimedia Retrieval (ICMR) 2017
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL
Comments Supplementary files at: http://www.cs.cmu.edu/~haohanw/document/sal_supp.pdf
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 14 pages, 12 figures, under review
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments 7 pages, double column, IEEE format, accepted at IEEE HPEC 2015
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL
Comments Submitted to NIPS Workshop. arXiv admin note: text overlap with arXiv:1608.02977 by other authors
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
Journal ref Dans PPSN'06, 4193 (2006) 382-391
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 9 pages, 2 tables, 1 figure, Peer reviewed in ACM TIST
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract)
Comments ICML 2024. Previously published as a workshop paper for the AAAI 2023 Workshop MLmDS as "Multimodal Teacher Forcing for Reconstructing Nonlinear Dynamical Systems"
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
机构 * Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系) ; Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所) ; Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文针对自动驾驶中多模态大语言模型,提出D3VL框架,整合2D和3D时间序列数据,回答交通场景相关问题,在KITTI问答数据集上性能提升11%,并引入Waymo QA数据集扩展以评估模型在多样驾驶条件下处理3D和时间序列数据的能力。
Comments Accepted to IEEE IV 2026
CMTFormer:融合Transformer与层次化信息交互的RGB-事件目标检测
专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 提出CMTFormer,通过浅层对齐、中层增强和深层可学习融合的层次化交互机制,有效融合RGB帧与事件流,在DSEC-Detection和PKU-DAVIS-SOD基准上超越现有方法。
Comments 15 pages
VideoLatent: 通过潜在自强制进行视频-语言学习
机构 * The Chinese University of Hong Kong(香港中文大学) ; Weitu AI(微图AI)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 提出VideoLatent,一种通过潜在自强制训练范式(包括潜在对齐和潜在多样性目标)进行视觉潜在推理的多模态大语言模型,仅依赖标准视频-问答三元组,在14个基准上优于现有模型,并大幅降低训练/推理开销。
DPC-VQA: 解耦质量感知与残差校准用于视频质量评估
机构 * Shanghai Jiao Tong University(上海交通大学) ; Baidu Inc.(百度公司) ; Xinjiang University(新疆大学)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM
AI总结 提出DPC-VQA框架,通过冻结多模态大模型提供感知先验,轻量校准分支预测残差修正,实现低训练成本下视频质量评估,在UGC和AIGC基准上取得竞争性能。
通过设计理解视频:数据集如何塑造视频模型
机构 * School of Engineering and Built Environment, Electrical and Electronic Engineering, Griffith University(工程与建筑环境学院,电气与电子工程学院,格里菲斯大学) ; School of Computer Science and Engineering, University of New South Wales(计算机科学与工程学院,新南威尔士大学)
专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
AI总结 本文从数据集视角出发,提出统一框架连接数据集结构、归纳偏差与架构设计,分析数据集特性如何驱动视频理解架构创新,并讨论不同数据体制下的表征偏差。
Comments Research report
CoCoVideo: 基于商业模型的高质量对比基准用于AI生成视频检测
机构 * School of Informatics, Xiamen University(厦门大学信息学院) ; China Academy of Information and Communications Technology(中国信息通信技术研究院) ; AI Transcend Pte. Ltd.(AI Transcend有限公司)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 针对现有数据集依赖低质量开源模型且商业样本带水印的问题,提出包含13个商业生成器的CoCoVideo-26K对比数据集,并设计结合对比学习与置信门控多模态大语言模型的CoCoDetect检测框架,实现高保真AI生成视频的鲁棒检测。
Comments Accepected by CVPR 2026
人类如何处理AI生成的幻觉内容:一项神经影像学研究
机构 * Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系) ; Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China(复旦大学可信具身人工智能研究院)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CL、cs.AI
AI总结 通过EEG实验,研究人类在处理多模态大语言模型生成的幻觉与非幻觉内容时的神经动力学差异,揭示误判的幻觉内容未能触发标准神经认知事实验证通路。
X2SAM:图像和视频中的任意分割
机构 * Sun Yat-Sen University(中山大学) ; Peng Cheng Laboratory(鹏城实验室)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 X2SAM是一种统一的分割多模态大语言模型,能够将图像分割能力扩展到视频,结合LLM和Mask Memory模块,支持通用、开放词汇、指称、推理、基于视觉的对话生成及交互式分割。
Comments Technical Report