GCM-Net: Graph-enhanced Cross-Modal Infusion with a Metaheuristic-Driven Network for Video Sentiment and Emotion Analysis
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 视频多模态 :MLLM(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Under review
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by WACV 2024; well-formatted PDF is in https://drive.google.com/file/d/1qvW52lamsvNGMCqPS7q8g8L4NaR_LlbR/view?usp=sharing. arXiv admin note: text overlap with arXiv:2401.04023
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
Comments Project webpage: https://drive-anywhere.github.io Explainer video: https://www.youtube.com/watch?v=4n-DJf8vXxo&feature=youtu.be
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted for publication at CVPR2022
专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments ACM Multimedia 2021
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted to WACV 2022
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV
Comments Accepted to ICLR 2021
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV
Comments To appear in CVPR 2017
LVSum:一个用于时间感知长视频摘要的基准测试
机构 * Apple(苹果公司)
专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出LVSum基准测试,用于评估长视频摘要中时间对齐的性能,通过引入新的评估指标揭示现有MLLM在时间理解上的系统性差距。
Comments 25 pages, 5 tables, 3 figures
从模态到命题:多模态智能的语言中心框架
机构 * NVIDIA(英伟达) ; University of Ottawa(渥太华大学)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 该研究提出多模态数据语言表示框架,将观察结果表示为原子命题,通过全局语义码本统一为共享词汇表,置于可解释空间,实现跨模态理解等,还在自动驾驶等数据上进行了展示。
MSAO:基于边缘-云协作的自适应模态稀疏性感知卸载框架用于高效多模态大语言模型推理
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract)
AI总结 本文提出MSAO框架,通过边缘-云协作和模态稀疏性分析,降低多模态大语言模型推理的延迟和资源消耗,提升吞吐量。
Comments 10 pages, 9 figures
DiFlowDubber:通过跨模态对齐与同步实现的离散流匹配自动化视频配音
机构 * FPT Software AI Center, Vietnam(FPT软件人工智能中心,越南) ; KAIST, South Korea(韩国科学技术院) ; University of Alabama at Birmingham, USA(阿拉巴马大学伯明翰分校)
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 本文提出DiFlowDubber框架,通过离散流匹配和两阶段训练策略,解决视频配音中内容准确性、表达语气、高质量音频和精确唇同步的问题。
Comments Accepted at CVPR 2026 Findings
ResAdapt:面向高效多模态推理的自适应分辨率
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 ResAdapt通过自适应输入分辨率框架,在保持高空间分辨率的同时提升多模态推理效率,尤其在压缩条件下显著提升性能。
Comments work in progress
通过自增强对比对齐缓解多模态大语言模型中的对象和动作幻觉
机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究所) ; NVIDIA
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 SANTA框架通过自增强对比对齐方法,有效缓解多模态大语言模型中的对象和动作幻觉问题。
Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. Project page: https://kpc0810.github.io/santa/
机构 * NVIDIA
专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Technical Report. Code: https://github.com/NVlabs/OmniVinci
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)
Comments Accepted by KDD 2025
机构 * TikTok(字节跳动)
专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract)
Comments Camera Ready for ACL 2025
机构 * New York University(纽约大学)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能国家重点实验室) ; School of Computer Science, Liupanshui Normal University(黎平师范学院计算机科学学院) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted at IJCAI 2025, 9pages, 6figures
机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) ; School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) ; School of Engineering, Westlake University(工程学院,西湖大学)
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 12 pages, 5 figures, 4 tables
机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) ; Monash University(墨尔本大学)
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS
Comments 4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted to PACIS 2024. 15 pages, 3 figures
Journal ref https://aisel.aisnet.org/pacis2024/track07_secprivacy/track07_secprivacy/2
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments Accepted at the 30th International Conference on Multimedia Modeling (MMM 2024)