arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-02 至 2026-03-02 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 6 篇

2602.24021 2026-03-02 cs.CV 83%

Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection

在冻结的多模态大语言模型中引导和校正潜在表示流形以进行视频异常检测

Zhaolin Cai, Fan Li, Huiyu Duan, Lijun He, Guangtao Zhai

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SteerVAD通过引导和校正冻结多模态大语言模型的潜在表示流形,提升视频异常检测的性能,仅需1%训练数据即达最优效果。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14896 2026-03-02 cs.CV 83%

Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection

利用多模态大语言模型对活动的描述进行可解释的半监督视频异常检测

Furkan Mumcu, Michael J. Jones, Anoop Cherian, Yasin Yilmaz

机构 * Department of Electrical Engineering University of South Florida(电气工程系 佛罗里达州立大学) Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室(MERL))

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出利用多模态大语言模型描述活动,以实现可解释的半监督视频异常检测,有效检测复杂交互异常并提升模型可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23937 2026-03-02 cs.RO cs.CV 79%

Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos

通过真实世界室内游览视频的多模态事件知识增强视觉语言导航

Haoxuan Xu, Tianfu Li, Wenbo Chen, Yi Liu, Xingxing Zuo, Yaoxian Song, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学) Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(马尔代夫 bin Zayed 大学人工智能学院) Hangzhou City University(杭州城市学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于多模态事件知识的视觉语言导航增强方法,通过构建大规模时空知识图谱并结合层次检索机制,提升长视界推理和粗粒度指令处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08578 2026-03-02 cs.CV 79%

Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities

多模态知识蒸馏用于抗缺失模态的自体视觉动作识别

Maria Santos-Villafranca, Dustin Carrión-Ojeda, Alejandro Perez-Yus, Jesus Bermudez-Cameo, Jose J. Guerrero, Simone Schaub-Meyer

机构 * I3A – University of Zaragoza(萨拉戈萨大学I3A研究中心) Technical University of Darmstadt, Department of Computer Science(达姆施塔特技术大学计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 KARMMA通过多模态知识蒸馏实现抗缺失模态的自体视觉动作识别,以轻量模型提升机器人部署效率。

Comments Project Page: https://visinf.github.io/KARMMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00805 2026-03-02 cs.CV 70%

Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding

通过草稿思考:用于高效长视频理解的推测性时间推理

Pengfei Hu, Meng Cao, Yingyao Wang, Yi Wang, Jiahua Dong, Jun Song, Yu Cheng, Bo Zheng, Xiaodan Liang

机构 * MBZUAI(马克斯·普朗克智能研究院) Alibaba Group Holding Limited(阿里巴巴集团控股有限公司) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 SpecTemp通过协作双模型设计,实现高效长视频理解,平衡效率与准确性,提升推理速度。

Comments Accepted by CVPR 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23784 2026-03-02 cs.LG cs.AI q-fin.CP q-fin.TR 57%

TradeFM: A Generative Foundation Model for Trade-flow and Market Microstructure

TradeFM: 一种用于交易流和市场微观结构的生成基础模型

Maxime Kawawa-Beaudan, Srijan Sood, Kassiani Papasotiriou, Daniel Borrajo, Manuela Veloso

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 TradeFM是一种基于生成Transformer的市场微观结构基础模型,通过学习大规模交易数据实现跨资产泛化,有效捕捉市场特征并提升合成数据生成与交易策略研究能力。

Comments 29 pages, 17 figures, 6 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏