arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-06-17 至 2026-06-17 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 5 篇

2507.14632 2026-06-17 cs.CV 版本更新 91%

BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM

BusterX++: 迈向基于MLLM的统一跨模态AI生成内容检测与解释

Haiquan Wen, Tianxiao Li, Zhenglin Huang, Yiwei He, Guangliang Cheng

机构 * University of Liverpool, UK(利物浦大学)

专题命中 视频多模态 :MLLM(title,title_cn);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 提出统一多模态大模型BusterX++,通过纯强化学习策略实现图像与视频伪造检测的跨模态能力迁移,性能超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16978 2026-06-17 cs.CV 版本更新 85%

A Benchmark for Omni-Modal Reasoning in Long Videos

长视频全模态推理基准

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou, Sahal Shaji Mullappilly, Mohammad Almansoori, Noor Ahsan, Beknur Kalmakhanbet, Sambal Shikhar, Rishabh Lalla, Jean Lahoud, Mariette Awad, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) American University of Beirut(贝鲁特美国大学) Linköping University(利尔贝里大学)

专题命中 视频多模态 :omni-modal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 提出LongShOTBench基准,用于评估长视频中视觉、语音和环境音频的全模态推理,并引入无训练的全模态证据搜索代理LongShOTAgent,在105个模型上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17298 2026-06-17 cs.CV 新提交 57%

Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins

面向手术室视频的推理式文本-视频检索:基于动作驱动数字孪生

Yiqing Shen, Hao Ding, Mathias Unberath

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

AI总结 提出OR3方法,通过动作驱动数字孪生(ActDT)将视频片段转化为结构化表示,并利用大语言模型生成假设ActDT进行检索,结合证据修正实现隐式查询推理,在手术室视频检索中显著优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12899 2026-06-17 eess.SP 新提交 50%

LGVSC: A Large-Model-Driven Generative Video Semantic Communication Framework

LGVSC:一种大模型驱动的生成式视频语义通信框架

Yu Ma, Hang Yin, Li Qiao, Shuo Sun, Zhen Gao, Yin Xu, Wenjun Zhang

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出大模型驱动的生成式视频语义通信框架LGVSC,通过解耦编解码器、引入概率语义相似度评分和语义引导关键帧提取,在极低带宽下实现高效视频语义传输,并保持零样本泛化能力。

Comments Accepted by IEEE Transactions on Vehicular Technology; v2: Added appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27583 2026-06-17 q-bio.NC cs.RO 版本更新 50%

Simulating Infant First-Person Sensorimotor Experience via Motion Retargeting from Babies to Humanoids

通过从婴儿到类人机器人的运动重定向模拟婴儿第一人称感觉运动经验

Francisco M. López, Hoshinori Kanazawa, Ondrej Fiala, Yakov Balashov, Valentin Marcel, Lukas Rustler, Miles Lenz, Dongmin Kim, Yasuo Kuniyoshi, Jochen Triesch, Matej Hoffmann

机构 * German Research Foundation(德国研究基金会) Johanna Quandt foundation(约翰娜·克文特基金会) Czech Science Foundation(捷克科学基金)

专题命中 视频多模态 :multimodal(abstract)

AI总结 提出一种从单视频重建婴儿3D姿态并映射到物理/虚拟类人平台的方法,实现亚厘米级精度的多感觉流模拟,为发育研究和神经发育障碍早期检测提供新工具。

Comments Accepted at IEEE ICDL 2026. 8 pages, 6 figures. Cite as: F. M. López, H. Kanazawa, O. Fiala, Y. Balashov, V. Marcel, L. Rustler, M. Lenz, D. Kim, Y. Kuniyoshi, J. Triesch, and M. Hoffmann, "Simulating infant first-person sensorimotor experience via motion retargeting from babies to humanoids'', in 2026 IEEE International Conference on Development and Learning (ICDL). IEEE, 2026, pp. 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏