arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-07-20 至 2026-07-20 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 6 篇

2607.15592 2026-07-20 cs.AI 新提交 93%

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

MGDT:具有关系自适应专家混合的MLLM引导扩散变压器用于多模态知识图谱补全

Xu Hou, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi, Kangkang Lu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Communication University of China(中国传媒大学)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 研究多模态知识图谱补全问题,提出MGDT框架,先通过RASR-MoE模块选路径、抑干扰,再用MLLM对齐表示,最后KGDT去噪生成,实验证明该框架在三个基准数据集上性能优于基线。

Comments 8pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15299 2026-07-20 cs.MM cs.CV cs.LG 新提交 92%

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation

MLLM-DataEngine:闭合多模态指令微调数据生成的循环

Zhiyuan Zhao, Bin Wang, Linke Ouyang, Yiqi Lin, Pan Zhang, Xiaoyi Dong, Jiaqi Wang, Conghui He

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title);分类 cs.CV、cs.MM

AI总结 本文提出MLLM-DataEngine闭环系统,通过自适应坏例采样模块分析模型弱点,为GPT-4提供信息以生成高质量增量数据集,能有针对性且自动地提升MLLMs能力,有望成为MLLMs数据管理通用方案。

Comments 6 pages, 4 figures, 7 tables; accepted by ICME 2026

Journal ref 2025 IEEE International Conference on Multimedia and Expo (ICME), Nantes, France, 30 June 2025 - 04 July 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15686 2026-07-20 cs.AI 新提交 79%

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

S1-Omni:用于科学理解、预测和生成的统一多模态推理模型

Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao

机构 * ScienceOne AI(科学一号人工智能) Wenge AI(文阁人工智能)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 研究针对科学人工智能模型能力分散问题,提出S1-Omni统一多模态推理模型,基于科学数据统一表示、知识对齐和解码三个核心组件,经训练和评估,在多基准测试中表现出色优于同类模型,为统一科学建模提供实用路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29573 2026-07-20 cs.CV 版本更新 79%

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

多模态大模型中可靠性优先的细粒度生成

Xiaomeng Fan, Wei Wu, Yuwei Wu, Zhi Gao, Shiyu Luo, Mingyang Gao, Haoyu Zhao, Zhenxin Diao, Yuxuan Ba, Lijia Feng, Yunde Jia, Mehrtash Harandi

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学) Department of Electrical and Computer System Engineering, Monash University(电子与计算机系统工程系,莫纳什大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 针对多模态大模型细粒度生成易出错的问题,提出GranFact基准和层次感知评估算法,并基于直接偏好优化提出可靠性优先的偏好优化方法,提升细粒度生成同时保持可靠性。

Comments Equal contribution: Xiaomeng Fan and Wu Wei. Corresponding authors: Zhi Gao and Yunde Jia

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19235 2026-07-20 cs.CV cs.RO 版本更新 57%

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

生成模型了解空间:解锁隐式3D先验用于场景理解

Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia, Yumeng Zhang, Xiaofan Li, Xiao Tan, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出通过大规模视频生成模型的隐式空间先验提升场景理解,引入VEGA-3D框架,利用预训练视频扩散模型作为潜在世界模拟器,通过时空特征提取与语义融合增强大语言模型的几何信息,实验验证其在3D场景理解、空间推理和具身操控中的优越性。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15485 2026-07-20 cs.LG math.PR stat.ML 新提交 50%

Diffusion models recover accurate mixture weights despite score function insensitivity

扩散模型即使在得分函数不敏感的情况下也能恢复准确的混合权重

Andrew Dennehy, Ramchandran Muthukumar, Rebecca Willett, Nisha Chandramoorthy

机构 * University of Chicago(芝加哥大学) Data Science Institute(数据科学研究所) Department of Computer Science(计算机科学系)

专题命中 多模态生成 :multimodal(abstract)

AI总结 研究基于得分的生成模型在多模态分布中恢复混合权重的问题,通过定义扩散得分敏感性指数,证明其控制目标分布参数估计准确性,还展示了噪声调度对敏感性及模式放大的影响,此框架可用于恢复目标分布的定性参数。

Comments 37 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏