arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-15 至 2026-05-15 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 11 篇

2512.22317 2026-05-15 cs.LG cs.AI cs.CV 81%

LangPrecip: Language-Aware Multimodal Precipitation Nowcasting

LangPrecip: 基于语言的多模态降水现在预报

Xudong Ling, Chaorong Li, Tianxi Huang, Qian Dong, Guiduo Duan

机构 * Laboratory of Intelligent Collaborative Computing, University of Electronic Science(智能协同计算实验室,电子科学科技大学) School of Computer Science(计算机科学学院) Technology (School of Artificial Intelligence), Yibin University(技术(人工智能学院),宜宾大学) College of Humanities(人文学院) General Education, Chengdu Textile College(通识教育,成都纺织学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出LangPrecip框架,通过将气象文本作为语义运动约束,结合雷达信息进行降水演变预测,提升预测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02271 2026-05-15 cs.CV 79%

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework

医学报告生成:基于层次任务结构的跨模态因果干预框架

Yucheng Song, Yifan Ge, Junhao Li, Zhining Liao, Zhifang Liao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出HTSC-CIF框架,通过层次任务分解解决医学报告生成中领域知识不足、跨模态对齐差和虚假相关性三大问题,提升生成效果。

Comments Due to issues with the training epochs and training strategy in our paper, there are numerical errors in the result comparison table presented in the preprint. Therefore, we have decided to withdraw the manuscript for further revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14159 2026-05-15 cs.RO 78%

MIMIC-D: Multi-modal Imitation for MultI-agent Coordination with Decentralized Diffusion Policies

MIMIC-D: 多模态模仿用于多智能体协调的去中心化扩散策略

Dayi Dong, Maulik Bhatt, Seoyeon Choi, Negar Mehr

机构 * Department of Mechanical Engineering, University of California Berkeley(加州大学伯克利分校机械工程系)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出MIMIC-D,通过去中心化扩散策略实现多智能体多模态模仿学习,解决了传统方法在多模态任务中协调效率低的问题。

Comments 8 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14594 2026-05-15 cs.CV cs.GR 70%

TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation

TOPOS: 高保真且高效的工业级3D人脸生成

Bojun Xiong, Zoubin Bi, Xinghui Peng, Yunmu Wang, Junchen Deng, Jun Liang, Jing Li, Bowen Cai, Huan Fu

机构 * HUJING Digital Media & Entertainment Group(华景数字媒体与娱乐集团)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV

AI总结 TOPOS框架通过统一拓扑结构实现单图像条件下的3D人脸生成,生成具有固定拓扑的高保真人脸网格,提升顶点级对应一致性,优于传统方法。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14068 2026-05-15 cs.CV 70%

CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning

CoCoEdit:通过区域正则化的强化学习实现内容一致的图像编辑

Yuhui Wu, Chenxi Xie, Ruibin Li, Liyi Chen, Qiaosi Yi, Lei Zhang

机构 * The Hong Kong Polytechnic University, Hong Kong(香港理工大学) OPPO Research Institute, ShenZhen, China(OPPO研究院,深圳,中国)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出CoCoEdit框架,通过区域正则化强化学习提升图像编辑内容一致性,利用改进的指令和掩码生成高质量训练集,并引入像素相似度奖励和区域正则化器,提升编辑质量和一致性。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13609 2026-05-15 cs.CV cs.LG 57%

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Do-Undo基准:图像生成中动作理解的可逆性

Shweta Mahajan, Shreya Kadambi, Hoang Le, Rajeev Yasarla, Apratim Bhattacharyya, Munawar Hayat, Fatih Porikli

机构 * York University(约克大学) Vector Institute for AI(人工智能向量研究所) Qualcomm AI Research(高通人工智能研究)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Do-Undo任务和基准,解决视觉-语言模型在理解并生成由现实动作驱动的场景变换中的关键缺口。通过要求模型模拟现实动作的后果并逆转至原始状态,测试真正的因果理解。实验表明当前模型在动作可逆性上存在困难,凸显评估动作理解的必要性。

Comments Project page: https://s-mahajan.github.io/Do-Undo-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14842 2026-05-15 cs.CV 57%

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

编辑推荐:通过原子实体分析评估图像编辑中的抽象意图

Mor Ventura, Roy Hirsch, Yonatan Bitton, Regev Cohen, Roi Reichart

机构 * Technion – Israel Institute of Technology(技术ion – 以色列理工学院) Google Research(谷歌研究)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出通过原子实体分析评估图像编辑中抽象意图的方法,引入Entity-Rubrics框架并构建AbstractEdit基准,揭示现有模型在平衡意图与保留之间的挑战,强调结合先进LLM和迭代思维的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14626 2026-05-15 cs.CV 57%

UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation

UniTriGen: 一致的可见-红外-标签三元组生成用于少样本RGB-T语义分割

Ping Zhou, Haoyu Wang, Mengmeng Zheng, Lei Zhang, Wei Wei, Chen Ding, Fei Zhou

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) School of Computer Science & Technology, Xi’an University of Posts & Telecommunications(西安邮电大学计算机科学与技术学院) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 UniTriGen提出统一三元组生成框架,通过共享潜在空间和扩散过程生成一致的可见-红外-标签三元组,解决少样本RGB-T语义分割中的对齐问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09304 2026-05-15 cs.CV 57%

GeRM: A Generative Rendering Model From Physically Realistic to Photorealistic

GeRM:从物理真实到照片级的生成渲染模型

Jiayuan Lu, Rengan Xie, Xuancheng Jin, Zhizhen Wu, Qi Ye, Tian Xie, Hujun Bao, Rui Wang. Yuchi Huo

机构 * State Key Lab of CAD\&CG, Zhejiang University Zhejiang Lab China State Key Laboratory of Industrial Control Technology, Zhejiang University China Zhejiang University China Zhejiang Lab State Key Laboratory of Industrial Control Technology, Zhejiang University Zhejiang University

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 GeRM是首个多模态生成渲染模型,通过学习分布转移向量场实现从物理真实到照片级的过渡,结合控制网络和多代理框架提升图像生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17588 2026-05-15 cs.CV 57%

HERO: Hierarchical Extrapolation and Refresh for Efficient World Models

HERO:高效世界模型的分层外推与刷新

Quanjian Song, Xinyu Wang, Donghao Zhou, Jingyu Lin, Cunjian Chen, Yue Ma

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 HERO通过分层策略提升世界模型推理效率,采用补丁刷新和线性外推技术,在保持质量的同时实现1.73倍加速。

Comments 12 pages in total

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05319 2026-05-15 cs.LG 50%

Accelerated Sequential Flow Matching: A Bayesian Filtering Perspective

加速的序列流匹配:从贝叶斯过滤视角

Yinan Huang, Hans Hao-Hsun Hsu, Junran Wang, Bo Dai, Pan Li

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出基于贝叶斯过滤的序列贝叶斯流匹配框架,通过学习概率流在时间步间传输后验分布,实现高效采样,提升实时流处理中的推断效率。

详情

展开后加载摘要…

URL PDF HTML 收藏