arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70051 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70051 篇

2604.24913 2026-04-29 cs.LG q-bio.PE 82%

Generative diffusion models for spatiotemporal influenza forecasting

生成扩散模型用于时空流感预测

Joseph Lemaitre, Justin Lessler

机构 * Department of Epidemiology, Gillings School of Global Public Health, University of North Carolina at Chapel Hill(流行病学系,全球公共卫生学院,北卡罗来纳大学查佩尔希尔分校) Department of Epidemiology, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD 21205, USA(流行病学系,约翰·霍普金斯伯恩斯坦公共卫生学院,巴尔的摩,MD 21205, USA) Carolina Population Center, University of North Carolina at Chapel Hill(卡罗来纳人口中心,北卡罗来纳大学查佩尔希尔分校)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract)

AI总结 本文提出Influpaint模型,利用扩散模型对流感流行病进行时空预测,通过将流感季节编码为时空图像,学习疾病动态的丰富分布,并在回顾评估中实现了与领先集成方法相当的预测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24487 2026-04-28 cs.RO 82%

Guiding Vector Field Generation via Score-based Diffusion Model

通过基于分数的扩散模型引导向量场生成

Zirui Chen, Shiliang Guo, Shiyu Zhao

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) WINDY Lab, Department of Artificial Intelligence, Westlake University(西湖大学人工智能系WINDY实验室)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出基于分数的引导向量场(SGVF),利用生成模型直接从数据分布构建向量场,解决传统方法在复杂路径中的不足,实验表明其在机器人导航中表现优异。

Comments 8 pages, 6 figrues, ICRA2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13312 2026-04-28 cs.RO cs.AI cs.LG 82%

EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation

EL3DD:扩展的潜在3D扩散用于语言条件的多任务操作

Jonas Bode, Raphael Memmesheimer, Sven Behnke

机构 * Autonomous Intelligent Systems, University of Bonn(博恩大学自主智能系统)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出EL3DD模型,通过结合视觉和文本输入,利用扩散模型生成精确的机器人轨迹,提升多任务操作的性能和长周期成功率。

Comments 10 pages; 2 figures; 1 table

Journal ref European Robotics Forum 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13366 2026-04-21 cs.LG cs.RO cs.SY eess.SY 82%

Diffusion Sequence Models for Generative In-Context Meta-Learning of Robot Dynamics

扩散序列模型用于生成性上下文元学习机器人动力学

Angelo Moroncelli, Matteo Rufolo, Gunes Cagin Aydin, Asad Ali Shahid, Loris Roveda

机构 * University of Applied Science and Arts of Southern Switzerland, Department of Innovative Technologies, IDSIA-SUPSI(瑞士南方应用科学与艺术大学创新技术系,IDSIA-SUPSI) Università della Svizzera Italiana, Faculty of Informatics(瑞士意大利大学信息学院) Politecnico di Milano, Mechanical Department(米兰理工学院机械系)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract)

AI总结 本文提出利用扩散序列模型进行机器人动力学的生成性上下文元学习,通过对比确定性和生成性模型,展示扩散模型在分布偏移下的鲁棒性提升。

Comments Angelo Moroncelli, Matteo Rufolo and Gunes Cagin Aydin contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17566 2026-04-21 eess.SY cs.LG cs.SY physics.flu-dyn 82%

Target Parameterization in Diffusion Models for Nonlinear Spatiotemporal System Identification

扩散模型中的目标参数化用于非线性时空系统识别

Achraf El Messaoudi, Noureddine Khaous, Karim Cherifi

机构 * Université Marie et Louis Pasteur, SUPMICROTECH, CNRS, institut FEMTO-ST(玛丽-路易·巴斯蒂安大学,SUPMICROTECH,法国国家科学研究中心,FEMTO-ST研究所) LLF, CNRS, Université Paris Cité(LLF,法国国家科学研究中心,巴黎cité大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文探讨了扩散模型在非线性时空系统识别中的目标参数化问题,通过湍流模拟实验表明,清洁状态预测能提升 rollout 稳定性和减少长时误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20128 2026-04-14 cs.GR cs.AI cs.CV cs.MM 82%

KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation

KSDiff: 关键帧增强型语音感知双路径扩散模型用于面部动画

Tianle Lyu, Junchuan Zhao, Ye Wang

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.GR、cs.MM

AI总结 本文提出KSDiff框架,通过双路径语音编码器和关键帧建立学习模块,提升语音驱动面部动画的唇同步和头部姿态自然度。

Comments Paper accepted at ICASSP 2026, 5 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04335 2026-04-10 cs.DC 82%

GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads

GENSERVE:高效共服务异构扩散模型工作负载

Fanjiang Ye, Zhangke Li, Xinrui Zhong, Ethan Ma, Russell Chen, Kaijian Wang, Jingwei Zuo, Desen Sun, Ye Cao, Triston Cao, Myungjin Lee, Arvind Krishnamurthy, Yuke Wang

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

AI总结 GENSERVE通过利用扩散过程的可预测性,优化共服务效率,解决异构工作负载下的延迟SLA问题,实验表明其在多种配置下将SLA达成率提升44%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22729 2026-04-02 stat.ML cs.LG math.ST stat.TH 82%

Identifying Drift, Diffusion, and Causal Structure from Temporal Snapshots

从时间快照中识别漂移、扩散和因果结构

Vincent Guan, Joseph Janssen, Hossein Rahmani, Andrew Warren, Stephen Zhang, Elina Robeva, Geoffrey Schiebinger

机构 * University of British Columbia(不列颠哥伦比亚大学) University of Melbourne(墨尔本大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出了一种从时间边际联合识别SDE漂移和扩散的方法,证明了非识别性仅在初始分布具有广义旋转对称性时发生,并展示了因果图可从SDE参数中恢复。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00007 2026-04-02 cs.CL cs.AI 82%

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

Dynin-Omni: 多模态统一的大扩散语言模型

Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee, Sieun Hyeon, Mintaek Lim, Yunseok Han, Dogeun Kim, Hoeun Lee, Hyunggeun Kim, Jaeyoung Do

机构 * AIDAS Lab Seoul National University [1.5ex] Project Page Code Model Demo -2em

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 Dynin-Omni首次提出基于掩码扩散的多模态基础模型,统一文本、图像、语音的理解与生成,以及视频理解,通过共享离散令牌空间实现双向上下文迭代优化。

Comments Project Page: https://dynin.ai/omni/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28209 2026-03-31 cs.SD 82%

On the Usefulness of Diffusion-Based Room Impulse Response Interpolation to Microphone Array Processing

扩散基房间脉冲响应插值在微电极阵列处理中的实用性

Sagi Della Torre, Mirco Pezzoli, Fabio Antonacci, Sharon Gannot

机构 * Bar-Ilan University(巴伊兰大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract)

AI总结 本文探讨了扩散基插值在多麦克风阵列处理中的应用,验证了其在真实房间脉冲响应插值中的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13213 2026-03-31 stat.ML cs.LG 82%

Diffusion Models with Double Guidance: Generate with aggregated datasets

具有双重引导的扩散模型:利用聚合数据集生成

Yanfeng Yang, Kenji Fukumizu

机构 * Graduate University of Advanced Studies (SOKENDAI), Kanagawa, Japan(综合研究大学院大学(SOKENDAI),神奈川,日本) The Institute of Statistical Mathematics, Tokyo, Japan(统计数理研究所,东京,日本)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出一种新的扩散模型,通过双重引导实现精确条件生成,即使训练数据中没有同时包含所有条件。在分子和图像生成任务中,该方法在对齐目标条件分布和缺失条件控制方面优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24594 2026-03-26 cs.LG cs.NA math.NA stat.ML 82%

Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method

在扩散模型中使用多级欧拉-马尔可夫方法实现多项式加速

Arthur Jacot

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出多级欧拉-马尔可夫方法,通过不同精度的近似器高效求解SDE和ODE,针对HTMC区域实现比传统方法更高的计算效率,应用于扩散模型时可显著提升图像生成速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04058 2026-03-24 cs.LG 82%

Unlearning in Diffusion models under Data Constraints: A Variational Inference Approach

在数据约束下扩散模型的卸载:一种变分推断方法

Subhodip Panda, Varun M S, Shreyans Jain, Sarthak Kumar Maharana, Prathosh A. P

机构 * Department of ECE, Indian Institute of Science(印度科学研究院电子工程系)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出变分扩散卸载方法,用于在数据受限条件下防止预训练扩散模型生成不良内容,通过优化损失函数中的可塑性诱导器和稳定性正则化器来提高效率。

Journal ref Transaction on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02929 2026-03-24 q-bio.QM 82%

Microscopy image reconstruction with physics-informed denoising diffusion probabilistic model

利用物理信息的去噪扩散概率模型进行显微图像重建

Rui Li, Gabriel della Maggiora, Vardan Andriasyan, Anthony Petkidis, Artsemi Yushkevich, Mikhail Kudryashev, Artur Yakimovich

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract)

AI总结 本文提出将显微成像物理问题融入模型损失函数,通过合成数据训练,改进DDPM以减少伪影,提升显微图像重建质量。

Comments 16 pages, 5 figures

Journal ref Communications Engineering 3, no. 1 (2024): 186

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01003 2026-03-24 cs.LG cs.RO 82%

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

收缩扩散策略:通过微分方程的收缩分数基采样实现鲁棒动作扩散

Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh, Anas El Houssaini, David Meger, Gregory Dudek, Hsiu-Chin Lin

机构 * Department of Electrical and Computer Engineering, McGill University(麦吉尔大学电气与计算机工程系) Department of Computer Science, McGill University(麦吉尔大学计算机科学系) Mila–Quebec AI Institute(魁北克人工智能研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出收缩扩散策略,通过引入收缩行为提升扩散采样动态的鲁棒性,减少求解和分数匹配误差,降低动作方差,在数据稀缺情况下表现优异。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19752 2026-03-23 cs.LG 82%

Fast 3D Diffusion for Scalable Granular Media Synthesis

快速3D扩散用于可扩展的颗粒介质合成

Muhammad Moeeze Hassan, Régis Cottereau, Filippo Gatti, Patryk Dec

机构 * Aix Marseille Univ, CNRS, Centrale Med, LMA(艾克斯-马赛大学,法国国家科学研究中心,中央医疗学院,马赛实验室) Université Paris Saclay, CentraleSupélec, ENS Paris-Saclay, CNRS, Laboratoire de Mécanique Paris-Saclay UMR 9026(巴黎萨克雷大学,中央超算学院,巴黎-萨克雷高等师范学院,法国国家科学研究中心,巴黎机械实验室(UMR 9026)) Innovation and Research Department, SNCF(法国国家铁路公司创新与研究部)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract)

AI总结 本文提出基于3D扩散模型的新型生成管道,能够高效合成大规模颗粒介质,通过两阶段流程生成并拼接 voxel 网格,实现机械真实配置,显著提升模拟效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19360 2026-03-23 cs.LG 82%

Warm-Start Flow Matching for Guaranteed Fast Text/Image Generation

预热启动流匹配用于保证快速文本/图像生成

Minyoung Kim

机构 * Samsung AI Center Cambridge(三星AI研究中心剑桥)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract)

AI总结 本文提出预热启动流匹配(WS-FM)方法,通过利用轻量级生成模型快速生成初始样本,减少流匹配算法的生成时间,同时保证生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06405 2026-03-20 math.DG 82%

Diffusion-Shock PDEs for Deep Learning on Position-Orientation Space

扩散-冲击偏微分方程用于位置-方向空间的深度学习

Finn M. Sherry, Kristina Schaefer, Remco Duits

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract)

AI总结 本文将正则化扩散-冲击滤波器从欧几里得空间扩展到位置-方向空间,解决了交叉结构增强与修复的问题,并提出了基于规范框架的算法以提高鲁棒性,同时证明了泛化扩散的数学性质。

Comments Accepted in the Journal of Mathematical Imaging and Vision Special Issue on Scale Space and Variational Methods in Computer Vision 2025 (SSVM). arXiv admin note: text overlap with arXiv:2502.17146

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22629 2026-03-18 cs.CL cs.AI 82%

Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models

时间退火扰动采样:扩散语言模型的多样化生成

Jingxuan Wu, Zhenglin Wan, Xingrui Yu, Yuzhe Yang, Yiqiao Huang, Ivor Tsang, Yang You

机构 * The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) National University of Singapore(新加坡国立大学) CFAR, Agency for Science, Technology and Research(科技研究局CFAR) University of California, Santa Barbara(加州大学圣芭芭拉分校) Harvard University(哈佛大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出时间退火扰动采样方法,通过在扩散过程中早期鼓励语义分支并逐步减少扰动,提升生成多样性,同时保持生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04308 2026-03-13 cs.LG cs.AI cs.SI physics.soc-ph 82%

HOG-Diff: Higher-Order Guided Diffusion for Graph Generation

HOG-Diff: 基于更高阶引导的图生成扩散模型

Yiming Huang, Tolga Birdal

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 HOG-Diff通过高阶拓扑引导和扩散桥梁,实现图生成的高阶拓扑结构捕捉,展示了方法在多样领域和大规模设置中的优越性能。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04559 2026-03-09 cs.LG cs.AI 82%

Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function

通过软Q函数的重参数化策略梯度进行扩散微调

Hyeongyu Kang, Jaewoo Lee, Woocheol Shin, Kiyoung Om, Jinkyoo Park

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

AI总结 SQDF通过软Q函数的重参数化策略梯度方法,解决扩散模型奖励过优化问题,提升样本多样性和自然性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21739 2026-03-06 cs.SD cs.LG eess.AS 82%

Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription

噪声到音符:基于扩散的生成与细化用于自动鼓件转录

Michael Yeung, Keisuke Toyama, Toya Teramoto, Shusuke Takahashi, Tamaki Kojima

机构 * Sony Group Corporation(索尼集团)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract)

AI总结 本研究提出N2N框架,利用扩散模型将音频条件高斯噪声转化为鼓件事件,结合退火伪Huber损失和音乐基础模型特征提升鲁棒性,实现自动鼓件转录的生成与细化。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21472 2026-02-26 cs.LG 82%

The Design Space of Tri-Modal Masked Diffusion Models

三模态掩码扩散模型的设计空间

Louis Bethune, Victor Turrisi, Bruno Kacper Mlodozeniec, Pau Rodriguez Lopez, Lokesh Boominathan, Nikhil Bhendawade, Amitis Shidani, Joris Pelemans, Theo X. Olausson, Devon Hjelm, Paul Dixon, Joao Monteiro, Pierre Ablin, Vishnu Banna, Arno Blaas, Nick Henderson, Kari Noriy, Dan Busbridge, Josh Susskind, Marco Cuturi, Irina Belousova, Luca Zappella, Russ Webb, Jason Ramapuram

机构 * Apple(苹果公司) Google Deepmind(谷歌DeepMind) University of Cambridge(剑桥大学) MIT(麻省理工学院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

AI总结 本文提出了一种从头开始预训练的三模态掩码扩散模型,通过分析多模态扩展定律和批大小效应,实现了优化的推理采样,并在文本生成、图像生成和语音生成任务中取得了显著成果。

Comments 41 pages, 29 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20193 2026-02-25 cs.CR cs.AI 82%

When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks

当后门超越触发器:在编码器攻击下扩散模型的语义漂移

Shenyang Chen, Liuwan Zhu

机构 * Google(谷歌) Electrical and Computer Engineering Department University of Hawai‘i at Mānoa(电气与计算机工程系,夏威夷大学马诺阿分校)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

AI总结 研究揭示了在编码器攻击下扩散模型的语义漂移问题,提出SEMAX框架用于量化结构退化,强调几何审计的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08011 2026-02-24 cs.AI 82%

Training-Free Safe Denoisers for Safe Use of Diffusion Models

无需训练的安全去噪器用于安全使用扩散模型

Mingyu Kim, Dongjun Kim, Amman Yusuf, Stefano Ermon, Mijung Park

机构 * CS, UBC(计算机科学系,不列颠哥伦比亚大学) CS, Standford(计算机科学系,斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出无需训练的无训练安全去噪器,通过修改采样轨迹避免不安全数据区域,提升扩散模型的安全使用。

Comments NeurIPS2025, Code: https://github.com/MingyuKim87/Safe_Denoiser

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10040 2026-02-24 cs.RO 82%

Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation

用于长时程机器人操作的扩散轨迹引导策略

Shichao Fan, Quantao Yang, Yajie Liu, Kun Wu, Zhengping Che, Qingjie Liu, Min Wan

机构 * School of Mechanical Engineering and Automation, BeiHang University(机械工程与自动化学院,北航) School of Computer Science and Engineering, BeiHang University(计算机科学与工程学院,北航) Beijing Innovation Center of Humanoid Robotics(人形机器人创新中心) Division of Robotics, Perception and Learning (RPL), KTH Royal Institute of Technology(机器人、感知与学习 division,皇家理工学院)

专题命中 扩散模型 :diffusion(title,abstract);generative vision(abstract)

AI总结 本文提出DTP框架,通过生成轨迹减少模仿学习中的误差累积,提升长时程机器人任务的性能。

Comments 8 pages, 5 figures, accepted to IEEE Robotics and Automation Letters (RAL)

Journal ref IEEE Robotics and Automation Letters (Volume: 10, Issue: 12, December 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15500 2026-02-24 stat.ML cs.AI cs.LG math.ST stat.TH 82%

Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective

修正流的低维适应:从扩散和随机定位视角

Saptarshi Roy, Alessandro Rinaldo, Purnamrita Sarkar

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

AI总结 本文提出一种基于修正流的低维适应采样方法,通过时间离散化方案和随机定位理论,提升采样效率和精度。

Comments 32 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16570 2026-02-19 cs.LG cs.DS 82%

Steering diffusion models with quadratic rewards: a fine-grained analysis

通过二次奖励引导扩散模型:细粒度分析

Ankur Moitra, Andrej Risteski, Dhruv Rohatgi

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文通过二次奖励引导扩散模型,分析了不同奖励函数下的采样效率及计算复杂性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04468 2026-02-12 cs.LG eess.IV math.PR 82%

Iterative Importance Fine-tuning of Diffusion Models

扩散模型的迭代重要性微调

Alexander Denker, Shreyas Padhy, Francisco Vargas, Johannes Hertrich

机构 * University College London(伦敦大学学院) University of Cambridge(剑桥大学) Xaira Therapeutics ENS Paris - PSL & Inria Mokaplan(巴黎高等师范学院 - PSL & Inria Mokaplan)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

AI总结 本文提出了一种基于自监督学习的扩散模型微调方法,通过迭代优化控制实现高效的条件采样,适用于图像生成和逆问题解决。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01362 2026-02-03 cs.CL 82%

Balancing Understanding and Generation in Discrete Diffusion Models

在离散扩散模型中实现理解和生成的平衡

Yue Liu, Yuzhong Zhao, Zheyong Xie, Qixiang Ye, Jianbin Jiao, Yao Hu, Shaosheng Cao, Yunfan Liu

机构 * Xiaohongshu Inc.(小红书公司)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 XDLM通过平稳噪声核平衡理解和生成能力,实现MDLM和UDLM的理论统一并提升生成质量与理解能力。

Comments 32 pages, Code is available at https://github.com/MzeroMiko/XDLM

详情

展开后加载摘要…

URL PDF HTML 收藏