arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 69868 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 69868 篇

2309.00952 2026-03-19 cs.CL cs.AI 89%

Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities

桥扩散模型:将中文文本到图像扩散模型与英文社区相结合

Shanyuan Liu, Bo Cheng, Yuhang Ma, Liebucha Wu, Ao Ma, Xiaoyu Wu, Dawei Leng, Yuhui Yin

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(title,abstract);image generation(abstract)

AI总结 本文提出桥扩散模型,通过端到端结构学习中文语义并保持潜在空间与英文文本到图像模型的兼容性,实现中英文语义融合生成。

Comments Accepted as Oral at AAAI 2025. 8 pages, 5 figures. Published in Proceedings of the 39th AAAI Conference on Artificial Intelligence. Code: https://github.com/360CVGroup/Bridge_Diffusion_Model

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 39(5), 5541-5549 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04030 2026-08-06 cs.GR cs.AI cs.CV cs.CY cs.LG 新提交 89%

NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

NuclearDiffusion:用于学习核能概念的文生成像基础模型

Mohammed I. Radaideh, Jeremy Moon, Andre Gala-Garza, Emma Son, Yug Shah, Majdi I. Radaideh

专题命中 扩散模型 :text-to-image(title,abstract);diffusion(abstract,abstract_cn);image generation(abstract);image synthesis(abstract)

AI总结 本研究通过微调开源扩散模型构建核能文生成像模型,发现微调效果依赖生成架构,且微调后开源模型在专业核能图像生成上优于主流商业系统,证实领域特定微调是开发可信领域生成式AI的可行路径。

Comments 29 pages, 10 figures, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13273 2026-06-02 cs.AI cs.LG 89%

EMoE: Training-Free Expert Disagreement for Uncertainty-Aware Text-to-Image Diffusion

EMoE: 面向不确定性感知的文本到图像扩散的无训练专家分歧方法

Lucas Berry, Axel Brando, Wei-Di Chang, Juan Camilo Gamboa Higuera, David Meger

机构 * McGill University(麦吉尔大学) Barcelona Supercomputing Center (BSC)(巴塞罗那超级计算中心 (BSC)) Ideogram AI

专题命中 扩散模型 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

AI总结 提出EMoE方法,通过预训练MoE扩散模型中早期MoE层的专家分歧,无需训练即可估计认知不确定性,用于提示风险诊断和生成质量排序。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29809 2026-05-29 cs.CR cs.CV cs.GR cs.LG cs.MM 89%

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

Cert-LAS:通过层自适应平滑实现文本到图像扩散模型的认证模型所有权验证

Leyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu, Dacheng Tao

机构 * Generative AI Lab, College of Computing Data Science, Nanyang Technological University, Singapore Department of Computer Science Engineering, Texas A\&M University, USA

专题命中 扩散模型 :text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV、cs.GR、cs.MM

AI总结 提出Cert-LAS方法,基于层自适应平滑和扩散分类器嵌入水印,通过假设检验验证模型所有权,并证明在恶意移除攻击下仍能可靠验证。

Comments This paper has been accepted to the International Conference on Machine Learning (ICML) 2026. 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05598 2025-11-11 cs.CR eess.IV 89%

Diffusion-Based Image Editing: An Unforeseen Adversary to Robust Invisible Watermarks

Wenkai Fu, Finn Carter, Yue Wang, Emily Davis, Bo Zhang

专题命中 扩散模型 :diffusion(title,abstract);image editing(title,abstract);image generation(abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03491 2025-11-04 eess.IV 89%

LanPaint: Training-Free Diffusion Inpainting with Asymptotically Exact and Fast Conditional Sampling

Candi Zheng, Yuan Lan, Yang Wang

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);image generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15516 2025-08-29 cs.CY 89%

When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models

Matthias Schneider, Thilo Hagendorff

专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract);text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11972 2025-08-05 cs.DC 89%

MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models

Yuchen Xia, Divyam Sharma, Yichao Yuan, Souvik Kundu, Nishil Talati

专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract);text-to-image(abstract)

Comments To appear in ASPLOS'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02035 2024-11-05 cs.CR 89%

Robustness of Watermarking on Text-to-Image Diffusion Models

Xiaodong Wu, Xiangman Li, Jianbing Ni

专题命中 扩散模型 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

Comments We find an error in one of the proposed attack methods, which significantly impact the correctness. In addition, the experiment is not solid enough to support the results

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04859 2023-07-12 cs.CV cs.GR cs.LG 89%

Articulated 3D Head Avatar Generation using Text-to-Image Diffusion Models

Alexander W. Bergman, Wang Yifan, Gordon Wetzstein

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(title,abstract);分类 cs.CV、cs.GR

Comments Project website: http://www.computationalimaging.org/publications/articulated-diffusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04488 2023-06-21 cs.CV cs.GR cs.LG 89%

Multi-Concept Customization of Text-to-Image Diffusion

Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, Jun-Yan Zhu

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(title,abstract);分类 cs.CV、cs.GR

Comments Updated v2 with results on the new CustomConcept101 dataset https://www.cs.cmu.edu/~custom-diffusion/dataset.html Project webpage: https://www.cs.cmu.edu/~custom-diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23783 2026-08-13 cs.CV 版本更新 89%

Diffusion Probe: Generated Image Result Prediction Using CNN Probes

扩散探针:利用CNN探针进行生成图像结果预测

Benlei Cui, Bukun Huang, Zhizeng Ye, Xuemei Dong, Tuo Chen, Hui Xue, Dingkang Yang, Longtao Huang, Jingqun Tang, Haiwen Hong

机构 * Alibaba Group(阿里巴巴集团) Laboratory for Statistical Monitoring and Intelligent Governance of Common Prosperity, School of Statistics and Data Science, Zhejiang Gongshang University(浙江工商大学统计与数据科学学院共同富裕统计监测与智能治理实验室) Southeast University(东南大学) College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) ByteDance Inc.(字节跳动有限公司)

专题命中 扩散模型 :diffusion(title,summary_cn);text-to-image(abstract);分类 cs.CV

AI总结 本文提出Diffusion Probe框架,通过早期扩散交叉注意力分布预测最终图像质量,提升文本生成图像的效率与质量。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03284 2026-08-05 cs.CV cs.AI 新提交 89%

Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates

通过中间干净图像估计实现安全文本引导图像生成的测试时缩放

Jinya Sakurai, Shueicheng Yan, Xun Xu

专题命中 扩散模型 :diffusion(summary_cn,abstract);image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 该研究针对文本到图像扩散模型的安全问题,提出利用中间干净图像估计和稀疏边际目标的测试时缩放方法,在 Stable Diffusion 上实现了更优的安全性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19244 2026-07-17 cs.CV 版本更新 89%

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

Lavida-O:用于统一多模态理解与生成的弹性大掩码扩散模型

Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin, Zijun Wei, Aditya Grover, Jason Kuen

机构 * Adobe(Adobe公司) UCLA(加州大学洛杉矶分校)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)

AI总结 研究提出Lavida-O统一多模态MDM,采用Elastic-MoT架构,结合轻量级生成与理解分支,通过多种技术支持高效生成。该模型在多模态任务基准测试中性能领先,优于现有模型,还能加速推理,成为可扩展多模态推理和生成新范式。

Comments 31 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27813 2026-05-28 cs.CV cs.AI cs.LG 89%

Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

残差化时间稀疏自编码器用于解释扩散模型

Calvin Yeung, Prathyush Poduval, Ali Zakeri, Zhuowen Zou, Mohsen Imani

机构 * University of California, Irvine(加州大学 Irvine 分校)

专题命中 扩散模型 :diffusion(title,summary_cn);text-to-image(abstract);分类 cs.CV

AI总结 提出残差化时间稀疏自编码器,通过去噪时间步间的线性预测残差学习扩散激活轨迹中的可解释特征,并在Stable Diffusion 1.5上验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08774 2025-12-10 cs.CV cs.AI 89%

Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps

通过可解释人工智能基于的缺陷激活图来改进扩散模型中的视觉伪影

Seoyeon Lee, Gwangyeol Yu, Chaewon Kim, Jonghyuk Park

机构 * Kookmin University(韩国庆熙大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);inpainting(abstract)

AI总结 通过可解释人工智能生成缺陷激活图,改进扩散模型中的视觉伪影问题,提升图像生成质量。

Comments 10 pages, 9 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06308 2025-10-09 cs.CV 89%

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

Yi Xin, Qi Qin, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu, Xiaohong Liu, Yu Qiao, Yihao Liu

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Nanjing University(南京大学) The University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)

Comments 33 pages, 13 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09416 2024-12-02 cs.CV 89%

Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization

Qihao Liu, Zhanpeng Zeng, Ju He, Qihang Yu, Xiaohui Shen, Liang-Chieh Chen

专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract);分类 cs.CV

Comments Introducing DiMR, a new diffusion backbone that surpasses all existing image generation models of various sizes on ImageNet 256 with only 505M parameters. Project page: https://qihao067.github.io/projects/DiMR

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09865 2022-09-01 cs.CV 89%

RePaint: Inpainting using Denoising Diffusion Probabilistic Models

Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, Luc Van Gool

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV

Comments We missed out on other diffusion models that work on inpainting. We corrected that and apologize for this mistake

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14765 2026-03-04 cs.CV cs.AI cs.GR 88%

Inpainting the Red Planet: Diffusion Models for the Reconstruction of Martian Environments in Virtual Reality

为虚拟现实重建火星环境的红色星球修复:扩散模型

Giuseppe Lorenzo Catalano, Agata Marta Soccini

机构 * Computer Science Department, Università degli Studi di Torino(托斯纳大学计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV、cs.GR

AI总结 本文提出基于无条件扩散模型的火星表面重建方法,利用增强数据集和非均匀重缩放策略,优于传统空洞填充技术,在重建精度和感知相似性上表现更优。

Comments 21 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10860 2026-01-22 cs.CV cs.GR 88%

RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion Priors

RI3D: 少样本高斯点云渲染与修复与修复扩散先验

Avinash Paliwal, Xilong Zhou, Wei Ye, Jinhui Xiong, Rakesh Ranjan, Nima Khademi Kalantari

机构 * Texas A&M University(德克萨斯A&M大学) Meta Reality Labs(Meta现实实验室) Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV、cs.GR

AI总结 RI3D通过分离视图合成任务并结合修复与修复扩散模型,实现了高质量的少样本3D渲染与缺失区域重建。

Comments ICCV 2025, Project page: https://people.engr.tamu.edu/nimak/Papers/RI3D, Code: https://github.com/avinashpaliwal/RI3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07050 2025-08-13 cs.CV cs.AI cs.MM 88%

TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation

Victor Shea-Jay Huang, Le Zhuo, Yi Xin, Zhaokai Wang, Fu-Yun Wang, Yuchi Wang, Renrui Zhang, Peng Gao, Hongsheng Li

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);image editing(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13023 2025-08-05 cs.CV cs.AI cs.MM 88%

Anti-Inpainting: A Proactive Defense Approach against Malicious Diffusion-based Inpainters under Unknown Conditions

Yimao Guo, Zuomin Qu, Wei Lu, Xiangyang Luo

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12847 2025-06-17 cs.GR cs.CV 88%

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Zhelun Shen, Chenming Wu, Junsheng Zhou, Chen Zhao, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Wei He, Jingdong Wang

机构 * Department of Computer Vision Technology(VIS), Baidu Inc.(计算机视觉技术系(VIS),百度公司)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV、cs.GR

Comments Technical report, 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07199 2025-03-12 cs.CV cs.AI cs.GR cs.LG 88%

RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion

Jaidev Shriram, Alex Trevithick, Lingjie Liu, Ravi Ramamoorthi

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV、cs.GR

Comments Published at 3DV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18616 2024-11-28 cs.CV cs.AI cs.GR cs.LG 88%

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

Shengqu Cai, Eric Chan, Yunzhi Zhang, Leonidas Guibas, Jiajun Wu, Gordon Wetzstein

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project page: https://primecai.github.io/dsd/

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02051 2023-08-24 cs.CV cs.AI cs.MM 88%

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini, Rita Cucchiara

专题命中 扩散模型 :diffusion(title,abstract);image editing(title,abstract);分类 cs.CV、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01147 2023-08-03 cs.CV cs.MM eess.IV 88%

Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image Generation

Guojin Zhong, Jin Yuan, Pan Wang, Kailun Yang, Weili Guan, Zhiyong Li

专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract);分类 cs.CV、cs.MM

Comments Accepted to ACM MM 2023. The code will be released at https://github.com/zgj77/FSACDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09935 2023-06-19 cs.LG cs.CV cs.GR 88%

Drag-guided diffusion models for vehicle image generation

Nikos Arechiga, Frank Permenter, Binyang Song, Chenyang Yuan

专题命中 扩散模型 :image generation(title,abstract);diffusion(title,abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01732 2023-06-05 cs.CV cs.AI cs.GR 88%

Video Colorization with Pre-trained Text-to-Image Diffusion Models

Hanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li, Tien-Tsin Wong

专题命中 扩散模型 :text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV、cs.GR

Comments project page: https://colordiffuser.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏