arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 69868 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 69868 篇

2006.10406 2020-06-19 cs.CV 86%

Fourth-Order Anisotropic Diffusion for Inpainting and Image Compression

Ikram Jumakulyyev, Thomas Schultz

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title);分类 cs.CV

Comments Accepted for publication in Springer book "Anisotropy Across Fields and Scales"

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13247 2026-06-12 cs.AI 新提交 86%

EPIG: Emotion-Based Prompting for Personalised Image Generation

EPIG:基于情感提示的个性化图像生成

Emna Othmen, Mohamed Yassine Landolsi, Lotfi Ben Romdhane

机构 * MARS Research Lab LR17ES05, ISITCom, University of Sousse(苏塞大学ISITCom学院MARS研究实验室LR17ES05)

专题命中 扩散模型 :image generation(title,abstract);text-to-image(abstract,comments);diffusion(abstract,comments)

AI总结 提出EPIG方法,利用心理学效价-唤醒模型在提示层面增强情感表达,无需训练即可控制生成图像的唤醒度,在10个多样化提示上平均唤醒误差降低14%-17%。

Comments Submitted to arXiv. 20 pages, 4 figures. Work on emotion-based prompt engineering for text-to-image diffusion models with applications in personalized image generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16255 2026-04-06 astro-ph.IM cs.AI 86%

Category-based Galaxy Image Generation via Diffusion Models

基于类别的银河图像生成:通过扩散模型

Xingzhong Fan, Hongming Tang, Yue Zeng, M. B. N. Kouwenhoven, Guangquan Zeng

机构 * Department of Physics, Xi'an Jiaotong-Liverpool University(西交利物浦大学物理系) Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) Department of Physics, The Chinese University of Hong Kong(香港中文大学物理系)

专题命中 扩散模型 :diffusion(title,abstract);image generation(title)

AI总结 本文提出GalCatDiff框架,结合银河图像特征和天体物理属性,通过增强的U-Net和新型Astro-RAB模块提升生成质量,实现高效且物理一致的银河生成。

Comments 23 pages, 10 figures. Accepted by AAS Astronomical Journal (AJ) and has now been published on https://iopscience.iop.org/article/10.3847/1538-3881/ae5064. See another independent work for further reference -- Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation (Ma, Sun et al.). Comments are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01134 2026-08-14 cs.GR cs.CV 版本更新 86%

RealMat: Realistic Materials with Diffusion and Reinforcement Learning

RealMat:结合扩散模型与强化学习的真实材质生成器

Xilong Zhou, Pedro Figueiredo, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Yiwei Hu, Nima Khademi Kalantari

机构 * Max Planck Institute for Informatics Saarbrücken Germany Texas A\&M University College Station USA Adobe Research San Jose USA Adobe Research London UK Max Planck Institute for Informatics Texas A\&M University Adobe Research

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 针对材质生成的合成数据真实感不足、真实数据规模有限的问题,提出结合SDXL微调与强化学习的RealMat,有效提升了生成材质的真实性。

Comments 12 pages, 12 figures

Journal ref Computer Graphics Forum 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19344 2026-07-22 cs.CV cs.AI cs.GR 新提交 86%

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

外观指针——扩散变压器的多模态区域控制

Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha

机构 * Brown University(布朗大学) Adobe Research(Adobe 研究院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 针对可控图像生成难题,提出外观指针方法,通过区域对应网络和空间聚合机制生成并细化指针,为扩散变压器引入模态无关的局部多模态控制接口,单一模型性能达或超现有技术。

Comments 38 Pages, Preprint with supplement

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01929 2026-03-25 cs.GR cs.AI cs.CV cs.LG 86%

Image Generation from Contextually-Contradictory Prompts

从语境矛盾提示生成图像

Saar Huberman, Or Patashnik, Omer Dahary, Ron Mokady, Daniel Cohen-Or

机构 * Tel Aviv University(特拉维夫大学) BRIA AI(BRIA人工智能)

专题命中 扩散模型 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种阶段感知提示分解框架,通过代理提示引导去噪过程,解决提示中概念矛盾导致的语义不准确问题,提升图像生成的准确性。

Comments Project page: https://tdpc2025.github.io/SAP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02580 2026-03-18 cs.CV cs.AI cs.GR cs.LG 86%

TAUE: Training-free Noise Transplant and Cultivation Diffusion Model

TAUE:无需训练的噪声移植与培育扩散模型

Daichi Nagai, Ryugo Morita, Shunsuke Kitada, Hitoshi Iyatomi

机构 * Faculty of Science and Engineering, Hosei University(恒河大学科学与工程学院) RPTU Kaiserslautern-Landau & DFKI GmbH(凯撒斯劳滕-兰道大学与DFKI GmbH)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 TAUE提出无需训练的噪声移植与培育扩散模型,通过嵌入全局结构信息和语义线索,实现多层图像生成,提升跨层一致性并支持新应用。

Comments Accepted to CVPR 2026 Findings. The first two authors contributed equally. Project Page: https://iyatomilab.github.io/TAUE

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05106 2026-03-06 cs.CV cs.GR cs.LG cs.RO 86%

NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation

NeuralRemaster: 保留相位的扩散用于结构对齐生成

Yu Zeng, Charles Ochoa, Mingyuan Zhou, Vishal M. Patel, Vitor Guizilini, Rowan McAllister

机构 * Toyota Research Institute(丰田研究院) University of Texas, Austin(德克萨斯大学奥斯汀分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 NeuralRemaster通过保留相位并随机化幅度,实现结构对齐的图像和视频生成,提升模拟到现实的转换性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24108 2026-03-04 cs.GR cs.AI cs.CV cs.LG 86%

Navigating with Annealing Guidance Scale in Diffusion Space

在扩散空间中利用退火引导尺度导航

Shai Yehezkel, Omer Dahary, Andrey Voynov, Daniel Cohen-Or

机构 * Tel Aviv University(特拉维夫大学) Google DeepMind(谷歌DeepMind)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 本文提出了一种退火引导调度器,通过动态调整引导尺度提升文本到图像生成的质量和对齐度,无需额外资源消耗。

Comments SIGGRAPH Asia, 2025. Project page: https://annealing-guidance.github.io/annealing-guidance/

Journal ref ACM Trans. Graph., Vol. 44, No. 6, Article 5. Publication date: December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17927 2026-01-27 cs.CV cs.MM 86%

RemEdit: Efficient Diffusion Editing with Riemannian Geometry

RemEdit: 基于黎曼几何的高效扩散编辑

Eashan Adhikarla, Brian D. Davison

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image editing(abstract);分类 cs.CV、cs.MM

AI总结 RemEdit通过基于黎曼几何的潜在空间导航和任务特定注意力剪枝机制,实现了高效且保真的图像编辑,同时保持实时性能。

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17259 2026-01-27 cs.CV cs.GR cs.LG 86%

Inference-Time Loss-Guided Colour Preservation in Diffusion Sampling

推理时的损失引导颜色保持在扩散采样中

Angad Singh Ahuja, Aarush Ram Anandh

机构 * Constrained Image-Synthesis Lab(受限图像合成实验室)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种无需额外训练的推理时颜色保持方法,通过区域约束和复合损失引导扩散模型,实现精准颜色控制。

Comments 25 Pages, 12 Figures, 3 Tables, 5 Appendices, 8 Algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05532 2025-10-08 cs.CV cs.GR cs.LG 86%

Teamwork: Collaborative Diffusion with Low-rank Coordination and Adaptation

Sam Sartor, Pieter Peers

机构 * College of William \& Mary Williamsburg USA College of William \& Mary

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11054 2025-07-03 cs.GR cs.CV cs.LG 86%

LUSD: Localized Update Score Distillation for Text-Guided Image Editing

Worameth Chinchuthakun, Tossaporn Saengja, Nontawat Tritrong, Pitchaporn Rewatbowornwong, Pramook Khungurn, Supasorn Suwajanakorn

专题命中 扩散模型 :image editing(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments ICCV 2025. Project page: https://github.com/sincostanx/LUSD

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12072 2024-11-20 cs.CV cs.AI cs.LG cs.MM 86%

Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution

Brian B. Moser, Stanislav Frolov, Tobias C. Nauen, Federico Raue, Andreas Dengel

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23775 2024-11-06 cs.CV cs.GR 86%

In-Context LoRA for Diffusion Transformers

Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, Jingren Zhou

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Tech report. Project page: https://ali-vilab.github.io/In-Context-LoRA-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17911 2024-07-26 cs.MM cs.AI cs.CV 86%

ReCorD: Reasoning and Correcting Diffusion for HOI Generation

Jian-Yu Jiang-Lin, Kang-Yang Huang, Ling Lo, Yi-Ning Huang, Terence Lin, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2024. Project website: https://alberthkyhky.github.io/ReCorD/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10316 2024-05-17 cs.CV cs.GR 86%

Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model

Zheng Gu, Shiyuan Yang, Jing Liao, Jing Huo, Yang Gao

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments Project page: https://analogist2d.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12422 2024-05-07 cs.CV cs.GR cs.LG 86%

DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation

Yukun Huang, Jianan Wang, Yukai Shi, Boshi Tang, Xianbiao Qi, Lei Zhang

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16225 2023-12-08 cs.GR cs.CV 86%

ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models

Yuxin Zhang, Weiming Dong, Fan Tang, Nisha Huang, Haibin Huang, Chongyang Ma, Tong-Yee Lee, Oliver Deussen, Changsheng Xu

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01409 2023-12-05 cs.CV cs.AI cs.GR 86%

Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models

Shengqu Cai, Duygu Ceylan, Matheus Gadelha, Chun-Hao Paul Huang, Tuanfeng Yang Wang, Gordon Wetzstein

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project page: https://primecai.github.io/generative_rendering/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00398 2023-09-08 cs.CV cs.MM 86%

VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Xin Li, Wenqing Chu, Ye Wu, Weihang Yuan, Fanglong Liu, Qi Zhang, Fu Li, Haocheng Feng, Errui Ding, Jingdong Wang

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

Comments 8pages, 8figures, project page: https://videogen.github.io/VideoGen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07605 2023-08-16 cs.CV cs.AI cs.MM 86%

SGDiff: A Style Guided Diffusion Model for Fashion Synthesis

Zhengwentai Sun, Yanghong Zhou, Honghong He, P. Y. Mok

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM'23

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.03099 2022-12-07 cs.CV cs.CL cs.MM 86%

Semantic-Conditional Diffusion Networks for Image Captioning

Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao, Jianlin Feng, Hongyang Chao, Tao Mei

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

Comments Source code is available at \url{https://github.com/YehLi/xmodaler/tree/master/configs/image_caption/scdnet}

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11319 2022-11-22 cs.CV cs.AI cs.GR cs.LG 86%

VectorFusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Models

Ajay Jain, Amber Xie, Pieter Abbeel

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project webpage: https://ajayj.com/vectorfusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15185 2025-03-18 cs.CV 86%

Tiled Diffusion

Or Madar, Ohad Fried

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract);分类 cs.CV

Comments Please visit our website for more information and the code: https://madaror.github.io/tiled-diffusion.github.io/

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20971 2025-01-14 cs.LG cs.CV 86%

Amortizing intractable inference in diffusion models for vision, language, and control

Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Nikolay Malkin

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments NeurIPS 2024; code: https://github.com/GFNOrg/diffusion-finetuning

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20474 2024-11-04 cs.CV 86%

GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

Phillip Y. Lee, Taehoon Yoon, Minhyuk Sung

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2024. Project Page: https://groundit-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01954 2024-06-17 cs.CV 86%

Plug-and-Play Diffusion Distillation

Yi-Ting Hsiao, Siavash Khodadadeh, Kevin Duarte, Wei-An Lin, Hui Qu, Mingi Kwon, Ratheesh Kalarot

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 project page: https://5410tiffany.github.io/plug-and-play-diffusion-distillation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05135 2024-03-11 cs.CV 86%

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, Gang Yu

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments Project Page: https://ella-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16203 2023-09-14 cs.LG cs.AI cs.CV cs.NE cs.RO 86%

Your Diffusion Model is Secretly a Zero-Shot Classifier

Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, Deepak Pathak

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments In ICCV 2023. Website at https://diffusion-classifier.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏