arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 69953 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 69953 篇

2507.09441 2025-12-12 cs.GR cs.CV 84%

RectifiedHR: High-Resolution Diffusion via Energy Profiling and Adaptive Guidance Scheduling

RectifiedHR: 通过能量分析和自适应引导调度实现高分辨率扩散

Ankit Sanjyal

机构 * Fordham University(福特汉姆大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 RectifiedHR通过能量分析和自适应引导调度提升高分辨率扩散模型的稳定性和图像质量。

Comments 8 Pages, 10 Figures, Pre-Print Version, This version is under review for citation accuracy

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05198 2025-12-08 cs.CV cs.GR cs.LG 84%

Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models

你的潜在掩码是错的:用于扩散模型的像素等效潜在合成

Rowan Bradbury, Dazhi Zhong

机构 * Bradbury Group(布拉德伯格集团)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出PELC原则,通过DecFormer实现像素等效潜在合成,提升扩散模型在补全任务中的性能与保真度。

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08530 2025-10-10 cs.GR cs.CV 84%

X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering

Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang, Hao Zhu, Jing Liao

机构 * City University of Hong Kong(香港城市大学) WeChat, Tencent Inc(微信、腾讯公司) Manycore Tech Inc(很多核科技公司)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR

Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24369 2025-09-30 cs.CV cs.AI cs.MM 84%

From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis

Khawlah Bajbaa, Abbas Anwar, Muhammad Saqib, Hafeez Anwar, Nabin Sharma, Muhammad Usman

机构 * Department of Information and Computer Science, King Fahd University of Petroleum and Minerals(信息与计算机科学系,国王法赫德石油与矿物大学) NCMI, CSIRO(CSIRO国家科学机构) Department of Computer Science, National University of Computer and Emerging Sciences (FAST-NUCES)(计算机科学系,国家计算机与新兴科学大学(FAST-NUCES)) Faculty of Engineering and IT, University of Technology Sydney(工程与信息技术学院,技术大学悉尼) Faculty of Science, University of Ontario Institute of Technology(科学学院, Ontario Institute of Technology大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21887 2025-09-29 cs.CV cs.MM 84%

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He, Boshi Tang, Zhiyong Wu, Yang Huang, Yang Wu, Zhongqian Sun, Wei Yang, Helen Meng

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03680 2025-09-05 cs.GR cs.AI cs.CV 84%

LuxDiT: Lighting Estimation with Video Diffusion Transformer

Ruofan Liang, Kai He, Zan Gojcic, Igor Gilitschenski, Sanja Fidler, Nandita Vijaykumar, Zian Wang

机构 * NVIDIA University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://research.nvidia.com/labs/toronto-ai/LuxDiT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16827 2025-08-26 cs.GR cs.CV cs.LG 84%

Beyond Blur: A Fluid Perspective on Generative Diffusion Models

Grzegorz Gruszczynski, Jakub Meixner, Michal Jan Wlodarczyk, Przemyslaw Musialski

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments ICCV 2025 main conference, 8 pages paper, 20 pages appendix, 24 figures, supplementary pseudocode in appendix, https://iccv.thecvf.com/virtual/2025/poster/1176

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08384 2025-08-13 cs.GR cs.AI cs.CV 84%

Spatiotemporally Consistent Indoor Lighting Estimation with Diffusion Priors

Mutian Tong, Rundi Wu, Changxi Zheng

机构 * Columbia University(哥伦比亚大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments 11 pages. Accepted by SIGGRAPH 2025 as Conference Paper

Journal ref SIGGRAPH '25: ACM SIGGRAPH 2025 Conference Conference Papers, Article 107, pages1-11, July 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19853 2025-08-05 cs.CV cs.GR cs.LG 84%

Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation

Nadav Z. Cohen, Oron Nir, Ariel Shamir

机构 * Reichman University(里奇曼大学) Microsoft Corporation(微软公司)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments Conference paper at CVPR 2025. Project page: https://nadavc220.github.io/conditional-balance.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15399 2025-07-22 cs.GR cs.CV 84%

Blended Point Cloud Diffusion for Localized Text-guided Shape Editing

Etai Sella, Noam Atia, Ron Mokady, Hadar Averbuch-Elor

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments Accepted to ICCV 2025. Project Page: https://tau-vailab.github.io/BlendedPC/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16977 2025-05-23 cs.CV cs.MM 84%

Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On

Siqi Wan, Jingwen Chen, Yingwei Pan, Ting Yao, Tao Mei

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.MM

Comments ICLR 2025. Code is publicly available at: https://github.com/HiDream-ai/SPM-Diff

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10558 2025-05-16 cs.GR cs.CV 84%

Style Customization of Text-to-Vector Generation with Image Diffusion Priors

Peiying Zhang, Nanxuan Zhao, Jing Liao

机构 * City University of Hong Kong(香港城市大学) Adobe Research(Adobe研究)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted by SIGGRAPH 2025 (Conference Paper). Project page: https://customsvg.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16938 2025-04-14 cs.CV cs.AI cs.GR 84%

Generative Object Insertion in Gaussian Splatting with a Multi-View Diffusion Model

Hongliang Zhong, Can Wang, Jingbo Zhang, Jing Liao

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments Accepted by Visual Informatics. Project Page: https://github.com/JiuTongBro/MultiView_Inpaint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21694 2025-03-28 cs.GR cs.AI cs.CV 84%

Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data

Zhiyuan Ma, Xinyue Liang, Rongyuan Wu, Xiangyu Zhu, Zhen Lei, Lei Zhang

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted to CVPR 2025. Code:https://github.com/theEricMa/TriplaneTurbo. Demo:https://huggingface.co/spaces/ZhiyuanthePony/TriplaneTurbo

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18590 2025-03-25 cs.CV cs.GR 84%

DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models

Ruofan Liang, Zan Gojcic, Huan Ling, Jacob Munkberg, Jon Hasselgren, Zhi-Hao Lin, Jun Gao, Alexander Keller, Nandita Vijaykumar, Sanja Fidler, Zian Wang

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.GR

Comments CVPR 2025; project page: research.nvidia.com/labs/toronto-ai/DiffusionRenderer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04776 2025-03-10 cs.GR cond-mat.mtrl-sci cs.CV cs.LG 84%

GrainPaint: A multi-scale diffusion-based generative model for microstructure reconstruction of large-scale objects

Nathan Hoffman, Cashen Diniz, Dehao Liu, Theron Rodgers, Anh Tran, Mark Fuge

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04378 2025-02-10 cs.CV cs.GR cs.LG cs.SE 84%

DILLEMA: Diffusion and Large Language Models for Multi-Modal Augmentation

Luciano Baresi, Davide Yi Xian Hu, Muhammad Irfan Mas'udi, Giovanni Quattrocchi

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02225 2025-02-05 cs.CV cs.AI cs.MM 84%

Exploring the latent space of diffusion models directly through singular value decomposition

Li Wang, Boyan Gao, Yanran Li, Zhao Wang, Xiaosong Yang, David A. Clifton, Jun Xiao

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18390 2024-12-30 cs.CV cs.AI cs.LG cs.MM 84%

RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction

Xiaoping Wu, Jie Hu, Xiaoming Wei

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05694 2024-12-10 cs.MM cs.GR cs.SD eess.AS 84%

Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation

Leonardo Pina, Yongmin Li

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.GR、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14740 2024-11-25 cs.CV cs.AI cs.GR 84%

TEXGen: a Generative Diffusion Model for Mesh Textures

Xin Yu, Ze Yuan, Yuan-Chen Guo, Ying-Tian Liu, JianHui Liu, Yangguang Li, Yan-Pei Cao, Ding Liang, Xiaojuan Qi

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments Accepted to SIGGRAPH Asia Journal Article (TOG 2024)

Journal ref ACM Transactions on Graphics (TOG) 2024, Volume 43, Issue 6, Article No.: 213, Pages 1-14

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07232 2024-11-13 cs.CV cs.AI cs.GR cs.LG 84%

Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

Yoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon, Lior Wolf, Gal Chechik

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.GR

Comments Project page is at https://research.nvidia.com/labs/par/addit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19989 2024-10-01 cs.CV cs.GR 84%

RoCoTex: A Robust Method for Consistent Texture Synthesis with Diffusion Models

Jangyeong Kim, Donggoo Kang, Junyoung Choi, Jeonga Wi, Junho Gwon, Jiun Bae, Dumim Yoon, Junghyun Han

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments 11 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08258 2024-09-13 cs.CV cs.MM 84%

Improving Virtual Try-On with Garment-focused Diffusion Models

Siqi Wan, Yehao Li, Jingwen Chen, Yingwei Pan, Ting Yao, Yang Cao, Tao Mei

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

Comments ECCV 2024. Source code is available at https://github.com/siqi0905/GarDiff/tree/master

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07452 2024-09-12 cs.CV cs.MM 84%

Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models

Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao, Zhineng Chen, Chong-Wah Ngo, Tao Mei

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2024. Source code is available at \url{https://github.com/yanghb22-fdu/Hi3D-Official}

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07447 2024-09-12 cs.CV cs.GR 84%

StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos

Sijie Zhao, Wenbo Hu, Xiaodong Cun, Yong Zhang, Xiaoyu Li, Zhe Kong, Xiangjun Gao, Muyao Niu, Ying Shan

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments 11 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14826 2024-08-28 cs.CV cs.MM 84%

Alfie: Democratising RGBA Image Generation With No $$$

Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Rita Cucchiara

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.MM

Comments Accepted at ECCV AI for Visual Arts Workshop and Challenges

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09702 2024-08-20 cs.CV cs.AI cs.GR 84%

Photorealistic Object Insertion with Diffusion-Guided Inverse Rendering

Ruofan Liang, Zan Gojcic, Merlin Nimier-David, David Acuna, Nandita Vijaykumar, Sanja Fidler, Zian Wang

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments ECCV 2024, Project page: https://research.nvidia.com/labs/toronto-ai/DiPIR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13220 2024-08-12 cs.CV cs.GR cs.LG 84%

TetraDiffusion: Tetrahedral Diffusion Models for 3D Shape Generation

Nikolai Kalischek, Torben Peters, Jan D. Wegner, Konrad Schindler

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR

Comments This version introduces further improvements compared to v2. Project page https://tetradiffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03178 2024-08-07 cs.CV cs.GR cs.LG 84%

An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion

Xingguang Yan, Han-Hung Lee, Ziyu Wan, Angel X. Chang

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR

Comments Project Page: https://omages.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏