arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 69989 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 69989 篇

2411.15580 2025-08-05 cs.CV 83%

TKG-DM: Training-free Chroma Key Content Generation Diffusion Model

Ryugo Morita, Stanislav Frolov, Brian Bernhard Moser, Takahiro Shirakawa, Ko Watanabe, Andreas Dengel, Jinjia Zhou

机构 * Faculty of Science and Engineering(科学与工程学部) RPTU Kaiserslautern-Landau & DFKI GmbH(科隆-兰道大学与DFKI GmbH)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted to CVPR2025(Highlight). Code at: https://github.com/ryugo417/TKG-DM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18092 2025-08-05 cs.CV cs.AI cs.RO 83%

DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models

Helin Cao, Sven Behnke

机构 * Autonomous Intelligent Systems group, Computer Science Institute VI – Intelligent Systems and Robotics – and the Center for Robotics and the Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn, Germany(自主智能系统组,计算机科学研究所VI——智能系统与机器人——和机器人中心以及拉马尔人工智能与机器学习研究所,波恩大学,德国)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025), Hangzhou, China, Oct 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01130 2025-08-05 cs.CV 83%

Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models

Bicheng Xu, Qi Yan, Renjie Liao, Lele Wang, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00438 2025-08-04 eess.IV cs.CV 83%

Diffusion-Based User-Guided Data Augmentation for Coronary Stenosis Detection

Sumin Seo, In Kyu Lee, Hyun-Woo Kim, Jaesik Min, Chung-Hwan Jung

机构 * Medipixel, Inc.(Medipixel公司) University of California San Diego(加州大学圣地亚哥分校)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted at MICCAI 2025. Dataset available at https://github.com/medipixel/DiGDA

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00413 2025-08-04 cs.CV cs.AI 83%

DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space

Junyu Chen, Dongyun Zou, Wenkun He, Junsong Chen, Enze Xie, Song Han, Han Cai

机构 * NVIDIA

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05846 2025-08-01 cs.CR cs.CV 83%

An Inversion-based Measure of Memorization for Diffusion Models

Zhe Ma, Qingming Li, Xuhong Zhang, Tianyu Du, Ruixiao Lin, Zonghui Wang, Shouling Ji, Wenzhi Chen

机构 * Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01654 2025-08-01 cs.LG cs.CV stat.ML 83%

Insights into Closed-form IPM-GAN Discriminator Guidance for Diffusion Modeling

Aadithya Srikanth, Siddarth Asokan, Nishanth Shetty, Chandra Sekhar Seelamantula

机构 * Microsoft Research #9 VIGYAN(微软研究院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00983 2025-07-31 eess.IV cs.CV 83%

DMCIE: Diffusion Model with Concatenation of Inputs and Errors to Improve the Accuracy of the Segmentation of Brain Tumors in MRI Images

Sara Yavari, Rahul Nitin Pandya, Jacob Furst

机构 * School of Computing, DePaul University(计算学院,德保罗大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21195 2025-07-30 cs.CR cs.AI cs.MM 83%

MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models

Po-Yuan Mao, Cheng-Chang Tsai, Chun-Shien Lu

机构 * IIS, Academia Sinica(中国台湾“中央研究院”资讯研究所)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17226 2025-07-29 cs.CV 83%

DDB: Diffusion Driven Balancing to Address Spurious Correlations

Aryan Yazdan Parast, Basim Azam, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05367 2025-07-24 cs.CV 83%

Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards

Aakash Garg, Libing Zeng, Andrii Tsarov, Nima Khademi Kalantari

机构 * Texas A&M University(德克萨斯A&M大学) Leia Inc(Leia公司)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07986 2025-07-24 cs.CV 83%

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen, Wangmeng Zuo, Ziwei Liu, Kwan-Yee K. Wong

机构 * The University of Hong Kong(香港大学) Nanjing University(南京大学) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted by ICCV 2025; Project Page: https://vchitect.github.io/TACA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16579 2025-07-23 eess.IV cs.AI cs.CV 83%

Pyramid Hierarchical Masked Diffusion Model for Imaging Synthesis

Xiaojiao Xiao, Qinmin Vivian Hu, Guanghui Wang

机构 * Department of Computer Science(计算机科学系) Toronto Metropolitan University(多伦多 Metropolitan 大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01738 2025-07-23 cs.CV cs.AI 83%

VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models

Kailai Feng, Yabo Zhang, Haodong Yu, Zhilong Ji, Jinfeng Bai, Hongzhi Zhang, Wangmeng Zuo

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments https://github.com/Carlofkl/VitaGlyph

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15361 2025-07-22 eess.IV cs.AI cs.CV 83%

Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation

Muhammad Aqeel, Maham Nazir, Zanxi Ruan, Francesco Setti

机构 * Dept. of Engineering for Innovation Medicine, University of Verona(创新医学工程系,威尼斯大学) Department of Computer Science, Beihang University(计算机科学系,北航大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted to CVGMMI Workshop at ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15216 2025-07-22 cs.CV 83%

Improving Joint Embedding Predictive Architecture with Diffusion Noise

Yuping Qiu, Rui Zhu, Ying-cong Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14797 2025-07-22 cs.CV 83%

Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models

Beier Zhu, Ruoyu Wang, Tong Zhao, Hanwang Zhang, Chi Zhang

机构 * Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments To appear in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14575 2025-07-22 cs.CV cs.AI 83%

Benchmarking GANs, Diffusion Models, and Flow Matching for T1w-to-T2w MRI Translation

Andrea Moschetto, Lemuel Puglisi, Alec Sargood, Pierluigi Dell'Acqua, Francesco Guarnera, Sebastiano Battiato, Daniele Ravì

机构 * Università degli Studi di Catania(卡塔尼亚大学) Università degli Studi di Messina(梅斯纳大学) University College London(伦敦大学学院)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12170 2025-07-21 cs.RO cs.CV 83%

DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving

Tao Wang, Cong Zhang, Xingguang Qu, Kun Li, Weiwei Liu, Chang Huang

机构 * Carizon Beihang University(北航大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 8 pages, 6 figures; Code released

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06848 2025-07-21 cs.LG cs.CL cs.CV 83%

A General Framework for Inference-time Scaling and Steering of Diffusion Models

Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, Rajesh Ranganath

机构 * Department of Computer Science, New York University(纽约大学计算机科学系) Columbia University(哥伦比亚大学) Center for Data Science, New York University(纽约大学数据科学中心)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12933 2025-07-18 cs.CV cs.AI cs.LG 83%

DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

Dongyeun Lee, Jiwan Hur, Hyounguk Shon, Jae Young Lee, Junmo Kim

机构 * KAIST(韩国科学技术院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03558 2025-07-18 cs.CV 83%

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Zehuan Huang, Yuan-Chen Guo, Xingqiao An, Yunhan Yang, Yangguang Li, Zi-Xin Zou, Ding Liang, Xihui Liu, Yan-Pei Cao, Lu Sheng

机构 * Beihang University(北京航空航天大学) VAST Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Project page: https://huanngzh.github.io/MIDI-Page/

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 23646 - 23657

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09502 2025-07-18 cs.LG cs.CV 83%

Golden Noise for Diffusion Models: A Learning Framework

Zikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang, Zhiqiang Xu, Bo Han, Zeke Xie

机构 * HKUST-GZ(香港科技大学-广州) Shanghai AI Lab(上海人工智能实验室) MBZUAI(穆罕默德·本·拉希德人工智能学院) HKBU(香港都会大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10745 2025-07-17 cs.CV 83%

Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition

Jeonghyeok Do, Munchurl Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments ICCV 2025 (camera-ready version). Please visit our project page at https://kaist-viclab.github.io/TDSM_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04692 2025-07-15 cs.CV 83%

Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow Removal

Wanchang Yu, Qing Zhang, Rongjia Zheng, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(教育部人工智能与先进计算重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07878 2025-07-14 cs.CV 83%

Single-Step Latent Diffusion for Underwater Image Restoration

Jiayi Wu, Tianfu Wang, Md Abu Bakr Siddique, Md Jahidul Islam, Cornelia Fermuller, Yiannis Aloimonos, Christopher A. Metzler

机构 * University of Maryland(马里兰大学) University of Florida(佛罗里达大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07104 2025-07-14 cs.CV 83%

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

Tiezheng Zhang, Yitong Li, Yu-cheng Chou, Jieneng Chen, Alan Yuille, Chen Wei, Junfei Xiao

机构 * Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学) Rice University(Rice 大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Project Page: https://lambert-x.github.io/Vision-Language-Vision/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07591 2025-07-11 cs.CV 83%

Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model

Kuiyuan Sun, Yuxuan Zhang, Jichao Zhang, Jiaming Liu, Wei Wang, Niculae Sebe, Yao Zhao

机构 * Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院) School of Computer Science, Ocean University of China(中国海洋大学计算机学院) Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Artificial Intelligence department, Tiamat AI(Tiamat AI人工智能部门) Department of Information Engineering and Computer Science (DISI), University of Trento(特伦托大学信息工程与计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09932 2025-07-11 cs.CV cs.AI 83%

HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations

Marco Federici, Riccardo Del Chiaro, Boris van Breugel, Paul Whatmough, Markus Nagel

机构 * Qualcomm(高通)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 8 Pages, 6 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07106 2025-07-10 cs.CV cs.LG 83%

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

Vatsal Agarwal, Matthew Gwilliam, Gefen Kohavi, Eshan Verma, Daniel Ulbricht, Abhinav Shrivastava

机构 * University of Maryland(马里兰大学) Apple(苹果公司)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Website: see https://vatsalag99.github.io/mustafar/

详情

展开后加载摘要…

URL PDF HTML 收藏