arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4201 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4201 篇

2311.12342 2024-03-27 cs.CV 85%

LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis

Peiang Zhao, Han Li, Ruiyang Jin, S. Kevin Zhou

专题命中 可控生成 :image synthesis(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

Comments Demo: https://huggingface.co/spaces/Pusheen/LoCo; Project page: https://momopusheen.github.io/LoCo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11878 2024-03-19 cs.CV 85%

InTeX: Interactive Text-to-texture Synthesis via Unified Depth-aware Inpainting

Jiaxiang Tang, Ruijie Lu, Xiaokang Chen, Xiang Wen, Gang Zeng, Ziwei Liu

专题命中 可控生成 :inpainting(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

Comments Project Page: https://me.kiui.moe/intex/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13314 2024-01-09 cs.CV cs.AI cs.LG 85%

Unlocking Pre-trained Image Backbones for Semantic Image Synthesis

Tariq Berrada, Jakob Verbeek, Camille Couprie, Karteek Alahari

专题命中 可控生成 :image synthesis(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09252 2023-12-15 cs.CV 85%

FineControlNet: Fine-level Text Control for Image Generation with Spatially Aligned Text Control Injection

Hongsuk Choi, Isaac Kasahara, Selim Engin, Moritz Graule, Nikhil Chavan-Dafle, Volkan Isler

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

Comments Hongsuk Choi and Isaac Kasahara have eqaul contributions. 19 pages, 15 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11147 2023-11-03 cs.CV cs.AI 85%

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, Stefano Ermon, Yun Fu, Ran Xu

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06700 2023-10-27 cs.CV cs.LG 85%

Control3Diff: Learning Controllable 3D Diffusion Models from Single-view Images

Jiatao Gu, Qingzhe Gao, Shuangfei Zhai, Baoquan Chen, Lingjie Liu, Josh Susskind

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract);分类 cs.CV

Comments Accepted by 3DV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13078 2023-06-23 cs.CV 85%

Continuous Layout Editing of Single Images with Diffusion Models

Zhiyuan Zhang, Zhitong Huang, Jing Liao

专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21116 2026-05-21 eess.IV 85%

GeoDiff-SAR II: 3D-Driven Foundation Diffusion Models for SAR Generation via Decoupled Control

GeoDiff-SAR II: 3D-Driven Foundation Diffusion Models for SAR Generation via Decoupled Control

Xuanting Wu, Fan Zhang, Fei Ma, Yingbing Liu, Lingxiao Peng, Qiang Yin, Yongsheng Zhou

专题命中 可控生成 :diffusion(title,title_cn);image generation(abstract)

AI总结 本文提出GeoDiff-SAR II,一种基于3D模型引导的解耦框架,用于通过解耦控制生成合成孔径雷达图像,通过物理基础的几何-电磁线索实现对关键成像参数的可控生成。

Comments 23 pages,14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05730 2026-04-08 cs.LG 85%

Controllable Image Generation with Composed Parallel Token Prediction

通过组合并行令牌预测实现可控图像生成

Jamie Stirling, Noura Al-Moubayed, Chris G. Willcocks, Hubert P. H. Shum

机构 * Durham University(杜伦大学)

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract)

AI总结 本文提出一种理论指导的离散生成过程组合方法,通过掩码生成(吸收扩散)实现多输入条件的精确组合,相比现有方法在误差率和FID上均有显著提升,并实现高效的文本到图像生成控制。

Comments 8 pages + references, 7 figures, accepted to CVPR Workshops 2026 (LoViF). arXiv admin note: substantial text overlap with arXiv:2405.06535

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22421 2025-12-05 cs.NI 85%

Semantic-Aware Caching for Efficient Image Generation in Edge Computing

面向边缘计算的语义感知缓存用于高效图像生成

Hanshuai Cui, Zhiqing Tang, Zhi Yao, Weijia Jia, Wei Zhao

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract)

AI总结 CacheGenius通过语义感知缓存和调度算法,在边缘计算中高效生成图像,减少41%的延迟和48%的计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17609 2024-07-16 cs.CV cs.GR cs.LG 84%

Curved Diffusion: A Generative Model With Optical Geometry Control

Andrey Voynov, Amir Hertz, Moab Arar, Shlomi Fruchter, Daniel Cohen-Or

专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project page at https://anylens-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03152 2024-01-09 cs.CV cs.LG 84%

Controllable Image Synthesis of Industrial Data Using Stable Diffusion

Gabriele Valvano, Antonino Agostino, Giovanni De Magistris, Antonino Graziano, Giacomo Veneri

专题命中 可控生成 :diffusion(title);image synthesis(title);分类 cs.CV

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2024, pp. 5354-5363

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01571 2026-05-15 cs.GR cs.AI cs.CV cs.LG 84%

Pro-DG: Procedural Diffusion Guidance for Architectural Facade Generation

Pro-DG:基于过程扩散引导的建筑立面生成

Aleksander Plocharski, Jan Swidzinski, Przemyslaw Musialski

机构 * Warsaw University of Technology(华沙技术大学) Akces NCBR Imperial College London(伦敦帝国理工学院) New Jersey Institute of Technology(新泽西理工学院)

专题命中 可控生成 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出Pro-DG方法,通过层次化过程规则生成控制图,实现建筑立面的真实感生成,支持结构编辑如楼层复制和窗户重排,经评估验证其在保持建筑身份和精确编辑方面的优越性。

Comments 17 pages, 15 figures, Computer Graphics Forum 2026 Journal Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07871 2026-03-11 cs.CV cs.MM cs.SD eess.AS 84%

Controllable Dance Generation with Style-Guided Motion Diffusion

具有风格引导的舞动生成

Hongsong Wang, Ying Zhu, Xin Geng, Liang Wang

专题命中 可控生成 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.MM

AI总结 本文提出Style-Guided Motion Diffusion方法,通过整合Transformer架构与风格调制模块,实现风格引导的舞蹈生成,提升舞蹈生成的可控性和风格一致性。

Journal ref Machine Intelligence Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03714 2026-03-05 cs.CL cs.AI cs.CV cs.MM 84%

Order Is Not Layout: Order-to-Space Bias in Image Generation

秩序并非布局:图像生成中的秩序到空间偏差

Yongkang Zhang, Zonglin Zhao, Yuechen Zhang, Fei Ding, Pei Li, Wenxuan Wang

机构 * Renmin University of China, China(中国人民大学) Huazhong Agricultural University, China(华中农业大学) Huazhong University of Science and Technology, China(华中科技大学) Jiangnan University, China(江南大学) Nanchang University, China(南昌大学)

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文研究了图像生成中因文本实体顺序导致的空间布局偏差问题,提出OTS-Bench进行量化评估,并通过微调和干预策略减少该偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16750 2025-11-19 cs.CV cs.CL cs.LG cs.MM 84%

Iris: Integrating Language into Diffusion-based Monocular Depth Estimation

Ziyao Zeng, Jingcheng Ni, Daniel Wang, Patrick Rim, Younjoon Chung, Fengyu Yang, Byung-Woo Hong, Alex Wong

机构 * Yale University(耶鲁大学) Brown University(布朗大学) Chung-Ang University(Chung-Ang 大学)

专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13434 2025-10-02 cs.CV cs.AI cs.MM 84%

BlobCtrl: Taming Controllable Blob for Element-level Image Editing

Yaowei Li, Lingen Li, Zhaoyang Zhang, Xiaoyu Li, Guangzhi Wang, Hongxiang Li, Xiaodong Cun, Ying Shan, Yuexian Zou

机构 * SECE, Peking University(北京大学SECE学院) The Chinese University of Hong Kong(香港中文大学) ARC Lab, Tencent(腾讯ARC实验室) The Hong Kong University of Science and Technology(香港科学与技术大学) GVC Lab, Great Bay University(Great Bay大学GVC实验室)

专题命中 可控生成 :image editing(title,abstract);diffusion(abstract);分类 cs.CV、cs.MM

Comments Project Webpage: https://liyaowei-stu.github.io/project/BlobCtrl/ This version presents a major update with rephrased writing. Accepted to SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00704 2025-07-22 cs.GR cs.CV 84%

Controllable Weather Synthesis and Removal with Video Diffusion Models

Chih-Hao Lin, Zian Wang, Ruofan Liang, Yuxuan Zhang, Sanja Fidler, Shenlong Wang, Zan Gojcic

机构 * NVIDIA University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 可控生成 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.GR

Comments International Conference on Computer Vision (ICCV) 2025, Project Website: https://research.nvidia.com/labs/toronto-ai/WeatherWeaver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03068 2025-03-06 cs.GR cs.CV 84%

Multi-View Depth Consistent Image Generation Using Generative AI Models: Application on Architectural Design of University Buildings

Xusheng Du, Ruihan Gui, Zhengyang Wang, Ye Zhang, Haoran Xie

专题命中 可控生成 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments 10 pages, 7 figures, in Proceedings of CAADRIA2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18598 2025-02-25 cs.CV cs.GR 84%

Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation

Tianyidan Xie, Rui Ma, Qian Wang, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang, Lanjun Wang, Zili Yi

专题命中 可控生成 :image generation(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments 18 pages, 15 figures, project page: https://anywheremultiagent.github.io, Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04191 2025-02-18 cs.CV cs.AI cs.GR 84%

GazeFusion: Saliency-Guided Image Generation

Yunxiang Zhang, Nan Wu, Connor Z. Lin, Gordon Wetzstein, Qi Sun

专题命中 可控生成 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments ACM Transactions on Applied Perception (ACM Symposium on Applied Perception 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00416 2024-08-19 cs.CV cs.AI cs.GR 84%

Interactive Character Control with Auto-Regressive Motion Diffusion Models

Yi Shi, Jingbo Wang, Xuekun Jiang, Bingkun Lin, Bo Dai, Xue Bin Peng

专题命中 可控生成 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19468 2024-07-30 cs.CV cs.MM 84%

MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability

Buyu Liu, Kai Wang, Yansong Liu, Jun Bao, Tingting Han, Jun Yu

专题命中 可控生成 :image generation(title);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM24

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08910 2021-04-20 cs.CV cs.MM 84%

Towards Open-World Text-Guided Face Image Generation and Manipulation

Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

Comments arXiv admin note: substantial text overlap with arXiv:2012.03308

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.03308 2021-03-30 cs.CV cs.AI cs.MM 84%

TediGAN: Text-Guided Diverse Face Image Generation and Manipulation

Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

Comments CVPR 2021. Code: https://github.com/weihaox/TediGAN Data: https://github.com/weihaox/Multi-Modal-CelebA-HQ Video: https://youtu.be/L8Na2f5viAM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11439 2025-04-08 cs.CV 84%

A Simple Approach to Unifying Diffusion-based Conditional Generation

Xirui Li, Charles Herrmann, Kelvin C. K. Chan, Yinxiao Li, Deqing Sun, Chao Ma, Ming-Hsuan Yang

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Project page: https://lixirui142.github.io/unicon-diffusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04287 2023-09-11 eess.SP cs.AI 83%

Sequential Semantic Generative Communication for Progressive Text-to-Image Generation

Hyelin Nam, Jihong Park, Jinho Choi, Seong-Lyun Kim

专题命中 可控生成 :image generation(title);text-to-image(title)

Comments 4 pages, 2 figures, to be published in IEEE International Conference on Sensing, Communication, and Networking, Workshop on Semantic Communication for 6G (SC6G-SECON23)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06751 2026-08-10 cs.CV cs.AI 新提交 83%

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

超越《星夜》:面向艺术家基础的文本到图像生成的捷径感知控制状态规划

Kuan Xing, Ye Wang, Changyi Gan, Yuheng Li, Thao Nguyen, Yi Chang, Yilin Wang

机构 * Jilin University(吉林大学) Adobe(奥多比公司) University of Wisconsin(威斯康星大学)

专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 该研究针对艺术家基础的文本到图像生成存在的捷径偏差问题,提出 Atelier 框架,结合 ArtIntentBench 基准测试,提升了风格保真度与结构保留度,减少了捷径替换。

Comments 47 pages, 13 figures, including appendices. Kuan Xing and Ye Wang contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00083 2026-08-04 cs.CV 新提交 83%

Beyond Edge Maps: Wavelet-Domain Conditioning for Multi-Adapter Map-to-Satellite Diffusion

超越边缘图:面向多适配器地图到卫星扩散的小波域条件控制

Arisha Prasain

机构 * Tribhuvan University(特里布文大学) Pulchowk Engineering Campus(普尔乔克工程校区)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 本文针对资源匮乏地区卫星底图缺失问题,提出基于OSM栅格地图和SWT子带的多适配器ControlNet扩散框架,在尼泊尔数据集和Pix2Pix基准上验证了其合成卫星图像的有效性。

Comments 10 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18267 2026-07-27 cs.CV 版本更新 83%

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation

SRC-Flow:紧凑语义表示实现归一化流用于图像生成

Longtao Jiang, Jianmin Bao, Zhendong Wang, Xin Tao, Pengfei Wan, Zhihui Li, Xiaojun Chang

机构 * University of Science and Technology of China(中国科学技术大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 可控生成 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 提出SRC-Flow,通过语义表示压缩器将高维RAE特征压缩到低维语义空间,降低归一化流建模负担,在ImageNet上实现最优生成质量,同时保持精确似然计算和确定性可逆采样。

详情

展开后加载摘要…

URL PDF HTML 收藏