arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4200 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4200 篇

1909.07083 2019-12-20 cs.CV cs.CL cs.LG 88%

Controllable Text-to-Image Generation

Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. Torr

专题命中 可控生成 :image generation(title,abstract);text-to-image(title,abstract);分类 cs.CV

Comments NeurIPS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.11459 2019-07-22 cs.CV 88%

Coordinate-based Texture Inpainting for Pose-Guided Image Generation

Artur Grigorev, Artem Sevastopolsky, Alexander Vakhitov, Victor Lempitsky

专题命中 可控生成 :inpainting(title,abstract);image generation(title);image synthesis(abstract);分类 cs.CV

Comments Published in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11680 2026-04-14 cs.RO 88%

Dual-Control Frequency-Aware Diffusion Model for Depth-Dependent Optical Microrobot Microscopy Image Generation

双控频率感知扩散模型用于深度依赖光学微机器人显微图像生成

Lan Wei, Zongcai Tan, Kangyi Lu, Jian-Qing Zheng, Dandan Zhang

机构 * Department of Bioengineering, Imperial-X AI Initiative, Imperial College London(帝国理工学院生物工程系、Imperial-X人工智能计划) CAMS-Oxford Institute, University of Oxford(牛津大学CAMS-Oxford研究所)

专题命中 可控生成 :diffusion(title,abstract);image generation(title);image synthesis(abstract)

AI总结 本文提出Du-FreqNet,通过双控频率感知扩散模型生成物理一致的显微图像,提升微机器人自主操作的3D感知能力,改进SSIM并增强下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00345 2026-05-04 cs.CV 87%

Pose-Aware Diffusion for 3D Generation

姿态感知扩散用于3D生成

Zihan Zhou, Luxi Chen, Jingzhi Zhou, Yuhao Wan, Min Zhao, Baoyu Fan, Chongxuan Li

机构 * Gaoling School of AI, Renmin University of China(中国人民大学 Gallagher人工智能学院) VCIP, School of Computer Science, Nankai University(南开大学计算机学院 VCIP) THU-Bosch MLCenter, Tsinghua University(清华大学 THU-Bosch 机器学习中心) Inspur Group(Inspur集团)

专题命中 可控生成 :diffusion(title,summary_cn);分类 cs.CV

AI总结 本文提出Pose-Aware Diffusion,通过直接在观测空间生成3D几何,解决姿态对齐难题,实现高保真姿态一致的3D资产生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11693 2025-11-18 cs.AI cs.CR cs.CV cs.LG 87%

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

Xin Zhao, Xiaojun Chen, Bingshan Liu, Zeyao Liu, Zhendong Zhao, Xiaoyan Gu

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);generative vision(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12919 2025-08-19 cs.CV 87%

7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models

Elena Izzo, Luca Parolari, Davide Vezzaro, Lamberto Ballan

机构 * Department of Mathematics, University of Padova(帕多瓦大学数学系) Brain Technologies Innovation Division, Brain Technologies srl(Brain Technologies创新部门)

专题命中 可控生成 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)

Comments Accepted to ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10424 2025-08-15 cs.CV 87%

NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

Shanyuan Liu, Jian Zhu, Junda Lu, Yue Gong, Liuzhuozheng Li, Bo Cheng, Yuhang Ma, Liebucha Wu, Xiaoyu Wu, Dawei Leng, Yuhui Yin

机构 * AI Research(360人工智能研究院) Nanjing University of Science and Technology(南京理工大学) University of Science and Technology Beijing(北京科技大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学)

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04151 2025-07-08 cs.CV 87%

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

Fernando Gabriela Garcia, Spencer Burns, Ryan Shaw, Hunter Young

机构 * Autonomous University of Nuevo León(新莱昂自治大学)

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00895 2025-03-21 cs.CV 87%

Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model

Chenyang Liu, Keyan Chen, Rui Zhao, Zhengxia Zou, Zhenwei Shi

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06419 2025-03-11 cs.CV 87%

Consistent Image Layout Editing with Diffusion Models

Tao Xia, Yudi Zhang, Ting Liu Lei Zhang

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01097 2025-01-31 cs.CV 87%

EliGen: Entity-Level Controlled Image Generation with Regional Attention

Hong Zhang, Zhongjie Duan, Xingjun Wang, Yingda Chen, Yu Zhang

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);inpainting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01284 2024-12-19 cs.CV cs.AI 87%

MFTF: Mask-free Training-free Object Level Layout Control Diffusion Model

Shan Yang

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05252 2024-01-11 cs.CV 87%

PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models

Junsong Chen, Yue Wu, Simian Luo, Enze Xie, Sayak Paul, Ping Luo, Hang Zhao, Zhenguo Li

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11392 2023-12-19 cs.CV 87%

SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing

Zeyinzi Jiang, Chaojie Mao, Yulin Pan, Zhen Han, Jingfeng Zhang

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04675 2024-05-09 cs.CV cs.GR 87%

TexControl: Sketch-Based Two-Stage Fashion Image Generation Using Diffusion Model

Yongming Zhang, Tianyu Zhang, Haoran Xie

专题命中 可控生成 :image generation(title,abstract);diffusion(title);分类 cs.CV、cs.GR

Comments 5 pages, 8 figures, accepted in NICOGRAPH International 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12046 2026-07-07 cs.CV cs.AI 版本更新 86%

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

通过结构化掩码的布局条件自回归文本到图像生成

Zirui Zheng, Takashi Isobe, Tong Shen, Xu Jia, Jianbin Zhao, Xiaomin Li, Mengmeng Ge, Baolu Li, Qinghe Wang, Haiwen Diao, Dong Li, Dong Zhou, Yunzhi Zhuge, Huchuan Lu, Emad Barsoum

机构 * Advanced Micro Devices Inc.(超威半导体公司) Dalian University of Technology(大连理工大学) S-Lab, Nanyang Technological University(南洋理工大学S-Lab)

专题命中 可控生成 :image generation(title,abstract);text-to-image(title);分类 cs.CV

AI总结 提出SMARLI框架,通过结构化掩码策略将布局约束注入自回归生成过程,并采用GRPO后训练缓解曝光偏差,实现高质量布局控制图像生成。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00121 2026-06-02 cs.CV cs.AI 86%

Versatile Framework with Semantic and Structural guidance for Image Reconstruction from Brain Activity

基于语义和结构引导的大脑活动图像重建通用框架

Yizhuo Lu, Changde Du, Qiongyi Zhou, Liuyun Jiang, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学未来技术学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 可控生成 :diffusion(summary_cn,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出MindDiffuser两阶段框架,结合CLIP文本嵌入和视觉特征,通过Stable Diffusion生成语义图像并迭代优化结构信息,在fMRI、EEG、MEG三种模态上显著提升图像重建性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19348 2026-02-24 cs.CV cs.AI 86%

MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose

MultiDiffSense: 基于扩散的多模态视觉-触觉图像生成,基于物体形状和接触姿态

Sirine Bhouri, Lan Wei, Jian-Qing Zheng, Dandan Zhang

机构 * Department of Bioengineering, Imperial-X Initiative, Imperial College London(生物工程系、Imperial-X计划、帝国理工学院伦敦分校) CAMS-Oxford Institute, University of Oxford(CAMS-牛津研究所、牛津大学)

专题命中 可控生成 :diffusion(title,abstract);image generation(title);分类 cs.CV

AI总结 MultiDiffSense是一种基于扩散的多模态视觉-触觉图像生成模型,通过双条件化实现可控且物理一致的多模态生成,提升了触觉传感数据集的生成效率和跨模态学习能力。

Comments Accepted by 2026 ICRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17141 2025-10-06 cs.CV 86%

Filter-Guided Diffusion for Controllable Image Generation

Zeqi Gu, Ethan Yang, Abe Davis

机构 * Cornell University(康奈尔大学)

专题命中 可控生成 :diffusion(title,abstract);image generation(title);分类 cs.CV

Comments First two listed authors have equal contribution. The latest version has been accepted to SIGGRAPH 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14906 2025-09-04 eess.IV cs.CV 86%

FetalFlex: Anatomy-Guided Diffusion Model for Flexible Control on Fetal Ultrasound Image Synthesis

Yaofei Duan, Tao Tan, Zhiyuan Zhu, Yuhao Huang, Yuanji Zhang, Rui Gao, Patrick Cheong-Iao Pang, Xinru Gao, Guowei Tao, Xiang Cong, Zhou Li, Lianying Liang, Guangzhi He, Linliang Yin, Xuedong Deng, Xin Yang, Dong Ni

机构 * organization= Faculty of Applied Sciences, Macao Polytechnic University , city= Macao , country= China organization= National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University , city= Shenzhen , state= Guangdong , country= China organization= Shenzhen RayShape Medical Technology Co., Ltd , city= Shenzhen , state= Guangdong , country= China organization= Department of Ultrasound, Shenzhen Guangming District People’s Hospital , city= Shenzhen , state= Guangdong , country= China organization= Jinan University , city= Guangzhou , state= Guangdong , country= China organization= Center for Medical Ultrasound, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University , city= Suzhou , state= Jiangsu , country= China organization= Medical Ultrasound Image Computing (MUSIC) Laboratory, Shenzhen University , city= Shenzhen , state= Guangdong , country= China organization= Qilu Hospital of Shandong University , city= Jinan , state= Shandong , country= China organization= Northwest Women \& Children Hospital , city= Xian , state= Shaanxi , country= China

专题命中 可控生成 :diffusion(title);image synthesis(title);image generation(abstract);分类 cs.CV

Comments 18 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11904 2025-08-19 cs.CV 86%

SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress

Lingyun Zhang, Yu Xie, Yanwei Fu, Ping Chen

专题命中 可控生成 :text-to-image(title,abstract);diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14779 2025-06-04 cs.CV 86%

DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models

Hongji Yang, Wencheng Han, Yucheng Zhou, Jianbing Shen

机构 * University of Macau(澳门大学)

专题命中 可控生成 :image generation(title,abstract);diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10451 2025-02-21 cs.LG cs.GR 86%

FlexControl: Computation-Aware ControlNet with Differentiable Router for Text-to-Image Generation

Zheng Fang, Lichuan Xiang, Xu Cai, Kaicheng Zhou, Hongkai Wen

专题命中 可控生成 :image generation(title);text-to-image(title);diffusion(abstract);分类 cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05833 2024-12-10 cs.CV eess.IV 86%

CSG: A Context-Semantic Guided Diffusion Approach in De Novo Musculoskeletal Ultrasound Image Generation

Elay Dahan, Hedda Cohen Indelman, Angeles M. Perez-Agosto, Carmit Shiran, Gopal Avinash, Doron Shaked, Nati Daniel

专题命中 可控生成 :image generation(title,abstract);diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12850 2024-10-10 cs.CV cs.CR cs.LG 86%

PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware Pretraining

Kecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao, Xinwen Hou, Tianhao Wang

专题命中 可控生成 :image generation(title);diffusion(title);image synthesis(abstract);分类 cs.CV

Comments Accepted at USENIX Security 2024. The first two authors contributed equally. We communicated with the author of DPSDA and have added an explanation for why the FID scores in the DPSDA table are lower than those reported in the original paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09048 2024-01-18 cs.CV 86%

Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis

Jonghyun Lee, Hansam Cho, Youngjoon Yoo, Seoung Bum Kim, Yonghyun Jeong

专题命中 可控生成 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00986 2023-06-13 cs.CV cs.LG stat.ML 86%

Diffusion Self-Guidance for Controllable Image Generation

Dave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros, Aleksander Holynski

专题命中 可控生成 :diffusion(title,abstract);image generation(title);分类 cs.CV

Comments Project page at https://dave.ml/selfguidance/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22945 2026-06-23 cs.GR cs.CV 新提交 86%

Controllable Texture Tiling with Transformed RoPE-Enhanced Diffusion Models

可控纹理平铺:基于变换RoPE增强扩散模型

Junrong Huang, Zhiyuan Zhang, Rui Tang, Hongbo Fu, Jnig Liao

机构 * City University of Hong Kong(香港城市大学) Manycore Tech Inc.(Manycore科技公司) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 可控生成 :diffusion(title,abstract);image editing(abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 提出基于扩散变换器的可控纹理平铺框架,通过坐标变换旋转位置编码实现精确平铺控制,并利用分离注意力掩码保持参考纹理结构,同时与场景光照几何无缝融合。

Comments The code and dataset are publicly accessible at https://github.com/junrongh/ControlTile

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09864 2025-12-01 cs.GR cs.CV cs.LG 86%

Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion Models

利用语义属性绑定实现扩散模型中的免费午餐颜色控制

Héctor Laria, Alexandra Gomez-Villa, Jiang Qin, Muhammad Atif Butt, Bogdan Raducanu, Javier Vazquez-Corral, Joost van de Weijer, Kai Wang

机构 * Computer Vision Center, Spain(西班牙计算机视觉中心) Universitat Autònoma de Barcelona(巴塞罗那自治大学) Harbin Institute of Technology(哈尔滨工业大学) City University of Hong Kong(香港城市大学) Program of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学项目)

专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 ColorWave通过语义属性绑定实现扩散模型中精确的颜色控制,无需微调,提升颜色指定的准确性和适用性。

Comments WACV 2026. Project page: https://hecoding.github.io/colorwave-page

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01424 2025-01-03 cs.CV cs.AI cs.GR 86%

Object-level Visual Prompts for Compositional Image Generation

Gaurav Parmar, Or Patashnik, Kuan-Chieh Wang, Daniil Ostashev, Srinivasa Narasimhan, Jun-Yan Zhu, Daniel Cohen-Or, Kfir Aberman

专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments Project: https://snap-research.github.io/visual-composer/

详情

展开后加载摘要…

URL PDF HTML 收藏