Controllable Text-to-Image Generation
专题命中 可控生成 :image generation(title,abstract);text-to-image(title,abstract);分类 cs.CV
Comments NeurIPS 2019
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 可控生成 :image generation(title,abstract);text-to-image(title,abstract);分类 cs.CV
Comments NeurIPS 2019
专题命中 可控生成 :inpainting(title,abstract);image generation(title);image synthesis(abstract);分类 cs.CV
Comments Published in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2019
双控频率感知扩散模型用于深度依赖光学微机器人显微图像生成
机构 * Department of Bioengineering, Imperial-X AI Initiative, Imperial College London(帝国理工学院生物工程系、Imperial-X人工智能计划) ; CAMS-Oxford Institute, University of Oxford(牛津大学CAMS-Oxford研究所)
专题命中 可控生成 :diffusion(title,abstract);image generation(title);image synthesis(abstract)
AI总结 本文提出Du-FreqNet,通过双控频率感知扩散模型生成物理一致的显微图像,提升微机器人自主操作的3D感知能力,改进SSIM并增强下游任务性能。
姿态感知扩散用于3D生成
机构 * Gaoling School of AI, Renmin University of China(中国人民大学 Gallagher人工智能学院) ; VCIP, School of Computer Science, Nankai University(南开大学计算机学院 VCIP) ; THU-Bosch MLCenter, Tsinghua University(清华大学 THU-Bosch 机器学习中心) ; Inspur Group(Inspur集团)
专题命中 可控生成 :diffusion(title,summary_cn);分类 cs.CV
AI总结 本文提出Pose-Aware Diffusion,通过直接在观测空间生成3D几何,解决姿态对齐难题,实现高保真姿态一致的3D资产生成。
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);generative vision(abstract)
机构 * Department of Mathematics, University of Padova(帕多瓦大学数学系) ; Brain Technologies Innovation Division, Brain Technologies srl(Brain Technologies创新部门)
专题命中 可控生成 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)
Comments Accepted to ICIAP 2025
机构 * AI Research(360人工智能研究院) ; Nanjing University of Science and Technology(南京理工大学) ; University of Science and Technology Beijing(北京科技大学) ; Beijing University of Aeronautics and Astronautics(北京航空航天大学)
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image synthesis(abstract)
机构 * Autonomous University of Nuevo León(新莱昂自治大学)
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);inpainting(abstract)
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)
Comments 8 pages, 7 figures
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
Comments Technical Report
专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image synthesis(abstract)
专题命中 可控生成 :image generation(title,abstract);diffusion(title);分类 cs.CV、cs.GR
Comments 5 pages, 8 figures, accepted in NICOGRAPH International 2024
通过结构化掩码的布局条件自回归文本到图像生成
机构 * Advanced Micro Devices Inc.(超威半导体公司) ; Dalian University of Technology(大连理工大学) ; S-Lab, Nanyang Technological University(南洋理工大学S-Lab)
专题命中 可控生成 :image generation(title,abstract);text-to-image(title);分类 cs.CV
AI总结 提出SMARLI框架,通过结构化掩码策略将布局约束注入自回归生成过程,并采用GRPO后训练缓解曝光偏差,实现高质量布局控制图像生成。
Comments ECCV 2026
基于语义和结构引导的大脑活动图像重建通用框架
机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术国家重点实验室) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学未来技术学院) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中 可控生成 :diffusion(summary_cn,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 提出MindDiffuser两阶段框架,结合CLIP文本嵌入和视觉特征,通过Stable Diffusion生成语义图像并迭代优化结构信息,在fMRI、EEG、MEG三种模态上显著提升图像重建性能。
MultiDiffSense: 基于扩散的多模态视觉-触觉图像生成,基于物体形状和接触姿态
机构 * Department of Bioengineering, Imperial-X Initiative, Imperial College London(生物工程系、Imperial-X计划、帝国理工学院伦敦分校) ; CAMS-Oxford Institute, University of Oxford(CAMS-牛津研究所、牛津大学)
专题命中 可控生成 :diffusion(title,abstract);image generation(title);分类 cs.CV
AI总结 MultiDiffSense是一种基于扩散的多模态视觉-触觉图像生成模型,通过双条件化实现可控且物理一致的多模态生成,提升了触觉传感数据集的生成效率和跨模态学习能力。
Comments Accepted by 2026 ICRA
机构 * Cornell University(康奈尔大学)
专题命中 可控生成 :diffusion(title,abstract);image generation(title);分类 cs.CV
Comments First two listed authors have equal contribution. The latest version has been accepted to SIGGRAPH 2024
机构 * organization= Faculty of Applied Sciences, Macao Polytechnic University , city= Macao , country= China ; organization= National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University , city= Shenzhen , state= Guangdong , country= China ; organization= Shenzhen RayShape Medical Technology Co., Ltd , city= Shenzhen , state= Guangdong , country= China ; organization= Department of Ultrasound, Shenzhen Guangming District People’s Hospital , city= Shenzhen , state= Guangdong , country= China ; organization= Jinan University , city= Guangzhou , state= Guangdong , country= China ; organization= Center for Medical Ultrasound, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University , city= Suzhou , state= Jiangsu , country= China ; organization= Medical Ultrasound Image Computing (MUSIC) Laboratory, Shenzhen University , city= Shenzhen , state= Guangdong , country= China ; organization= Qilu Hospital of Shandong University , city= Jinan , state= Shandong , country= China ; organization= Northwest Women \& Children Hospital , city= Xian , state= Shaanxi , country= China
专题命中 可控生成 :diffusion(title);image synthesis(title);image generation(abstract);分类 cs.CV
Comments 18 pages, 10 figures
专题命中 可控生成 :text-to-image(title,abstract);diffusion(title);分类 cs.CV
机构 * University of Macau(澳门大学)
专题命中 可控生成 :image generation(title,abstract);diffusion(title);分类 cs.CV
专题命中 可控生成 :image generation(title);text-to-image(title);diffusion(abstract);分类 cs.GR
专题命中 可控生成 :image generation(title,abstract);diffusion(title);分类 cs.CV
专题命中 可控生成 :image generation(title);diffusion(title);image synthesis(abstract);分类 cs.CV
Comments Accepted at USENIX Security 2024. The first two authors contributed equally. We communicated with the author of DPSDA and have added an explanation for why the FID scores in the DPSDA table are lower than those reported in the original paper
专题命中 可控生成 :diffusion(title,abstract);image synthesis(title);分类 cs.CV
Comments ICLR 2024
专题命中 可控生成 :diffusion(title,abstract);image generation(title);分类 cs.CV
Comments Project page at https://dave.ml/selfguidance/
可控纹理平铺:基于变换RoPE增强扩散模型
机构 * City University of Hong Kong(香港城市大学) ; Manycore Tech Inc.(Manycore科技公司) ; Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 可控生成 :diffusion(title,abstract);image editing(abstract);inpainting(abstract);分类 cs.CV、cs.GR
AI总结 提出基于扩散变换器的可控纹理平铺框架,通过坐标变换旋转位置编码实现精确平铺控制,并利用分离注意力掩码保持参考纹理结构,同时与场景光照几何无缝融合。
Comments The code and dataset are publicly accessible at https://github.com/junrongh/ControlTile
利用语义属性绑定实现扩散模型中的免费午餐颜色控制
机构 * Computer Vision Center, Spain(西班牙计算机视觉中心) ; Universitat Autònoma de Barcelona(巴塞罗那自治大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; City University of Hong Kong(香港城市大学) ; Program of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学项目)
专题命中 可控生成 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR
AI总结 ColorWave通过语义属性绑定实现扩散模型中精确的颜色控制,无需微调,提升颜色指定的准确性和适用性。
Comments WACV 2026. Project page: https://hecoding.github.io/colorwave-page
专题命中 可控生成 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments Project: https://snap-research.github.io/visual-composer/