PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract)
Comments Accepted to UIST2023
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract)
Comments Accepted to UIST2023
专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract)
Comments Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)
专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)
Comments Accepted at Generative AI in HCI workshop, CHI '23
专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)
Comments 15 pages, 3 figures
Journal ref 28th International Conference on Intelligent User Interfaces (IUI '23), March 27--31, 2023, Sydney, NSW, Australia
专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract)
专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)
专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract)
Comments Accepted to BMVC 2021
RankT2I:一种用于发现文本到图像模型中可解释且多样化语义的子模框架
机构 * Virginia Tech(弗吉尼亚理工大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)
AI总结 本文提出无训练、模型无关的RankT2I框架,通过子模方法自动发现T2I模型的可编辑语义,性能优于现有方法,可高效获取多领域文本到图像编辑所需的多样化语义。
Comments ECCV 2026
Mural: 通过混合Transformer将LLM知识迁移到图像生成
机构 * Amazon AGI(亚马逊人工智能研究院)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
AI总结 提出Mixture-of-Transformers架构,将冻结的LLM与扩散图像生成器结合,仅用文本-图像对训练,在多个基准上取得优异性能,并涌现出跨语言生成、颜色引导构图等新能力。
嵌入算术:一种轻量级、无需调优的文本到图像模型后处理偏见缓解框架
机构 * AIMotion Bavaria, Technische Hochschule Ingolstadt(AIMotion巴伐利亚、英格尔施泰特技术大学)
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract,abstract_cn);image generation(abstract);分类 cs.CV
AI总结 本文提出一种轻量级、无需调优的后处理框架,通过嵌入算术分析并修正嵌入空间中的偏见,保持语义和视觉上下文,提升多样性的同时维持高概念一致性。
Comments A demo notebook with basic implementations can be found at \url{https://github.com/cvims/EMBEDDING-ARITHMETIC}
统一思考者:一种通用图像生成的推理模块核心
机构 * Zhejiang University(浙江大学) ; Alibaba Group(阿里巴巴集团) ; Nanjing University(南京大学) ; Fudan University(复旦大学)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image editing(abstract);image synthesis(abstract)
AI总结 本文提出统一思考者,一种通用图像生成的推理模块核心,旨在通过可执行推理解决生成模型中的推理-执行差距,提升图像生成质量。
DreamLite:一种轻量级的设备端统一模型用于图像生成和编辑
机构 * Intelligent Creation Lab, ByteDance(字节跳动智能创作实验室)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)
AI总结 本文提出DreamLite,一种轻量级设备端统一扩散模型,支持图像生成和文本引导的图像编辑,通过剪枝移动U-Net骨干网络和潜在空间内的上下文空间拼接实现统一条件化,达到高效生成和编辑效果。
SNCE:面向可扩展离散图像生成的几何感知监督
机构 * Adobe ; UCLA
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image editing(abstract);image synthesis(abstract)
AI总结 本文提出SNCE,一种新的训练目标,通过构建软类别分布解决大代码书离散图像生成的优化问题,提升收敛速度和生成质量。
Comments 21 pages, 4 figures
Wukong框架用于文本到图像系统中的不安全内容检测
机构 * Nanyang Technological University(南洋理工大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)
AI总结 Wukong通过利用扩散过程早期的去噪步骤和U-Net的预训练交叉注意力参数,实现高效的不安全内容检测,优于传统文本和图像过滤方法。
Comments Accepted by KDD'26 (round 1)
Color Bind: 探索文本到图像模型中的颜色感知
机构 * Tel Aviv University(特拉维夫大学) ; Cornell University(康奈尔大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)
AI总结 本文提出了一种专门的图像编辑技术,用于解决多对象提示中颜色属性的语义对齐问题,并在多种指标上提升了性能。
Comments Project webpage: https://tau-vailab.github.io/color-edit/
机构 * Salesforce Research(Salesforce研究部) ; University of Maryland(马里兰大学) ; Virginia Tech(弗吉尼亚理工大学) ; New York University(纽约大学) ; UC Davis(加州大学戴维斯分校)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)
机构 * Northeastern University(东北大学) ; Meta GenAI(Meta 生成人工智能) ; Meta FAIR ; National University of Singapore(国立新加坡大学) ; The Chinese University of Hong Kong(香港中文大学) ; University of Washington(华盛顿大学)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
Comments Project Page: https://ma-xu.github.io/token-shuffle/ Add related works
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);inpainting(abstract)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)
Comments Accepted by ICLR 2025
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
Comments 16 pages
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)
Comments This paper is being withdrawn due to issues of misconduct in the experiments presented in Table 2 and Figures 6, 7, and 8. We recognize this as an ethical concern and sincerely apologize to the research community for any inconvenience it may have caused
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)
Comments Update the paper for OmniGen-v1
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)
专题命中 文生图 :image generation(title,abstract);diffusion(abstract);inpainting(abstract);image synthesis(abstract)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)
Comments Accepted to Conference on Neural Information Processing Systems (NeurIPS) 2023. 20 pages, 15 figures. Code: https://github.com/yandex-research/DVAR
专题命中 文生图 :image synthesis(title);image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments Project available at https://ljzycmd.github.io/projects/MasaCtrl
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments Accepted to CVPR 2023