Leveraging Visual Question Answering to Improve Text-to-Image Synthesis
专题命中 文生图 :text-to-image(title,abstract);image synthesis(title);分类 cs.CV
Comments Accepted to the LANTERN workshop at COLING 2020
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :text-to-image(title,abstract);image synthesis(title);分类 cs.CV
Comments Accepted to the LANTERN workshop at COLING 2020
专题命中 文生图 :image generation(title,abstract);text-to-image(title);分类 cs.CV
专题命中 文生图 :image generation(title,abstract);text-to-image(title);分类 cs.CV
Comments 5 papers, 5 figures, Published in 2018 25th IEEE International Conference on Image Processing (ICIP)
专题命中 文生图 :text-to-image(title);image synthesis(title);image editing(abstract);分类 cs.CV
Comments To appear in the proceedings of IEEE Winter Conference on Applications of Computer Vision, WACV-2019
机构 * Information Engineering University Zhengzhou China ; School of Computer Science, \ University Wuhan China ; Xi’an Jiaotong University Xi’an China ; Wuhan University Wuhan China ; Information Engineering University ; School of Computer Science, \ University ; Xi’an Jiaotong University ; Wuhan University
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);inpainting(abstract);分类 cs.CV、cs.MM
Comments This paper has been accepted by ACM MM 2025
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV、cs.MM
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV、cs.MM
Journal ref Published at NeurIPS 2023
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.MM
Comments ICCV 2023
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments Tech Report
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV
Comments TLDR: We show for the first time that normalizing flows can be scaled for high-resolution and text-conditioned image synthesis
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV
Comments Multi-turn interactive image generation
专题命中 文生图 :image synthesis(title,abstract);diffusion(abstract,comments);image generation(abstract);分类 cs.CV
Comments v3(same as v2) version, update structure (add foreground generation, stable diffusion), add more experiments
多语言文本到图像生成中跨语言一致性的局限性研究
机构 * Khalifa University(哈利法大学) ; Queen Mary University of London(伦敦玛丽女王大学) ; The Chinese University of Hong Kong(香港中文大学) ; The University of Western Australia(西澳大学) ; The University of Melbourne(墨尔本大学) ; University of Central Florida(中佛罗里达大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
AI总结 本文针对多语言文本到图像生成的跨语言一致性局限,构建含10种语言、3.3万条提示词的LingT2I基准,经分析揭示语言相关生成规律,为开发更鲁棒的多语言T2I模型奠定基础。
Comments Accepted to ACM MM 2026
代理流引导与并行展开搜索用于空间感知的文本到图像生成
机构 * Harbin Institute of Technology(哈尔滨工业大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
AI总结 本文提出AFS-Search框架,通过闭合环路机制和空间感知引导提升文本到图像生成的精度与效率,实现性能和速度的双重优化。
PEPPER:基于感知的扰动用于文本到图像扩散模型的鲁棒后门防御
机构 * Texas A&M University(德克萨斯A&M大学) ; National Taiwan University(台湾国立大学) ; University of Michigan(密歇根大学)
专题命中 文生图 :diffusion(title,abstract);text-to-image(title)
AI总结 PEPPER通过重写提示生成语义上不同但视觉上相似的描述,干扰输入提示中的触发器,提高文本到图像扩散模型的鲁棒性,尤其有效对抗基于文本编码器的攻击。
文本到图像生成中默认图像的探索
机构 * University of Oulu Oulu Finland ; Carleton University Ottawa Canada ; IST, University of Lisbon \& IHA-NOVA FCSH / IN2PAST Lisbon Portugal ; University of Oulu ; Carleton University ; IST, University of Lisbon \& IHA-NOVA FCSH / IN2PAST
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
AI总结 本文首次探讨了文本到图像生成中默认图像的特性,通过实验分析揭示了默认图像的一致性及其对用户满意度的影响。
Comments 21 pages, 10 figures
SoS:多语言文本到图像生成中的表面与语义分析
机构 * Trustworthy AI Lab, University of Hamburg(汉堡大学可信人工智能实验室) ; Language Technology Group, University of Hamburg(汉堡大学语言技术小组)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
AI总结 本文分析了多语言文本到图像生成中表面与语义的偏差问题,揭示了T2I模型在不同语言下生成刻板印象图像的倾向,并提出新的度量方法来量化这种现象。
机构 * TikTok ; University of Maryland, College Park(马里兰大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
机构 * School of Cyber Science and Technology, Shandong University(山东大学网络科学与技术学院) ; Netflix Eyeline Studios ; State Key Laboratory of Cryptography and Digital Economy Security, Shandong University(山东大学密码与数字经济安全国家重点实验室) ; Shandong Key Laboratory of Artificial Intelligence Security, Shandong University(山东省人工智能安全重点实验室)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
机构 * University of Sheffield, UK(谢菲尔德大学)
专题命中 文生图 :diffusion(title,abstract);text-to-image(title)
机构 * University of Zurich(苏黎世大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
Comments Accepted to ICASSP 2025
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
Comments To be published at the CEGIS (critical evaluation of generative models and their impact on society) workshop at ECCV 2024
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
Comments Findings of the EMNLP 2024
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
Comments camera-ready version
专题命中 文生图 :text-to-image(title,abstract);image generation(title)
Comments EMNLP 2023
专题命中 文生图 :image synthesis(title,abstract);text-to-image(title)
TAMF-VTON:通过高保真图像合成实现纹理感知无掩码虚拟试穿
机构 * State Key Lab of CAD & CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室) ; Style3D Research China(中国凌迪科技风格实验室)
专题命中 文生图 :image synthesis(title,abstract);diffusion(abstract);inpainting(abstract);分类 cs.CV
AI总结 TAMF-VTON旨在解决虚拟试穿方法局限,通过统一生成管道,含轻量级专家混合适应方案、频域监督机制和强大数据管理管道,实现无掩码、多服装组合、纹理保留及高效推理,优于现有方法,为数字时尚提供可行方案。
用于文本到图像上下文学习的思维树推理
机构 * Korea University(韩国大学)
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);image synthesis(abstract);分类 cs.CV
AI总结 研究文本到图像上下文学习中模型的推理问题,提出思维树推理框架,通过多阶段推理和选择层减轻提示模糊性与组合错误,实验表明该结构化多分支推理能提升图像生成效果,无需额外训练或微调。
Comments 6 pages, 3 figures, 4 tables. Accepted at IEEE SMC 2026. Code available at https://github.com/Pandastep/ToT-T2I-ICL