arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 3474 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3474 篇

2308.05184 2023-08-11 cs.HC cs.AI 88%

PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions

John Joon Young Chung, Eytan Adar

专题命中 文生图 :text-to-image(title,abstract);image generation(title);diffusion(abstract)

Comments Accepted to UIST2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05564 2023-07-13 cs.CL 88%

Augmenters at SemEval-2023 Task 1: Enhancing CLIP in Handling Compositionality and Ambiguity for Zero-Shot Visual WSD through Prompt Augmentation and Text-To-Image Diffusion

Jie S. Li, Yow-Ting Shiue, Yong-Siang Shih, Jonas Geiping

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract)

Comments Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13530 2023-05-03 cs.HC 88%

Text-to-Image Generation: Perceptions and Realities

Jonas Oppenlaender, Aku Visuri, Ville Paananen, Rhema Linder, Johanna Silvennoinen

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)

Comments Accepted at Generative AI in HCI workshop, CHI '23

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08477 2023-02-17 cs.HC 88%

Large-scale Text-to-Image Generation Models for Visual Artists' Creative Works

Hyung-Kwon Ko, Gwanmo Park, Hyeon Jeon, Jaemin Jo, Juho Kim, Jinwook Seo

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)

Comments 15 pages, 3 figures

Journal ref 28th International Conference on Intelligent User Interfaces (IUI '23), March 27--31, 2023, Sydney, NSW, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01902 2023-01-06 cs.HC 88%

What is in a Text-to-Image Prompt: The Potential of Stable Diffusion in Visual Arts Education

Nassim Dehouche, Kullathida Dehouche

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02907 2022-07-08 cs.NE 88%

Exploring Generative Adversarial Networks for Text-to-Image Generation with Evolution Strategies

Victor Costa, Nuno Lourenço, João Correia, Penousal Machado

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.02423 2021-11-30 cs.LG 88%

Improving Text-to-Image Synthesis Using Contrastive Learning

Hui Ye, Xiulong Yang, Martin Takac, Rajshekhar Sunderraman, Shihao Ji

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract)

Comments Accepted to BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14226 2026-08-17 cs.CV 新提交 87%

RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models

RankT2I:一种用于发现文本到图像模型中可解释且多样化语义的子模框架

Ritika Allada, Pinar Yanardag

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)

AI总结 本文提出无训练、模型无关的RankT2I框架,通过子模方法自动发现T2I模型的可编辑语义,性能优于现有方法,可高效获取多领域文本到图像编辑所需的多样化语义。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29013 2026-06-30 cs.CV 87%

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers

Mural: 通过混合Transformer将LLM知识迁移到图像生成

Achin Jain, Jie An, Siddharth Chaudhary, Davide Modolo

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

AI总结 提出Mixture-of-Transformers架构,将冻结的LLM与扩散图像生成器结合,仅用文本-图像对训练,在多个基准上取得优异性能,并涌现出跨语言生成、颜色引导构图等新能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18167 2026-04-21 cs.CV 87%

Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models

嵌入算术:一种轻量级、无需调优的文本到图像模型后处理偏见缓解框架

Venkatesh Thirugnana Sambandham, Torsten Schön

机构 * AIMotion Bavaria, Technische Hochschule Ingolstadt(AIMotion巴伐利亚、英格尔施泰特技术大学)

专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract,abstract_cn);image generation(abstract);分类 cs.CV

AI总结 本文提出一种轻量级、无需调优的后处理框架,通过嵌入算术分析并修正嵌入空间中的偏见,保持语义和视觉上下文,提升多样性的同时维持高概念一致性。

Comments A demo notebook with basic implementations can be found at \url{https://github.com/cvims/EMBEDDING-ARITHMETIC}

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03127 2026-04-06 cs.CV cs.AI 87%

Unified Thinker: A General Reasoning Modular Core for Image Generation

统一思考者:一种通用图像生成的推理模块核心

Sashuai Zhou, Qiang Zhou, Jijin Hu, Hanqing Yang, Yue Cao, Junpeng Ma, Yinchao Ma, Jun Song, Tiezheng Ge, Cheng Yu, Bo Zheng, Zhou Zhao

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Nanjing University(南京大学) Fudan University(复旦大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image editing(abstract);image synthesis(abstract)

AI总结 本文提出统一思考者,一种通用图像生成的推理模块核心,旨在通过可执行推理解决生成模型中的推理-执行差距,提升图像生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28713 2026-03-31 cs.CV 87%

DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing

DreamLite:一种轻量级的设备端统一模型用于图像生成和编辑

Kailai Feng, Yuxiang Wei, Bo Chen, Yang Pan, Hu Ye, Songwei Liu, Chenqian Yan, Yuan Gao

机构 * Intelligent Creation Lab, ByteDance(字节跳动智能创作实验室)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

AI总结 本文提出DreamLite,一种轻量级设备端统一扩散模型,支持图像生成和文本引导的图像编辑,通过剪枝移动U-Net骨干网络和潜在空间内的上下文空间拼接实现统一条件化,达到高效生成和编辑效果。

Comments https://carlofkl.github.io/dreamlite/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15150 2026-03-17 cs.CV 87%

SNCE: Geometry-Aware Supervision for Scalable Discrete Image Generation

SNCE:面向可扩展离散图像生成的几何感知监督

Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin, Aditya Grover, Jason Kuen

机构 * Adobe UCLA

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image editing(abstract);image synthesis(abstract)

AI总结 本文提出SNCE,一种新的训练目标,通过构建软类别分布解决大代码书离散图像生成的优化问题,提升收敛速度和生成质量。

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00591 2026-01-21 cs.CV cs.AI cs.CR 87%

Wukong Framework for Not Safe For Work Detection in Text-to-Image systems

Wukong框架用于文本到图像系统中的不安全内容检测

Mingrui Liu, Sixiao Zhang, Cheng Long

机构 * Nanyang Technological University(南洋理工大学)

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)

AI总结 Wukong通过利用扩散过程早期的去噪步骤和U-Net的预训练交叉注意力参数,实现高效的不安全内容检测,优于传统文本和图像过滤方法。

Comments Accepted by KDD'26 (round 1)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19791 2025-12-16 cs.CV 87%

Color Bind: Exploring Color Perception in Text-to-Image Models

Color Bind: 探索文本到图像模型中的颜色感知

Shay Shomer-Chai, Wenxuan Peng, Bharath Hariharan, Hadar Averbuch-Elor

机构 * Tel Aviv University(特拉维夫大学) Cornell University(康奈尔大学)

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)

AI总结 本文提出了一种专门的图像编辑技术,用于解决多对象提示中颜色属性的语义对齐问题,并在多种指标上提升了性能。

Comments Project webpage: https://tau-vailab.github.io/color-edit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15857 2025-10-20 cs.CV 87%

BLIP3o-NEXT: Next Frontier of Native Image Generation

Jiuhai Chen, Le Xue, Zhiyang Xu, Xichen Pan, Shusheng Yang, Can Qin, An Yan, Honglu Zhou, Zeyuan Chen, Lifu Huang, Tianyi Zhou, Junnan Li, Silvio Savarese, Caiming Xiong, Ran Xu

机构 * Salesforce Research(Salesforce研究部) University of Maryland(马里兰大学) Virginia Tech(弗吉尼亚理工大学) New York University(纽约大学) UC Davis(加州大学戴维斯分校)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17789 2025-04-29 cs.CV 87%

Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

Xu Ma, Peize Sun, Haoyu Ma, Hao Tang, Chih-Yao Ma, Jialiang Wang, Kunpeng Li, Xiaoliang Dai, Yujun Shi, Xuan Ju, Yushi Hu, Artsiom Sanakoyeu, Felix Juefei-Xu, Ji Hou, Junjiao Tian, Tao Xu, Tingbo Hou, Yen-Cheng Liu, Zecheng He, Zijian He, Matt Feiszli, Peizhao Zhang, Peter Vajda, Sam Tsai, Yun Fu

机构 * Northeastern University(东北大学) Meta GenAI(Meta 生成人工智能) Meta FAIR National University of Singapore(国立新加坡大学) The Chinese University of Hong Kong(香港中文大学) University of Washington(华盛顿大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments Project Page: https://ma-xu.github.io/token-shuffle/ Add related works

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13842 2025-03-11 cs.CV 87%

Detecting Human Artifacts from Text-to-Image Models

Kaihong Wang, Lingzhi Zhang, Jianming Zhang

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);inpainting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02237 2025-02-25 cs.CV cs.AI 87%

Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models

Jungwon Park, Jungmin Ko, Dongnam Byun, Jangwon Suh, Wonjong Rhee

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10797 2025-02-20 cs.CV 87%

STAR: Scale-wise Text-conditioned AutoRegressive image generation

Xiaoxiao Ma, Mohan Zhou, Tao Liang, Yalong Bai, Tiejun Zhao, Biye Li, Huaian Chen, Yi Jin

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12642 2024-12-30 cs.CV cs.AI 87%

Zero-shot Text-guided Infinite Image Synthesis with LLM guidance

Soyeong Kwon, Taegyeong Lee, Taehwan Kim

专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);diffusion(abstract);image editing(abstract)

Comments This paper is being withdrawn due to issues of misconduct in the experiments presented in Table 2 and Figures 6, 7, and 8. We recognize this as an ethical concern and sincerely apologize to the research community for any inconvenience it may have caused

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18928 2024-12-30 cs.CV cs.LG 87%

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation

Lunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Mingming Gong, Gui-Song Xia

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11340 2024-11-22 cs.CV cs.AI 87%

OmniGen: Unified Image Generation

Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Chaofan Li, Shuting Wang, Tiejun Huang, Zheng Liu

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

Comments Update the paper for OmniGen-v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03206 2024-03-06 cs.CV 87%

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, Robin Rombach

专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08056 2023-12-14 cs.CV cs.AI 87%

Knowledge-Aware Artifact Image Synthesis with LLM-Enhanced Prompting and Multi-Source Supervision

Shengguang Wu, Zhenglun Chen, Qi Su

专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03692 2023-12-07 cs.CR cs.CV cs.LG 87%

Memory Triggers: Unveiling Memorization in Text-To-Image Generative Models through Word-Level Duplication

Ali Naseh, Jaechul Roh, Amir Houmansadr

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01985 2023-12-05 cs.CV 87%

UniGS: Unified Representation for Image Generation and Segmentation

Lu Qi, Lehan Yang, Weidong Guo, Yu Xu, Bo Du, Varun Jampani, Ming-Hsuan Yang

专题命中 文生图 :image generation(title,abstract);diffusion(abstract);inpainting(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04841 2023-11-02 cs.CV cs.LG 87%

Is This Loss Informative? Faster Text-to-Image Customization by Tracking Objective Dynamics

Anton Voronov, Mikhail Khoroshikh, Artem Babenko, Max Ryabinin

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);image synthesis(abstract)

Comments Accepted to Conference on Neural Information Processing Systems (NeurIPS) 2023. 20 pages, 15 figures. Code: https://github.com/yandex-research/DVAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08465 2023-04-18 cs.CV 87%

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, Yinqiang Zheng

专题命中 文生图 :image synthesis(title);image generation(abstract);text-to-image(abstract);diffusion(abstract)

Comments Project available at https://ljzycmd.github.io/projects/MasaCtrl

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14412 2023-03-28 cs.CV 87%

Freestyle Layout-to-Image Synthesis

Han Xue, Zhiwu Huang, Qianru Sun, Li Song, Wenjun Zhang

专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏