arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86306 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3472 篇

2312.03043 2023-12-07 eess.IV cs.AI cs.CV q-bio.TO 91%

Navigating the Synthetic Realm: Harnessing Diffusion-based Models for Laparoscopic Text-to-Image Generation

Simeon Allmendinger, Patrick Hemmer, Moritz Queisner, Igor Sauer, Leopold Müller, Johannes Jakubik, Michael Vössing, Niklas Kühl

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10816 2023-08-22 cs.CV 91%

BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion

Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, Mike Zheng Shou

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image synthesis(title);分类 cs.CV

Comments Accepted by ICCV 2023. Code is available at: https://github.com/showlab/BoxDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10893 2023-07-18 cs.LG cs.AI cs.CV cs.CY cs.HC 91%

Fair Diffusion: Instructing Text-to-Image Generation Models on Fairness

Felix Friedrich, Manuel Brack, Lukas Struppek, Dominik Hintersdorf, Patrick Schramowski, Sasha Luccioni, Kristian Kersting

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13308 2023-05-23 cs.CV 91%

If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection

Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep Akata

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13378 2026-04-10 cs.GR cs.CV 91%

SMPL-GPTexture: Dual-View 3D Human Texture Estimation using Text-to-Image Generation Models

SMPL-GPTexture:利用文本到图像生成模型进行双视角3D人体纹理估计

Mingxiao Tu, Shuchang Ye, Hoijoon Jung, Jinman Kim

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);inpainting(abstract)

AI总结 本文提出SMPL-GPTexture方法,通过文本提示生成双视角图像,结合人体网格恢复模型和反向光栅化技术,生成高精度纹理映射,解决3D人体纹理生成中的隐私和数据获取难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11337 2025-01-31 cs.CV cs.CL cs.MM 91%

DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter

Ziyi Dong, Pengxu Wei, Liang Lin

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image editing(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19390 2024-12-02 cs.CV cs.GR cs.LG 91%

DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models

Shwetha Ram, Tal Neiman, Qianli Feng, Andrew Stuart, Son Tran, Trishul Chilimbi

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);image synthesis(abstract)

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04175 2023-10-24 cs.CR cs.CV cs.MM 91%

Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data Poisoning

Shengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu, Yuejian Fang, Hang Su

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);image synthesis(abstract)

Comments Carmera-ready version. To appear in ACM MM 2023. Code will be released at: https://github.com/sf-zhai/BadT2I

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06311 2023-10-11 cs.CV cs.MM 91%

Improving Compositional Text-to-image Generation with Large Vision-Language Models

Song Wen, Guian Fang, Renrui Zhang, Peng Gao, Hao Dong, Dimitris Metaxas

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image editing(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00642 2026-01-07 cs.DC 90%

HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation

HADIS:混合自适应扩散模型服务用于高效的文本到图像生成

Qizheng Yang, Tung-I Chen, Siyu Zhao, Ramesh K. Sitaraman, Hui Guan

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(title)

AI总结 HADIS通过混合自适应架构优化扩散模型服务,提升响应质量和降低延迟违规率

Comments 15 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14945 2026-08-12 cs.CV 版本更新 90%

Introspective Attention Modulation for Safe Text-to-Image Generation

用于安全文本到图像生成的内省注意力调制

Basim Azam, Hossein Rahmani, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学) Lancaster University(兰卡斯特大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

AI总结 研究基于流的文本到图像模型易产生不安全内容的问题,提出通过推理时内省调节注意力动态实现安全的方法,该方法在保持或提升质量的同时显著提高安全分数,为安全图像生成提供新途径。

Comments Accepted at ECCV 2026. 20 pages, 7 figures. Project page: https://basim-azam.github.io/iam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03067 2026-03-06 cs.CV 90%

EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models

编辑:用于文本到图像扩散模型的有效且可解释的提示倒置

Mingzhe Li, Kejing Xia, Gehao Zhang, Zhenting Wang, Guanhong Tao, Siqi Pan, Juan Zhai, Shiqing Ma

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校) Georgia Institute of Technology(佐治亚理工学院) Rutgers University(罗格斯大学) University of Utah(犹他大学) Dolby Laboratories(杜比实验室)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);image synthesis(abstract)

AI总结 本文提出\sys技术,通过预训练模型初始化、潜在空间反向工程和嵌入到文本转换,提升文本到图像扩散模型的提示倒置效果,实现更高的图像相似性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10993 2025-11-17 cs.CV 90%

CLUE: Controllable Latent space of Unprompted Embeddings for Diversity Management in Text-to-Image Synthesis

Keunwoo Park, Jihye Chae, Joong Ho Ahn, Jihoon Kweon

机构 * Departments of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center(融合医学部门、首尔大学医学院、Asan医院) Department of Otorhinolaryngology-Head and Neck Surgery, University of Ulsan College of Medicine, Asan Medical Center(耳鼻喉科-头颈外科部门、首尔大学医学院、Asan医院)

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19291 2025-11-11 cs.CV cs.AI 90%

TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis

Kazi Mahathir Rahman, Showrin Rahman, Sharmin Sultana Srishty

机构 * BRAC University(布拉克大学)

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);diffusion(abstract)

Comments 19 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20667 2025-08-12 cs.CV cs.AI cs.LG 90%

Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA

Ojonugwa Oluwafemi Ejiga Peter, Md Mahmudur Rahman, Fahmi Khalifa

机构 * Department of Computer Science, SCMNS School, Morgan State University, Baltimore, Maryland 21251, USA Electrical \& Computer Engineering Dept., School of Engineering, Morgan State University, Baltimore, Maryland 21251, USA

专题命中 文生图 :diffusion(title,abstract);image synthesis(title,abstract);image generation(abstract);text-to-image(abstract)

Journal ref Conference and Labs of the Evaluation Forum (CLEF) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15466 2025-06-05 cs.CV 90%

Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator

Chaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh Yoon

机构 * Data Science and AI Laboratory, ECE, Seoul National University(数据科学与人工智能实验室、电子与计算机工程系、首尔国立大学) AIIS, ASRI, INMC, ISRC, and Interdisciplinary Program in AI, Seoul National University(人工智能研究所、人工智能研究室、智能网络与计算中心、信息科学与工程系以及人工智能交叉学科项目、首尔国立大学)

专题命中 文生图 :text-to-image(title,abstract);inpainting(title,abstract);image generation(abstract);image editing(abstract)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06512 2025-05-16 cs.CV 90%

HCMA: Hierarchical Cross-model Alignment for Grounded Text-to-Image Generation

Hang Wang, Zhi-Qi Cheng, Chenhao Lin, Chao Shen, Lei Zhang

机构 * The Hong Kong Polytechnic University(香港理工大学) University of Washington(华盛顿大学) Xi’an Jiaotong University(西安交通大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13162 2025-04-18 cs.CV 90%

Personalized Text-to-Image Generation with Auto-Regressive Models

Kaiyue Sun, Xian Liu, Yao Teng, Xihui Liu

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

Comments Project page: https://github.com/KaiyueSun98/T2I-Personalization-with-AR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13920 2025-01-24 cs.CV cs.CL cs.LG 90%

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models

Jiayi Lei, Renrui Zhang, Xiangfei Hu, Weifeng Lin, Zhen Li, Wenjian Sun, Ruoyi Du, Le Zhuo, Zhongyu Li, Xinyue Li, Shitian Zhao, Ziyu Guo, Yiting Lu, Peng Gao, Hongsheng Li

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image editing(abstract)

Comments 75 pages, 73 figures, Evaluation scripts: https://github.com/jylei16/Imagine-e

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08159 2025-01-24 cs.CV cs.LG 90%

DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Jiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang, Dinghuai Zhang, Navdeep Jaitly, Josh Susskind, Shuangfei Zhai

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

Comments Accepted by ICLR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17310 2024-11-27 cs.CV cs.LG 90%

Reward Incremental Learning in Text-to-Image Generation

Maorong Wang, Jiafeng Mao, Xueting Wang, Toshihiko Yamasaki

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07546 2024-08-14 cs.CV cs.AI cs.CL 90%

Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

Xingyu Fu, Muyu He, Yujie Lu, William Yang Wang, Dan Roth

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

Comments COLM 2024, Project Url: https://zeyofu.github.io/CommonsenseT2I/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18958 2024-07-19 cs.CV 90%

AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation

Yanan Sun, Yanchen Liu, Yinhao Tang, Wenjie Pei, Kai Chen

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image synthesis(abstract)

Comments Accepted by ECCV 2024, code and dataset available in https://github.com/open-mmlab/AnyControl

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10683 2024-07-16 cs.CV cs.AI 90%

Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval

Youngsun Lim, Hyunjung Shim

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);image editing(abstract)

Comments This paper has been accepted for oral presentation at the IJCAI 2024 Workshop on Trustworthy Interactive Decision-Making with Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00637 2024-05-31 cs.CV 90%

Wuerstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models

Pablo Pernias, Dominic Rampas, Mats L. Richter, Christopher J. Pal, Marc Aubreville

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);image synthesis(abstract)

Comments Corresponding to "Würstchen v2"

Journal ref The Twelfth International Conference on Learning Representations (ICLR), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11929 2024-03-19 cs.CV 90%

LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model

Runhui Huang, Kaixin Cai, Jianhua Han, Xiaodan Liang, Renjing Pei, Guansong Lu, Songcen Xu, Wei Zhang, Hang Xu

专题命中 文生图 :diffusion(title,abstract);image synthesis(title,abstract);image generation(abstract);image editing(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11577 2024-03-05 cs.CV 90%

LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model

Chenjie Cao, Yunuo Cai, Qiaole Dong, Yikai Wang, Yanwei Fu

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);inpainting(abstract);image synthesis(abstract)

Comments Accepted by CVPR2024. Codes and models are released at https://github.com/ewrfcas/LeftRefill, Project page: https://ewrfcas.github.io/LeftRefill

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13490 2024-02-22 cs.CV 90%

Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models

Chen Wu, Fernando De la Torre

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.08891 2024-01-10 cs.CV cs.AI cs.CY cs.LG 90%

Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis

Lukas Struppek, Dominik Hintersdorf, Felix Friedrich, Manuel Brack, Patrick Schramowski, Kristian Kersting

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);diffusion(abstract)

Comments Published in the Journal of Artificial Intelligence Research (JAIR)

Journal ref Journal of Artificial Intelligence Research (JAIR), Vol. 78 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17489 2023-11-09 cs.CV 90%

Text-to-image Editing by Image Information Removal

Zhongping Zhang, Jian Zheng, Jacob Zhiyuan Fang, Bryan A. Plummer

专题命中 文生图 :text-to-image(title,abstract);image editing(title,abstract);image generation(abstract);diffusion(abstract)

Comments Full paper is accepted by WACV2024; Best paper runner-up of AI4CC@CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏