arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 3474 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3474 篇

2306.02236 2023-06-06 cs.CV cs.AI cs.LG 89%

Detector Guidance for Multi-Object Text-to-Image Generation

Luping Liu, Zijian Zhang, Yi Ren, Rongjie Huang, Xiang Yin, Zhou Zhao

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17870 2023-05-24 cs.CV 89%

GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Jian Ma, Mingjun Zhao, Chen Chen, Ruichen Wang, Di Niu, Haonan Lu, Xiaodong Lin

专题命中 文生图 :image generation(title,abstract);text-to-image(title);diffusion(abstract);image synthesis(abstract)

Comments 24 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14573 2023-05-01 cs.CV cs.AI 89%

SceneGenie: Scene Graph Guided Diffusion Models for Image Synthesis

Azade Farshad, Yousef Yeganeh, Yu Chi, Chengzhi Shen, Björn Ommer, Nassir Navab

专题命中 文生图 :diffusion(title,abstract);image synthesis(title);image generation(abstract);text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12073 2023-04-27 cs.CV 89%

Towards Equitable Representation in Text-to-Image Synthesis Models with the Cross-Cultural Understanding Benchmark (CCUB) Dataset

Zhixuan Liu, Youeun Shin, Beverley-Claire Okogwu, Youngsik Yun, Lia Coleman, Peter Schaldenbrand, Jihie Kim, Jean Oh

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);diffusion(abstract);分类 cs.CV

Comments Still on going work

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14704 2023-04-04 cs.CV 89%

Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models

Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, Shenghua Gao

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by CVPR 2023. Project page: https://bluestyle97.github.io/dream3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17591 2023-03-31 cs.CV cs.AI cs.LG 89%

Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion Models

Eric Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, Humphrey Shi

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14517 2023-03-28 cs.CV 89%

Indonesian Text-to-Image Synthesis with Sentence-BERT and FastGAN

Made Raharja Surya Mahadi, Nugraha Priya Utama

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13439 2023-03-24 cs.CV 89%

Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, Humphrey Shi

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments The project is available at: https://github.com/Picsart-AI-Research/Text2Video-Zero

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01324 2023-03-15 cs.CV cs.LG 89%

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, Ming-Yu Liu

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12959 2023-01-31 cs.CV cs.AI 89%

GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis

Ming Tao, Bing-Kun Bao, Hao Tang, Changsheng Xu

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);diffusion(abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09515 2023-01-24 cs.LG cs.CV 89%

StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

Axel Sauer, Tero Karras, Samuli Laine, Andreas Geiger, Timo Aila

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);diffusion(abstract);分类 cs.CV

Comments Project page: https://sites.google.com/view/stylegan-t/

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06998 2023-01-10 cs.CR cs.CV cs.LG 89%

DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation Models

Zeyang Sha, Zheng Li, Ning Yu, Yang Zhang

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16031 2022-11-04 cs.CV cs.CL 89%

UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance

Wei Li, Xue Xu, Xinyan Xiao, Jiachen Liu, Hu Yang, Guohao Li, Zhanpeng Wang, Zhifan Feng, Qiaoqiao She, Yajuan Lyu, Hua Wu

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments First Version, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14124 2022-10-26 cs.CV cs.AI 89%

Lafite2: Few-shot Text-to-Image Generation

Yufan Zhou, Chunyuan Li, Changyou Chen, Jianfeng Gao, Jinhui Xu

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09596 2022-08-23 cs.CV 89%

Vision-Language Matching for Text-to-Image Synthesis via Generative Adversarial Networks

Qingrong Cheng, Keyu Wen, Xiaodong Gu

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02035 2022-04-06 cs.CV 89%

DT2I: Dense Text-to-Image Generation from Region Descriptions

Stanislav Frolov, Prateek Bansal, Jörn Hees, Andreas Dengel

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04928 2021-12-10 cs.CV cs.CL cs.LG 89%

Self-Supervised Image-to-Text and Text-to-Image Synthesis

Anindya Sundar Das, Sriparna Saha

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments ICONIP 2021 : The 28th International Conference on Neural Information Processing

Journal ref ICONIP 2021. Lecture Notes in Computer Science, vol 13111, pp 415-426. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.09983 2021-10-07 cs.CV 89%

Adversarial Text-to-Image Synthesis: A Review

Stanislav Frolov, Tobias Hinz, Federico Raue, Jörn Hees, Andreas Dengel

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments Published at Neural Networks Journal, available at https://www.sciencedirect.com/science/article/pii/S0893608021002823

Journal ref Neural Networks, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01361 2021-09-14 cs.CV 89%

Cycle-Consistent Inverse GAN for Text-to-Image Synthesis

Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

专题命中 文生图 :text-to-image(title,abstract);image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted at ACM MM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04750 2026-08-06 cs.CV cs.CL cs.MM 新提交 89%

Simile Understanding in Text-to-Image Models: An Evaluation Framework

文本到图像模型中的明喻理解:一个评估框架

Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito, Katsuhiko Hayashi, Jingun Kwon, Manabu Okumura, Taro Watanabe

机构 * The University of Tokyo(东京大学) Nara Institute of Science and Technology(奈良科学技术研究所) Chungnam National University(忠南国立大学) Institute of Science Tokyo(东京科学大学)

专题命中 文生图 :diffusion(summary_cn,abstract);text-to-image(title,abstract);分类 cs.CV、cs.MM

AI总结 针对文本到图像模型常混淆明喻喻体与本体的问题,提出含受控数据集、YOLO指标及Diffusion Lens分析的评估框架,实验发现模型存在字面化失败模式并讨论了缓解策略。

Comments Accepted as a full paper at ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08123 2026-04-10 cs.DC cs.AI 89%

LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows

LegoDiffusion:微服务文本到图像扩散工作流

Lingyun Yang, Suyi Li, Tianyu Feng, Xiaoxiao Jiang, Zhipeng Di, Weiyi Lu, Kan Liu, Yinghao Yu, Tao Lan, Guodong Yang, Lin Qu, Liping Zhang, Wei Wang

机构 * Hong Kong University of Science and Technology(香港科技大学) Alibaba Group(阿里巴巴集团)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

AI总结 LegoDiffusion通过将扩散工作流分解为松耦合的模型执行节点,实现高效管理与调度,提升请求率和突发流量容忍度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07970 2026-03-04 cs.LG 89%

Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective

文本到图像扩散模型的持续反学习:一种正则化视角

Justin Lee, Zheda Mai, Jinsu Yoo, Chongyu Fan, Cheng Zhang, Wei-Lun Chao

机构 * The Ohio State University(俄亥俄州立大学) Michigan State University(密歇根州立大学) Texas A&M University(德克萨斯大学安德森分校) Boston University(波士顿大学)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

AI总结 本文提出通过正则化方法解决文本到图像扩散模型中持续反学习的问题,通过减轻参数漂移和增强语义意识,提升反学习性能并推动生成式AI的安全发展。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12501 2025-12-16 cs.AI 89%

SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation

SafeGen: 在文本到图像生成中嵌入伦理保障

Dang Phuong Nam, Nguyen Kieu, Pham Thanh Hieu

机构 * Posts and Telecommunications Institute of Technology(电信技术研究所)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract)

AI总结 SafeGen通过整合BGE-M3和Hyper-SD,实现文本到图像生成中的伦理保障,有效过滤有害提示并生成高质量图像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23974 2025-10-29 cs.LG cs.AI 89%

Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models

Byeonghu Na, Minsang Park, Gyuwon Sim, Donghyeok Shin, HeeSun Bae, Mina Kang, Se Jung Kwon, Wanmo Kang, Il-Chul Moon

机构 * KAIST(韩国科学技术院) NAVER Cloud(NAVER云)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image editing(abstract)

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16853 2025-09-30 cs.LG 89%

Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models

Semin Kim, Yeonwoo Cha, Jaehoon Yoo, Seunghoon Hong

机构 * KAIST(韩国科学技术院)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

Comments 29 pages, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01605 2025-08-05 cs.CR 89%

Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models

Haoran Dai, Jiawen Wang, Ruo Yang, Manali Sharma, Zhonghao Liao, Yuan Hong, Binghui Wang

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18214 2025-07-22 cs.LG 89%

Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey

Yi Zhang, Zhen Chen, Chih-Hong Cheng, Wenjie Ruan, Xiaowei Huang, Dezong Zhao, David Flynn, Siddartha Khastgir, Xingyu Zhao

机构 * WMG, The University of Warwick(沃里克大学工学院) Dept. of Computer Science, University of Liverpool(利物浦大学计算机科学系) Dept. of Computer Science & Engineering, Chalmers University of Technology(楚科奇斯技术大学计算机科学与工程系) James-Watt Engineering School, University of Glasgow(格拉斯哥大学詹姆斯-沃特工程学院)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24360 2025-07-14 cs.LG 89%

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko, Abdur Raheem Ali, Yixiong Hao, Arthur Conmy

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

Comments 10 pages, 10 figures, Mechanistic Interpretability for Vision at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15381 2025-05-13 cs.DC 89%

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

Sohaib Ahmad, Qizheng Yang, Haoliang Wang, Ramesh K. Sitaraman, Hui Guan

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract)

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11358 2024-08-22 cs.CY 89%

Gender Bias Evaluation in Text-to-image Generation: A Survey

Yankun Wu, Yuta Nakashima, Noa Garcia

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏