arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86504 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

1711.11585 2018-08-21 cs.CV cs.GR cs.LG 81%

High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, Bryan Catanzaro

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments v2: CVPR camera ready, adding more results for edge-to-photo examples

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.10992 2018-05-01 cs.CV cs.AI cs.GR cs.LG 81%

Semi-parametric Image Synthesis

Xiaojuan Qi, Qifeng Chen, Jiaya Jia, Vladlen Koltun

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments Published at the Conference on Computer Vision and Pattern Recognition (CVPR 2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02823 2018-04-17 cs.CV cs.GR 81%

TextureGAN: Controlling Deep Image Synthesis with Texture Patches

Wenqi Xian, Patsorn Sangkloy, Varun Agrawal, Amit Raj, Jingwan Lu, Chen Fang, Fisher Yu, James Hays

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments CVPR 2018 spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05349 2017-08-18 cs.CV cs.GR cs.LG 81%

PixelNN: Example-based Image Synthesis

Aayush Bansal, Yaser Sheikh, Deva Ramanan

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments Project Page: http://www.cs.cmu.edu/~aayushb/pixelNN/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01888 2026-04-03 cs.CV 80%

Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters

低努力对抗攻击针对文本到图像安全过滤器

Ahmed B Mustafa, Zihan Ye, Yang Lu, Michael P Pound, Shreyank N Gowda

机构 * University of Nottingham(诺丁汉大学) Xi’an Jiaotong-Liverpool University(西交利物浦大学) Xiamen University(厦门大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 研究揭示文本到图像模型易受低努力对抗攻击影响,通过提示策略绕过安全过滤,展示多种视觉对抗技术,揭示表面提示过滤与深层语义理解的差距,攻击成功率高达74.47%。

Comments Text-to-Image version of the Anyone can Jailbreak paper. Accepted in CVPR-W AIMS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13619 2023-11-27 cs.CV cs.CR 80%

Steal My Artworks for Fine-tuning? A Watermarking Framework for Detecting Art Theft Mimicry in Text-to-Image Models

Ge Luo, Junqiang Huang, Manman Zhang, Zhenxing Qian, Sheng Li, Xinpeng Zhang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments A Watermarking Framework for Detecting Art Theft Mimicry in Text-to-Image Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15678 2021-12-10 cs.CV 80%

A Shading-Guided Generative Implicit Model for Shape-Accurate 3D-Aware Image Synthesis

Xingang Pan, Xudong Xu, Chen Change Loy, Christian Theobalt, Bo Dai

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS2021. We proposed ShadeGAN, which could perform shape-accurate 3D-aware image synthesis by modeling shading in generative implicit models

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.12666 2021-08-10 cs.CV 80%

Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification

Zefeng Ding, Changxing Ding, Zhiyin Shao, Dacheng Tao

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments A new database for text-to-image ReID is provided. Code will be released

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06761 2025-05-16 cs.LG cs.MA 80%

Learning Graph Representation of Agent Diffusers

Youcef Djenouri, Nassim Belmecheri, Tomasz Michalak, Jan Dubiński, Ahmed Nabil Belbachir, Anis Yazidi

机构 * University of South-Eastern Norway(南欧挪威大学) Norwegian Research Centre(挪威研究中心) Simula Laboratory Research(Simula实验室研究) University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学) University of Oslo(奥斯陆大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments Accepted at AAMAS2025 International Conference on Autonomous Agents and Multiagent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08042 2022-01-21 cs.IR cs.LG 80%

GAN-based Matrix Factorization for Recommender Systems

Ervin Dervishaj, Paolo Cremonesi

专题命中 文生图 :image generation(abstract);text-to-image(abstract);inpainting(abstract);image synthesis(abstract)

Comments Accepted at the 37th ACM/SIGAPP Symposium on Applied Computing (SAC '22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14976 2026-08-18 cs.CV 新提交 79%

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

基于图像-描述提示的前沿文本到图像模型基准测试

Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad

机构 * Perle(珀尔)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 该研究针对组合要求高的图像-描述提示,评估了四个前沿文本到图像模型的性能,发现 Gemini 3 Pro Image 表现最优,领先系统的主要问题是对象计数错误和几何伪影。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11452 2026-08-13 cs.CV cs.AI 新提交 79%

TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation

TangPoetryBench:面向诗歌到图像生成的多维度基准与基于评分规则的评估器

Haoqi Hu, Tongji Luo, Li Zhang, Boning Zhou

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 该研究推出TangPoetryBench多维度基准与PAE评估器,解决T2I模型生成诗歌插图的评估难题,PAE性能接近Claude且可泛化,相关资源已公开。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03357 2026-08-05 cs.CV 新提交 79%

Can Text-to-Image Models Draw from the Right Frame of Reference?

文本到图像模型能否从正确的参考框架中生成图像?

Zheyuan Gu, Ruihang Li, Yong Huang, Yiqian Zhang, XIangzhao Hao, Jiaxin Niu, Jiahao Hu, Zhenyu Zhang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 该研究引入FoR-T2I基准,发现现有22个T2I模型在参考框架提示下的布局理解准确率远低于相机视图提示,还提出VLM门控重写方法提升了参考框架下的生成准确率。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00598 2026-08-04 cs.MM cs.CY 新提交 79%

EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios

EmergencyBias:文本到图像模型在应急场景下的偏差

Haibo Tang, Linqi Zhang, Hongxin Huan, Chenwei Lin, Xian Xu

专题命中 文生图 :text-to-image(title,abstract);分类 cs.MM

AI总结 本文定义了T2I模型在应急场景下的EmergencyBias,构建评估框架发现其存在人口统计学与行为偏差,提出ActionAlign方法可减少行为差异并保留图像质量。

Comments 15 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27779 2026-07-31 cs.CV 新提交 79%

CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography

CXR-Retrieve:胸部X线摄影中的组合式文本到图像检索

Tomer Erez, Moshe Kimhi, Chaim Baskin, Ehud Rivlin

机构 * Technion – Israel Institute of Technology(以色列理工学院) Ben-Gurion University of the Negev(内盖夫本-古里安大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 针对胸部X线检索的目标不匹配问题,提出CXR-Retrieve基准与标签感知对比微调方法,在双病理组合和否定查询的Precision@5上较CXR-CLIP分别提升8.5和22.0个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24215 2026-07-28 cs.CV 新提交 79%

TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

TreeAdapter:用于细粒度物种图像生成的分层分类法引导适配器组合

Yuze Sun, Zhongjie Duan, Yingda Chen

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 针对通用文本到图像模型在生成稀有生物物种图像时性能下降的问题,提出TreeAdapter框架,利用分层分类数据,通过在分类树节点附加轻量级适配器及两阶段训练范式,实现准确的细粒度图像生成,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20155 2026-07-20 cs.CV cs.CL 版本更新 79%

NAMESAKES: Probing Identity Memorization in Text-to-Image Models

NAMESAKES: 探究文本到图像模型中的身份记忆

Morris Alper, Vasudha Varadarajan, Moran Yanuka, Angelina Wang, Hadar Averbuch-Elor

机构 * Carnegie Mellon University(卡内基梅隆大学) Tel Aviv University(特拉维夫大学) Cornell University(康奈尔大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 提出一种黑盒行为探针,无需参考照片或训练数据,即可区分文本到图像模型生成的图像是记忆还是虚构,并在NAMESAKES数据集上验证其有效性。

Comments Project page: https://namesakes-web.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08201 2026-07-10 cs.CV cs.AI 新提交 79%

TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation

TMI:文本到图像与图像到图像结合用于互补数据合成以促进长尾实例分割

Hyeonseop Song, Seokhun Choi, Hoseok Do

机构 * LG Electronics(LG电子)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 研究针对大词汇量实例分割受长尾分布和类间模糊性限制的问题,提出结合文本到图像生成与上下文感知图像到图像编辑的混合管道,引入VRAIN编辑器,在LVIS基准测试中超越基线,有效提升分割性能。

Comments Accepted to ECCV 2026. The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24548 2026-07-02 cs.CV 新提交 79%

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

文本到图像模型是归纳主义的火鸡吗?一个用于因果推理的反事实基准

Jiayi Lei, Yuandong Pu, Xingyu Han, Rongpeng Zhu, Jing Xu, Jinyao Wang, Zijian Zhou, Bin Fu, Yuewen Cao, Yihao Liu, Hongsheng Li

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 提出反事实基准CF-World,通过三个递进层级测试T2I模型在违反现实先验规则下的图像生成能力,发现所有模型在反事实设置下性能急剧下降,原因是模型将世界知识与视觉外观编码为紧密耦合的模式。

Comments 10 pages, 7 figures. Project page: https://github.com/jylei16/CF-World.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30458 2026-06-30 cs.CV 79%

Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

跨分辨率语义迁移用于低分辨率监控下的鲁棒文本-图像检索

Wenjie Qian, Bin Yang, Xiao Wang, Wenke Huang, Ling Mei, Xin Xu, Mang Ye

机构 * School of Computer Science and Technology, Wuhan University of Science and Technology(武汉科技大学计算机科学与技术学院) School of Computer Science, National Engineering Research Center for Multimedia Software, Wuhan University(武汉大学计算机学院,国家多媒体软件工程技术研究中心) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System, Wuhan University of Science and Technology(湖北省智能信息处理与实时工业系统重点实验室,武汉科技大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 针对低分辨率监控场景中文本-图像检索的可靠性崩溃和排序漂移问题,提出CLIP框架CRST,通过分辨率条件推理、文本引导精炼和跨分辨率邻域迁移,在三个数据集上平均提升超低分辨率Rank-1和mAP分别5.7%和5.3%。

Comments 10 pages,8 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27122 2026-06-30 cs.CV 79%

InterPartAbility: Phrase-Region Grounding for Interpretable Text-to-Image Person Re-Identification

InterPartAbility: 基于文本引导的部分匹配用于可解释的人员重识别

Shakeeb Murtaza, Aryan Shukla, Rajarshi Bhattacharya, Maguelonne Heritier, Eric Granger

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(LIVIA系统工程系,蒙特利尔ÉTS学院,加拿大) Genetec Inc.(Genetec公司)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出InterPartAbility,通过显式部分匹配和短语-区域绑定提升TI-ReID的可解释性,引入PPIM模块实现概念级指导,生成 grounded 解释图谱,实验表明在CUHK-PEDES等基准上达到SOTA可解释性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25297 2026-06-25 cs.CV 新提交 79%

Minimalist Preprocessing Approach for Image Synthesis Detection

图像合成检测的极简预处理方法

Hoai-Danh Vo, Trung-Nghia Le

机构 * University of Science, VNU-HCM(胡志明市国立大学理科大学) Vietnam National University, Ho Chi Minh City(胡志明市国立大学)

专题命中 文生图 :image synthesis(title);image generation(abstract);分类 cs.CV

AI总结 提出一种基于梯度计算相邻像素波动的轻量级预处理方法,作为高通滤波器突出关键特征,在低端设备上实现与先进技术相当的检测精度。

Comments SOICT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06021 2026-06-24 cs.CV cs.AI cs.LG 79%

Improved Sub-Visible Particle Classification in Flow Imaging Microscopy via Generative AI-Based Image Synthesis

通过生成式AI图像合成改进流成像显微镜下的亚可见粒子分类

Utku Ozbulak, Michaela Cohrs, Hristo L. Svilenov, Joris Vankerschaver, Wesley De Neve

机构 * Center for Biosystems and Biotech Data Science(生物系统与生物技术数据科学中心) Ghent University Global Campus(根特大学全球校区) Department of Electronics and Information Systems(电子与信息系统系) Faculty of Pharmaceutical Sciences(药学系) Biopharmaceutical Technology, TUM School of Life Sciences(生物制药技术,技术大学生命科学学院) Department of Mathematics, Computer Science and Statistics(数学、计算机科学与统计学系)

专题命中 文生图 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 本文提出基于生成式AI的图像合成方法,解决流成像显微镜下亚可见粒子分类中的数据不平衡问题,通过生成高保真图像提升多类分类性能,公开模型和工具促进研究复现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23669 2026-06-23 cs.CV 新提交 79%

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

GeoFidelity-Bench:评估文本到图像街景生成中的片段级地理保真度

Kaizhen Tan, Hanzhe Hong, Siru Tao

机构 * Heinz College of Information Systems and Public Policy(信息系统与公共政策学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 提出GeoFidelity-Bench基准,通过参考面板排名测试文本到图像模型能否生成特定道路片段而非通用城市街景,发现添加街道和社区名称可提升检索准确率,但目标与最近邻片段间相似度差距近乎为零。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15967 2026-06-23 cs.CR cs.CV 版本更新 79%

When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models

TwoHamsters:文本到图像模型中多概念组合不安全性的基准测试

Chaoshuo Zhang, Yibo Liang, Mengke Tian, Chenhao Lin, Zhengyu Zhao, Le Yang, Chong Zhang, Yang Zhang, Qian Wang, Chao Shen

机构 * School of Cyber Science and Engineering, Xi'an Jiaotong University(西安交通大学计算机科学与工程学院) CISPA Helmholtz Center for Information Security(信息安全研究中心) School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出TwoHamsters基准,通过17500个提示测试文本到图像模型在多概念组合不安全性的表现,揭示现有模型和防御机制在处理危险组合生成时的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17853 2026-06-23 cs.CV cs.AI 版本更新 79%

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

Detail++: 文本到图像扩散模型的免训练细节增强器

Lifeng Chen, Jiner Wang, Zihao Pan, Beier Zhu, Xiaofeng Yang, Chi Zhang

机构 * AGI Lab, Westlake University(AGI实验室,西lake大学) Nanyang Technological University(南洋理工大学)

专题命中 文生图 :diffusion(title);text-to-image(abstract);分类 cs.CV

AI总结 提出免训练框架Detail++,通过渐进式细节注入策略分解复杂提示词,利用自注意力布局控制与交叉注意力质心对齐损失,提升多主体复杂提示下的生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18555 2026-06-18 cs.CV 新提交 79%

Rethinking Text-to-Image as Semantic-Aware Data Augmentation for Indoor Scene Recognition

重新思考文本到图像作为室内场景识别的语义感知数据增强

Trong-Vu Hoang, Quang-Binh Nguyen, Dinh-Khoi Vo, Hoai-Danh Vo, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM, Vietnam(越南国立大学胡志明市理科大学) Vietnam National University, Ho Chi Minh City, Vietnam(越南国立大学胡志明市分校)

专题命中 文生图 :text-to-image(title);diffusion(abstract);分类 cs.CV

AI总结 针对室内图像数据不足,提出利用稳定扩散生成合成图像进行数据增强,并通过扩散重建误差防止滥用,在MIT室内场景数据集上验证了有效性。

Comments MAPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14657 2026-06-15 cs.CV 新提交 79%

HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities

HPSv3++:跨扩散模型能力全谱系扩展奖励模型

Yijun Liu, Jie Huang, Zeyue Xue, Yuming Li, Ruizhe He, Haoran Li, Shijia Ge, Siming Fu

机构 * Tsinghua University(清华大学) JD Explore Academy(京东探索研究院) Peking University(北京大学) Zhejiang University(浙江大学)

专题命中 文生图 :diffusion(title);text-to-image(abstract);分类 cs.CV

AI总结 提出HPSv3++奖励模型框架,通过双维度偏好数据集HPDv3++和两阶段训练(正交梯度投影+无监督引导),提升对各类T2I模型及RL迭代的偏好预测能力,在多个基准上达到最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14351 2026-06-15 cs.CV 新提交 79%

ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image Models

ForceForget: 通过强化概念移除增强文本到图像模型的安全性

Dong Han, Yong Li

机构 * Dong Han(董汉) Yong Li(李勇)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 针对文本到图像模型生成不安全内容的问题,提出基于强化学习优化概念擦除奖励的方法,通过安全适配器调节文本嵌入,在消除不安全内容的同时保持模型对安全语义的生成能力。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05478 2026-06-05 cs.CV cs.LG 79%

Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?

我们能否在生成之前预测文生图内容的人类偏好,以及这样做是否有用?

Joong Ho Kim, Keith G. Mills

机构 * LSU ATHENA Lab(LSU ATHENA实验室)

专题命中 文生图 :text-to-image(title);diffusion(abstract);分类 cs.CV

AI总结 研究在扩散模型生成图像前预测人类偏好评分(HPM)的可行性,并利用该预测提升生成质量,同时评估不同HPM的适用性。

Comments Code is available at https://github.com/LSU-ATHENA/HPM-Predict

详情

展开后加载摘要…

URL PDF HTML 收藏