arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 3480 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2505.19297 2026-03-09 cs.CV 79%

Alchemist: Turning Public Text-to-Image Data into Generative Gold

炼金术:将公共文本到图像数据转化为生成黄金

Valerii Startsev, Alexander Ustyuzhanin, Alexey Kirillov, Dmitry Baranchuk, Sergey Kastryulin

机构 * Yandex Research(Yandex研究院) HSE(莫斯科大学) Yandex MSU(莫斯科大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出了一种利用预训练模型生成高质量通用SFT数据集的方法,通过Alchemist数据集提升了文本到图像模型的生成质量并公开了微调权重。

Comments Accepted to the Datasets and Benchmarks Track of the 39th Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01579 2026-03-03 cs.CV cs.AI 79%

SkeleGuide: Explicit Skeleton Reasoning for Context-Aware Human-in-Place Image Synthesis

SkeleGuide: 基于显式骨骼推理的上下文感知人体图像合成

Chuqiao Wu, Jin Song, Yiyun Fei

机构 * Alibaba Group(阿里巴巴集团)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

AI总结 SkeleGuide通过显式骨骼推理提升上下文感知的人体图像合成质量,提供高保真且结构合理的生成结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03516 2026-03-03 cs.CV 79%

Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

绘画比思考更容易:文本到图像模型能否铺垫,却无法主导?

Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang, Rui Chen, Jiarong Ou, Xin Tao, Pengfei Wan, Xiaojuan Qi, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) The University of Hong Kong(香港大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出T2I-CoReBench基准测试,用于评估文本到图像模型的组合与推理能力,揭示现有模型在高组合场景和推理任务中的局限性。

Comments Accepted to ICLR 2026. Project Page: https://t2i-corebench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00445 2026-03-02 cs.CV 79%

AutoDebias: Automated Framework for Debiasing Text-to-Image Models

AutoDebias:文本到图像模型的自动化去偏框架

Hongyi Cai, Mohammad Mahdinur Rahman, Mingkang Dong, Muxin Pu, Moqyad Alqaily, Jie Li, Xinfeng Li, Jialie Shen, Meikang Qiu, Qingsong Wen

机构 * Universiti Malaya(马来大学) Monash University(莫纳什大学) United Arab Emirates University(阿拉伯联合酋长国大学) University of Science and Technology Beijing(北京科技大学) Nanyang Technological University (NTU)(南洋理工大学) City St George’s, University of London(伦敦大学圣乔治学院) Augusta University(奥古斯塔大学) Squirrel Ai Learning

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 AutoDebias提出了一种自动化框架,用于检测和减轻文本到图像模型中的恶意偏见,通过视觉语言模型和CLIP引导的训练过程,有效应对隐秘注入的刻板印象和多重攻击。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07687 2026-03-02 eess.IV cs.CV 79%

FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis

FermatSyn: 基于改进双向Mamba的多模态医学图像合成方法

Feng Yuan

机构 * USTC(中国科学技术大学) SII(上海信息研究所)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

AI总结 FermatSyn通过改进的双向Mamba结合Fermat螺旋扫描策略,解决多模态医学图像合成中全局一致性与局部细节的平衡问题,提升合成图像质量与临床应用价值。

Comments MICCAI 2026(under view)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20672 2026-02-25 cs.CV 79%

BBQ-to-Image: Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models

BBQ-to-Image: 数字边界框与颜色控制在大规模文本到图像模型中

Eliran Kachlon, Alexander Visheratin, Nimrod Sarid, Tal Hacham, Eyal Gutflaish, Saar Huberman, Hezi Zisman, David Ruppin, Ron Mokady

机构 * BRIA AI(BRIA人工智能)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 BBQ-to-Image通过引入数字边界框和RGB三元组的直接条件生成,实现了对文本到图像模型中物体位置、大小和颜色的精确控制,提升了生成图像的准确性和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25774 2026-02-25 cs.CV cs.AI cs.LG 79%

PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models

PCPO:比例信用政策优化用于对齐图像生成模型

Jeongjae Lee, Jong Chul Ye

机构 * KAIST(韩国科学技术院)

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 PCPO通过稳定的目标重构和时间步重新加权,解决图像生成模型训练中的不稳定性问题,从而提升收敛速度和图像质量。

Comments 35 pages, 20 figures. ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19350 2026-02-24 cs.CV 79%

PoseCraft: Tokenized 3D Body Landmark and Camera Conditioning for Photorealistic Human Image Synthesis

PoseCraft: 基于令牌化的3D人体姿态和相机条件化的人像图像合成

Zhilin Guo, Jing Yang, Kyle Fogarty, Jingyi Wan, Boqiao Zhang, Tianhao Wu, Weihao Xia, Chenliang Zhou, Sakar Khattar, Fangcheng Zhong, Cristina Nader Vasconcelos, Cengiz Oztireli

机构 * University of Cambridge(剑桥大学) Google(谷歌)

专题命中 文生图 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 PoseCraft通过令牌化3D姿态和相机信息,提升人像合成的逼真度和细节保留能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17160 2026-02-23 cs.CV cs.LG 79%

A Pragmatic Note on Evaluating Generative Models with Fréchet Inception Distance for Retinal Image Synthesis

关于使用Fréchet inception距离评估视网膜图像合成生成模型的务实笔记

Yuli Wu, Fucheng Liu, Rüveyda Yilmaz, Henning Konermann, Peter Walter, Johannes Stegmaier

机构 * RWTH Aachen University(亚琛RWTH大学) Uniklinik RWTH Aachen(亚琛RWTH医院) Heinrich Heine University Düsseldorf(多特蒙德海因里希·海涅大学)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

AI总结 本文探讨了在视网膜图像合成中,FID等指标与任务特定评估目标不一致的问题,并指出其在生物医学成像中的局限性。

Comments MIDL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13349 2026-02-17 cs.CV cs.AI 79%

From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models

从提示到生产:利用文本到图像模型自动化安全品牌营销图像

Parmida Atighehchian, Henry Wang, Andrei Kapustin, Boris Lerner, Tiancheng Jiang, Taylor Jensen, Negin Sokhandan

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出了一种自动化生成品牌安全营销图像的系统,通过文本到图像模型提升图像保真度和人类偏好。

Comments 17 pages, 12 figures, Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10053 2026-02-12 cs.CV 79%

DiCo: Disentangled Concept Representation for Text-to-image Person Re-identification

DiCo: 用于文本到图像人物重识别的解耦概念表示

Giyeol Kim, Chanho Eom

机构 * organization= Department of Imaging Science, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea organization= Department of Metaverse Convergence, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 DiCo通过解耦概念表示方法,提升文本到图像人物重识别的跨模态对齐和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16748 2026-02-10 cs.CV 79%

HyPlaneHead: Rethinking Tri-plane-like Representations in Full-Head Image Synthesis

HyPlaneHead:重新思考全头图像合成中的三平面表示

Heyuan Li, Kenkun Liu, Lingteng Qiu, Qi Zuo, Keru Zheng, Zilong Dong, Xiaoguang Han

机构 * Tongyi Lab, Alibaba Inc.(阿里云实验室) SSE, CUHK (Shenzhen)(城市大学(深圳)信息科学与工程学院) FNii-Shenzhen Guangdong Provincial Key Laboratory of Future Networks of Intelligence(未来网络智能化广东省重点实验室)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

AI总结 HyPlaneHead通过引入混合平面表示,解决三平面表示中的特征纠缠和映射不均问题,提升全头图像合成性能。

Comments Accepted by NeurIPS 2025. Project page: https://lhyfst.github.io/hyplanehead/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02227 2026-02-03 cs.CV 79%

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

展示,而非告知:将潜在推理转化为图像生成

Harold Haodong Chen, Xinxiang Yin, Wen-Jie Shu, Hongfei Zhang, Zixin Zhang, Chenfei Liao, Litao Guo, Qifeng Chen, Ying-Cong Chen

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 LatentMorph通过隐式潜在推理提升文本到图像生成的效率和效果,减少推理时间与标记消耗,同时提高与人类直觉的一致性。

Comments Code: https://github.com/EnVision-Research/LatentMorph

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04634 2026-01-26 cs.CV 79%

Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models

你所要求的是你所得到的吗?探究文本到图像模型中的概念关联

Salma Abdel Magid, Weiwei Pan, Simon Warchol, Grace Guo, Junsik Kim, Mahia Rahman, Hanspeter Pfister

机构 * Department of Computer Science(计算机科学系) Harvard University(哈佛大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出 Concept2Concept 框架,用于审计文本到图像模型中提示与生成内容之间的概念关联,通过可解释的概念和度量标准进行可视化分析。

Journal ref Trans. Mach. Learn. Res, 2835-8856, 2025, https://openreview.net/forum?id=mk1YIkVvTQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08835 2026-01-21 cs.CV cs.AI cs.CL 79%

CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics

CulturalFrames: 评估文本到图像模型与评估指标中的文化期望一致性

Shravan Nayak, Mehar Bhatia, Xiaofeng Zhang, Verena Rieser, Lisa Anne Hendricks, Sjoerd van Steenkiste, Yash Goyal, Karolina Stańczak, Aishwarya Agrawal

机构 * Mila – Quebec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) Samsung - SAIT AI Lab(三星-SAIT人工智能实验室) ETH AI Center(苏黎世联邦理工学院人工智能中心)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 CulturalFrames研究了文本到图像模型在文化期望一致性方面的表现,发现模型在显性和隐性文化期望上均存在显著遗漏,揭示了现有评估指标与人类判断的相关性不足,提出了改进文化意识模型的方向。

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09647 2026-01-15 cs.CV cs.CR cs.LG 79%

Identifying Models Behind Text-to-Image Leaderboards

识别文本到图像排行榜背后的模型

Ali Naseh, Yuefeng Peng, Anshuman Suri, Harsh Chaudhari, Alina Oprea, Amir Houmansadr

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Northeastern University(东北大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 研究揭示了文本到图像排行榜中通过图像嵌入空间聚类实现模型匿名性的突破,发现模型特定特征及提示对可区分性的影响,揭示了排行榜中的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04706 2026-01-09 cs.CV 79%

Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models

锻造与淬火:提升统一多模态模型中图像生成的高保真度

Yanbing Zeng, Jia Wang, Hanghang Ma, Junqiang Wu, Jie Zhu, Xiaoming Wei, Jie Hu

机构 * Meituan(美团)

专题命中 文生图 :image generation(title,abstract);分类 cs.CV

AI总结 本文提出Forge-and-Quench框架,通过利用理解模型增强生成图像的保真度和细节,提升多模态模型的生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01839 2026-01-09 cs.CR cs.AI cs.CL cs.CV 79%

Jailbreaking Safeguarded Text-to-Image Models via Large Language Models

通过大型语言模型对受保护的文本到图像模型进行劫持

Zhengyuan Jiang, Yuepeng Hu, Yuchen Yang, Yinzhi Cao, Neil Zhenqiang Gong

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出了一种利用微调大型语言模型来劫持受安全防护的文本到图像模型的方法,有效绕过安全防护并优于现有攻击技术。

Comments Accepted by EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15002 2026-01-07 cs.CV 79%

How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models

需要多少张图片?文本到图像模型中模仿阈值的估计

Sahil Verma, Royi Rassin, Arnav Das, Gantavya Bhatt, Preethi Seshadri, Chirag Shah, Jeff Bilmes, Hannaneh Hajishirzi, Yanai Elazar

机构 * University of Washington, Seattle(华盛顿大学) Bar-Ilan University(巴伊兰大学) University of California, Irvine(加州大学伊文斯顿分校) Allen Institute of AI(人工智能研究院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 研究通过评估文本到图像模型的模仿阈值,探讨其在版权和隐私合规中的应用。

Comments Accepted at TMLR 2025, ATTRIB, RegML, and SafeGenAI workshops at NeurIPS 2024 and NLLP Workshop 2024. https://openreview.net/forum?id=x0qJo7SPhs

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13754 2025-12-30 cs.CV 79%

Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval

跨模态全模式细粒度对齐用于文本到图像人物检索

Hao Yin, Xin Man, Feiyu Chen, Jie Shao, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳先进研究所,电子科学与技术大学) University of Electronic Science and Technology of China(电子科学与技术大学) Sichuan Artificial Intelligence Research Institute(四川人工智能研究院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出FMFA框架,通过显式细粒度对齐和隐式关系推理实现文本到图像人物检索的高精度匹配。

Comments accepted by ACM Transactions on Multimedia Computing Communications and Applications in December 2025

Journal ref ACM Transactions on Multimedia Computing Communications and Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18766 2025-12-23 cs.CV 79%

MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation

MaskFocus: 为掩码图像生成聚焦策略优化

Guohui Zhang, Hu Yu, Xiaoxiao Ma, Yaning Pan, Hang Xu, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 MaskFocus通过聚焦关键步骤优化策略,提升掩码图像生成模型的性能。

Comments Code is available at https://github.com/zghhui/MaskFocus

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14102 2025-12-17 cs.CV cs.AI cs.IR 79%

Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries

面向遥感文本到图像检索的神经符号推理:为复杂查询的基座模型

Emanuele Mezzi, Gertjan Burghouts, Maarten Kruithof

机构 * Vrije Universiteit Amsterdam(瓦赫宁根大学阿姆斯特丹) TNO(荷兰国防研究院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 RUNE通过结合大型语言模型和神经符号AI,利用逻辑推理提升遥感文本到图像检索的性能与可解释性,针对复杂查询和图像不确定性提出新的评估指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13427 2025-12-16 cs.CV cs.LG 79%

MineTheGap: Automatic Mining of Biases in Text-to-Image Models

MineTheGap: 自动挖掘文本到图像模型中的偏见

Noa Cohen, Nurit Spingarn-Eliezer, Inbar Huberman-Spiegelglas, Tomer Michaeli

机构 * Technion – Israel Institute of Technology(技术学院–以色列理工学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 MineTheGap通过遗传算法自动挖掘文本到图像模型中导致偏见的提示,利用偏见评分评估并优化提示池以暴露偏见。

Comments Code and examples are available on the project's webpage at https://noa-cohen.github.io/MineTheGap/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16329 2025-12-09 cs.CR cs.AI cs.CV 79%

DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling

通过分布建模实现文本到图像生成系统可扩展的红队测试

Boheng Li, Junjie Wang, Yiming Li, Zhiyang Hu, Leyi Qi, Jianshuo Dong, Run Wang, Han Qiu, Zhan Qin, Tianwei Zhang

机构 * School of Cyber Science(网络安全学院) State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 DREAM通过分布建模实现文本到图像生成系统的可扩展红队测试,有效发现多样化的有害提示,提升安全性和多样性。

Comments To appear in the IEEE Symposium on Security & Privacy, May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08701 2025-11-26 cs.CV cs.AI cs.LG 79%

SafeFix: Targeted Model Repair via Controlled Image Generation

SafeFix: 通过受控图像生成实现目标模型修复

Ouyang Xu, Baoming Zhang, Ruiyu Mao, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 SafeFix通过生成语义忠实的图像来修复模型,提升模型对罕见案例的鲁棒性,减少系统性错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18684 2025-11-25 cs.CV 79%

Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation

现在你看见它,现在你又看不见 - 用于安全文本到图像和视频生成的即时概念消除

Shristi Das Biswas, Arani Roy, Kaushik Roy

机构 * Purdue University(普渡大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 ICE通过无需训练的一次性权重修改方法,实现文本到图像和视频生成中精确且持久的概念消除,无需额外开销,提升安全性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16965 2025-11-24 cs.CV cs.LG 79%

Real-Time Cooked Food Image Synthesis and Visual Cooking Progress Monitoring on Edge Devices

边缘设备上的实时烹饪食品图像合成与视觉烹饪进程监控

Jigyasa Gupta, Soumya Goyal, Anil Kumar, Ishan Jindal

机构 * Samsung R&D Institute India(三星印度研发院)

专题命中 文生图 :image synthesis(title);image generation(abstract);分类 cs.CV

AI总结 本文提出了一种边缘高效的生成模型,通过引入基于烤箱的烹饪进程数据集和领域特定的CIS度量标准,实现了在边缘设备上实时合成逼真烹饪食品图像并监控烹饪进程。

Comments 13 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22300 2025-11-24 cs.CR cs.AI cs.CV 79%

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

T2I-RiskyPrompt:用于评估、攻击和防御文本到图像模型安全性的基准

Chenyu Zhang, Tairen Zhang, Lanjun Wang, Ruidong Chen, Wenhui Li, Anan Liu

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 T2I-RiskyPrompt提出了一种用于评估文本到图像模型安全性的综合基准,通过层次化风险分类和原因驱动的检测方法,全面评估了多种模型和防御策略的安全性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17987 2025-11-24 cs.CR cs.AI cs.CV 79%

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

通过LLM推理进行文本到图像模型的劫持:Reason2Attack

Chenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li, An-An Liu

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 Reason2Attack通过将劫持攻击融入LLM后训练过程,提升生成对抗性提示的推理能力,实现更高效的文本到图像模型攻击

Comments Noted that This paper includes model-generated content that may contain offensive or distressing material

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14449 2025-11-19 cs.CV 79%

DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval

Zongwei Zhen, Biqing Zeng

机构 * South China Normal University School of Artificial Intelligence(南方科技大学人工智能学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏