arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 3480 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2506.01493 2025-06-03 cs.CV cs.LG 79%

Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity

Yuya Kobayashi, Yuhta Takida, Takashi Shibuya, Yuki Mitsufuji

机构 * SonyAI(索尼人工智能)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Accepted at IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00786 2025-06-03 cs.CV 79%

Aiding Medical Diagnosis through Image Synthesis and Classification

Kanishk Choudhary

机构 * Independent Researcher(独立研究者) Fremont, CA 94539(美国弗里蒙特)

专题命中 文生图 :image synthesis(title);diffusion(abstract);分类 cs.CV

Comments 8 pages, 6 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21074 2025-05-28 cs.LG cs.AI cs.CR cs.CV stat.ML 79%

Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

Yichuan Cao, Yibo Miao, Xiao-Shan Gao, Yinpeng Dong

机构 * KLMM, UCAS, Academy of Mathematics and Systems Science, Chinese Academy of Sciences(UCAS信息科学学院、数学与系统科学研究院、中国科学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13981 2025-05-27 cs.CV cs.AI 79%

On the Fairness, Diversity and Reliability of Text-to-Image Generative Models

Jordan Vice, Naveed Akhtar, Leonid Sigal, Richard Hartley, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Melbourne(墨尔本大学) University of British Columbia(不列颠哥伦比亚大学) Australian National University(澳大利亚国立大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments This research is supported by the NISDRG project #20100007, funded by the Australian Government

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.03675 2025-05-27 cs.CV cs.AI cs.CL cs.CY 79%

Auditing Gender Presentation Differences in Text-to-Image Models

Yanzhe Zhang, Lu Jiang, Greg Turk, Diyi Yang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Preprint, 23 pages, 14 figures. Project page at https://salt-nlp.github.io/GEP/

Journal ref EAAMO '24: Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17280 2025-05-26 cs.CV 79%

Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models

Pushkar Shukla, Aditya Chinchure, Emily Diana, Alexander Tolbert, Kartik Hosanagar, Vineeth N Balasubramanian, Leonid Sigal, Matthew Turk

机构 * Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) University of British Columbia(不列颠哥伦比亚大学) Carnegie Mellon University, Tepper School of Business(卡内基梅隆大学商学院) Emory University(埃默里大学) University of Pennsylvania, The Wharton School(宾夕法尼亚大学沃顿商学院) Indian Institute of Technology Hyderabad(海得拉巴印度理工学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14341 2025-05-21 cs.CV cs.AI 79%

Replace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image

Sifan Li, Ming Tao, Hao Zhao, Ling Shao, Hao Tang

机构 * Liaoning University(辽宁大学) Nanjing University of Posts and Telecommunications(南京邮电大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00759 2025-05-14 cs.CV cs.AI cs.CL 79%

Multi-Modal Language Models as Text-to-Image Model Evaluators

Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Michal Drozdzal, Adriana Romero-Soriano

机构 * FAIR at Meta - Montreal, New York(Meta 蒙特利尔 FAIR) University of Texas at Austin(德克萨斯大学奥斯汀分校) Mila McGill University(麦吉尔大学) Canada CIFAR AI chair(加拿大 CIFAR 人工智能主席)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04105 2025-05-12 eess.IV cs.CV 79%

MAISY: Motion-Aware Image SYnthesis for Medical Image Motion Correction

Andrew Zhang, Hao Wang, Shuchang Ye, Michael Fulham, Jinman Kim

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02236 2025-05-06 cs.CV cs.AI 79%

Improving Physical Object State Representation in Text-to-Image Generative Systems

Tianle Chen, Chaitanya Chakka, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Runway

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Submitted to Synthetic Data for Computer Vision - CVPR 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16870 2025-04-24 cs.CV 79%

High-Quality Cloud-Free Optical Image Synthesis Using Multi-Temporal SAR and Contaminated Optical Data

Chenxi Duan

机构 * Faculty of Geo-Information Science and Earth Observation (ITC), University of Twente(地理信息科学与地球观测学院(ITC)、特文特大学)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07718 2025-04-11 cs.CV 79%

Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval

Zehong Ma, Hao Chen, Wei Zeng, Limin Su, Shiliang Zhang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments TMM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07556 2025-04-11 cs.CV 79%

TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs

Zijian Zhang, Xuhui Zheng, Xuecheng Wu, Chong Peng, Xuezhi Cao

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13989 2025-04-01 cs.CV cs.AI cs.CL cs.LG 79%

Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models

Masane Fuchi, Tomohiro Takagi

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments 21 pages, 8 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17831 2025-03-25 eess.IV cs.AI cs.CV 79%

FundusGAN: A Hierarchical Feature-Aware Generative Framework for High-Fidelity Fundus Image Generation

Qingshan Hou, Meng Wang, Peng Cao, Zou Ke, Xiaoli Liu, Huazhu Fu, Osmar R. Zaiane

专题命中 文生图 :image generation(title);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16820 2025-03-18 cs.CV 79%

Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Olivia Wiles, Chuhan Zhang, Isabela Albuquerque, Ivana Kajić, Su Wang, Emanuele Bugliarello, Yasumasa Onoe, Pinelopi Papalampidi, Ira Ktena, Chris Knutsen, Cyrus Rashtchian, Anant Nawalgaria, Jordi Pont-Tuset, Aida Nematzadeh

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Accepted to ICLR 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02150 2025-03-18 cs.CV cs.LG eess.IV 79%

This Intestine Does Not Exist: Multiscale Residual Variational Autoencoder for Realistic Wireless Capsule Endoscopy Image Generation

Dimitrios E. Diamantis, Panagiota Gatoula, Anastasios Koulaouzidis, Dimitris K. Iakovidis

专题命中 文生图 :image generation(title);image synthesis(abstract);分类 cs.CV

Comments 10 pages

Journal ref IEEE Access, 12, 25668-25683 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11481 2025-03-17 cs.CV 79%

T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation

Seyed Mohammad Hadi Hosseini, Amir Mohammad Izadi, Ali Abdollahi, Armin Saghafian, Mahdieh Soleymani Baghshah

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop EVAL-FoMo

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12761 2025-03-17 cs.CV cs.AI cs.LG 79%

SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation

Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao, Mohit Bansal

专题命中 文生图 :text-to-image(title);diffusion(abstract);分类 cs.CV

Comments ICLR 2025; The first two authors contributed equally; Project page: https://safree-safe-t2i-t2v.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09962 2025-03-14 cs.CV 79%

Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification

Jiayu Jiang, Changxing Ding, Wentao Tan, Junhong Wang, Jin Tao, Xiangmin Xu

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments CVPR 2025. Project website: https://github.com/sssaury/HAM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09763 2025-03-14 cs.CV cs.CL cs.LG 79%

BiasConnect: Investigating Bias Interactions in Text-to-Image Models

Pushkar Shukla, Aditya Chinchure, Emily Diana, Alexander Tolbert, Kartik Hosanagar, Vineeth N. Balasubramanian, Leonid Sigal, Matthew A. Turk

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08012 2025-03-12 cs.CV cs.AI 79%

Exploring Bias in over 100 Text-to-Image Generative Models

Jordan Vice, Naveed Akhtar, Richard Hartley, Ajmal Mian

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Accepted to ICLR 2025 Workshop on Open Science for Foundation Models (SCI-FM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22592 2025-03-12 cs.CV 79%

GRADE: Quantifying Sample Diversity in Text-to-Image Models

Royi Rassin, Aviv Slobodkin, Shauli Ravfogel, Yanai Elazar, Yoav Goldberg

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments For project page and code see https://royira.github.io/GRADE

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15211 2025-03-11 cs.CV 79%

"Stones from Other Hills can Polish Jade": Zero-shot Anomaly Image Synthesis via Cross-domain Anomaly Injection

Siqi Wang, Yuanze Hu, Xinwang Liu, Siwei Wang, Guangpu Wang, Chuanfu Xu, Jie Liu, Ping Chen

专题命中 文生图 :image synthesis(title);diffusion(abstract);分类 cs.CV

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00178 2025-03-11 cs.CV cs.AI eess.IV 79%

Clinical Evaluation of Medical Image Synthesis: A Case Study in Wireless Capsule Endoscopy

Panagiota Gatoula, Dimitrios E. Diamantis, Anastasios Koulaouzidis, Cristina Carretero, Stefania Chetcuti-Zammit, Pablo Cortegoso Valdivia, Begoña González-Suárez, Alessandro Mussetto, John Plevris, Alexander Robertson, Bruno Rosa, Ervin Toth, Dimitris K. Iakovidis

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

Comments This work has been submitted for possible journal publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04902 2025-03-07 eess.IV cs.CV 79%

HAGAN: Hybrid Augmented Generative Adversarial Network for Medical Image Synthesis

Zhihan Ju, Wanting Zhou, Longteng Kong, Yu Chen, Yi Li, Zhenan Sun, Caifeng Shan

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

Journal ref Machine Intelligence Research 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00020 2025-03-04 cs.CL cs.AI cs.CV 79%

A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety

Rakeen Rouf, Trupti Bavalatti, Osama Ahmed, Dhaval Potdar, Faraz Jawed

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments Accepted for publication in IEEE Access, DOI: 10.1109/ACCESS.2025.3539933

Journal ref IEEE Access 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17360 2025-02-25 eess.IV cs.AI cs.CV 79%

RELICT: A Replica Detection Framework for Medical Image Generation

Orhun Utku Aydin, Alexander Koch, Adam Hilbert, Jana Rieger, Felix Lohrke, Fujimaro Ishida, Satoru Tanioka, Dietmar Frey

专题命中 文生图 :image generation(title);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09408 2025-02-21 cs.CV cs.LG 79%

Data Attribution for Text-to-Image Models by Unlearning Synthesized Images

Sheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Jun-Yan Zhu, Richard Zhang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments NeurIPS 2024 camera ready version. Project page: https://peterwang512.github.io/AttributeByUnlearning Code: https://github.com/PeterWang512/AttributeByUnlearning

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06798 2025-02-12 cs.LG cs.DC cs.GR 79%

Prompt-Aware Scheduling for Efficient Text-to-Image Inferencing System

Shubham Agarwal, Saud Iqbal, Subrata Mitra

专题命中 文生图 :text-to-image(title,abstract);分类 cs.GR

Comments Poster presented at NSDI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏