arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86585 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2505.15877 2025-10-15 cs.CV cs.CL cs.LG 57%

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

Siting Li, Xiang Gao, Simon Shaolei Du

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments NeurIPS 2025; 27 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV 57%

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08951 2025-10-13 eess.IV cs.CV 57%

FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation

Yingtie Lei, Zimeng Li, Chi-Man Pun, Yupeng Liu, Xuhang Chen

机构 * Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院) School of Electronic and Communication Engineering, Shenzhen Polytechnic University(深圳职业技术学院电子与通信工程学院) Department of Cardiology, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University, Guangzhou, China(广东省人民医院心内科(广东省医学科学院)) Guangdong Cardiovascular Institute, Guangdong Provincial People’s Hospital, Guangdong Academy of Medical Sciences, Guangzhou, China(广东省心血管病研究院,广东省人民医院,广东省医学科学院,广州,中国)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted by BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10575 2025-10-07 cs.CV 57%

A Survey of Defenses Against AI-Generated Visual Media: Detection,Disruption, and Authentication

Jingyi Deng, Chenhao Lin, Zhengyu Zhao, Shuai Liu, Zhe Peng, Qian Wang, Chao Shen

机构 * School of Cyber Science and Engineering, Xi'an Jiaotong University(西安交通大学计算机科学与工程学院) Interdisciplinary Research Center of Frontier science and technology, Xi'an Jiaotong University(西安交通大学前沿科学与技术交叉研究中心) School of Software Engineering, Xi'an Jiaotong University(西安交通大学软件工程学院) Department of Industrial and Systems Engineering, Hong Kong Polytechnic University(香港理工大学工业与系统工程系) School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted by ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25122 2025-09-30 cs.CV 57%

Triangle Splatting+: Differentiable Rendering with Opaque Triangles

Jan Held, Renaud Vandeghen, Sanghyun Son, Daniel Rebain, Matheus Gadelha, Yi Zhou, Ming C. Lin, Marc Van Droogenbroeck, Andrea Tagliasacchi

机构 * University of Liège(利根大学) Simon Fraser University(西蒙弗雷泽大学) University of Maryland(马里兰大学) University of British Columbia(不列颠哥伦比亚大学) University of Toronto(多伦多大学) Adobe Research(Adobe研究)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 9 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23555 2025-09-30 cs.CV 57%

From Fields to Splats: A Cross-Domain Survey of Real-Time Neural Scene Representations

Javed Ahmad, Penggang Gao, Donatien Delehelle, Mennuti Canio, Nikhil Deshpande, Jesús Ortiz, Darwin G. Caldwell, Yonas Teodros Tefera

机构 * Advanced Robotics, Istituto Italiano di Tecnologia (IIT)(先进机器人技术,意大利技术研究院) Istituto Nazionale per l’Assicurazione contro gli Infortuni sul Lavoro (INAIL)(国家职业伤害保险研究所) School of Computer Science, University of Nottingham(计算机科学学院,诺丁汉大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22940 2025-09-30 cs.CL cs.CV 57%

LLMs Behind the Scenes: Enabling Narrative Scene Illustration

Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

机构 * Midjourney Northwestern University(西北大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18699 2025-09-30 cs.CV 57%

AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group Swapping

Zedong Zhang, Ying Tai, Jianjun Qian, Jian Yang, Jun Li

机构 * Nanjing University of Science and Technology(南京理工大学) Nanjing University(南京大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted to SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22394 2025-09-29 eess.IV cs.AI cs.CV 57%

Deep Learning-Based Cross-Anatomy CT Synthesis Using Adapted nnResU-Net with Anatomical Feature Prioritized Loss

Javier Sequeiro González, Arthur Longuefosse, Miguel Díaz Benito, Álvaro García Martín, Fabien Baldacci

机构 * Univ. Bordeaux, CNRS, LaBRI, UMR 5800, F-33400 Talence, France(波尔多大学,法国国家科学研究中心,LaBRI研究院,UMR 5800) Universidad Autónoma de Madrid, Madrid, 28049, Spain(马德里自治大学,西班牙) Pazmany Peter Catholic University, Budapest, 1083, Hungary(佩斯大学,匈牙利) RIKEN Center for Integrative Medical Sciences, Medical Data Deep Learning Team, Tokyo, Japan(日本再生医学科学中心,医学数据深度学习团队)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21997 2025-09-29 cs.CV 57%

Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors

Youxu Shi, Suorong Yang, Dong Liu

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21473 2025-09-29 cs.LG cs.AI cs.CL cs.CV stat.ML 57%

Are Hallucinations Bad Estimations?

Hude Liu, Jerry Yao-Chieh Hu, Jennifer Yuntong Zhang, Zhao Song, Han Liu

机构 * Center for Foundation Models and Generative AI, Northwestern University(基础模型与生成式人工智能中心,西北大学) Department of Computer Science, Northwestern University(计算机科学系,西北大学) Ensemble AI Engineering Science, University of Toronto(工程科学,多伦多大学) University of California, Berkeley(加州大学伯克利分校) Department of Statistics and Data Science, Northwestern University(统计与数据科学系,西北大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Code is available at https://github.com/MAGICS-LAB/hallucination

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21466 2025-09-29 cs.CV cs.AI cs.CL 57%

Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models

Khaloud S. AlKhalifah, Malak Mashaabi, Hend Al-Khalifa

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20171 2025-09-25 cs.CV 57%

Optical Ocean Recipes: Creating Realistic Datasets to Facilitate Underwater Vision Research

Patricia Schöntag, David Nakath, Judith Fischer, Rüdiger Röttgers, Kevin Köser

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 26 pages, 9 figures, submitted to IEEE Journal of Ocean Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 57%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10328 2025-09-23 eess.IV cs.CV 57%

Anatomical feature-prioritized loss for enhanced MR to CT translation

Arthur Longuefosse, Baudouin Denis de Senneville, Gael Dournes, Ilyes Benlala, Pascal Desbarats, Fabien Baldacci

机构 * Univ. Bordeaux, CNRS, Bordeaux INP, LaBRI, UMR 5800, 33400 Talence, France(波尔多大学) Univ. Bordeaux, CNRS, Bordeaux INP, IMB, UMR 5251, 33400 Talence, France(波尔多大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Journal ref 2025 Phys. Med. Biol. 70 145012

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12667 2025-09-23 cs.CV 57%

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

Zihan Su, Xuerui Qiu, Hongbin Xu, Tangyu Jiang, Junhao Zhuang, Chun Yuan, Ming Li, Shengfeng He, Fei Richard Yu

机构 * Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) South China University of Technology(华南理工大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室) Singapore Management University(新加坡管理学院)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Safa-Sora is accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15391 2025-09-22 cs.CV 57%

RaceGAN: A Framework for Preserving Individuality while Converting Racial Information for Image-to-Image Translation

Mst Tasnim Pervin, George Bebis, Fang Jiang, Alireza Tavakkoli

机构 * 1 2 4 Department of Computer Science \& Engineering, 3 Department of Psychology, University of Nevada, Reno, USA Email: 1 , 2 , 3 , 4

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Journal ref ICMLA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14746 2025-09-19 cs.CV cs.IR 57%

Chain-of-Thought Re-ranking for Image Retrieval Tasks

Shangrong Wu, Yanghong Zhou, Yang Chen, Feng Zhang, P. Y. Mok

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03486 2025-09-12 cs.CR cs.CV cs.SI 57%

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) TU Delft(代尔夫特理工大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments To Appear in the ACM Conference on Computer and Communications Security (CCS), October 13, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08818 2025-09-11 cs.CV 57%

GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts

Jenna Kang, Maria Silva, Patsorn Sangkloy, Kenneth Chen, Niall Williams, Qi Sun

机构 * New York University(纽约大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07963 2025-09-09 cs.AI cs.CL cs.CV 57%

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards

Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu, Ji-Rong Wen, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) School of Computer Science(计算机科学学院) Baidu Inc.(百度公司) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04298 2025-09-05 cs.CV 57%

Noisy Label Refinement with Semantically Reliable Synthetic Images

Yingxuan Li, Jiafeng Mao, Yusuke Matsui

机构 * The University of Tokyo(东京大学) CyberAgent, Inc.(CyberAgent公司)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted to ICIP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00752 2025-09-03 cs.CV 57%

Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification

Y Hop Nguyen, Doan Anh Phan Huu, Trung Thai Tran, Nhat Nam Mai, Van Toi Giap, Thao Thi Phuong Dao, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南胡志明市科学大学) Thong Nhat Hospital(通纳特医院)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00549 2025-09-03 cs.CV 57%

A Modality-agnostic Multi-task Foundation Model for Human Brain Imaging

Peirong Liu, Oula Puonti, Xiaoling Hu, Karthik Gopinath, Annabel Sorby-Adams, Daniel C. Alexander, W. Taylor Kimberly, Juan E. Iglesias

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19555 2025-08-28 cs.CV 57%

MonoRelief V2: Leveraging Real Data for High-Fidelity Monocular Relief Recovery

Yu-Wei Zhang, Tongju Han, Lipeng Gao, Mingqiang Wei, Hui Liu, Changbao Li, Caiming Zhang

机构 * School of Mechanical Engineering, Qilu University of Technology (Shandong Academy of Sciences)(机械工程学院、齐鲁工业大学(山东科学院)) School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院、南京航空航天大学) School of Computer Science and Technology, Shandong University of Finance and Economics(计算机科学与技术学院、山东财经大学) School of Computer Science and Technology, Shandong University(计算机科学与技术学院、山东大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15761 2025-08-27 cs.CV 57%

Waver: Wave Your Way to Lifelike Video Generation

Yifu Zhang, Hao Yang, Yuqi Zhang, Yifei Hu, Fengda Zhu, Chuang Lin, Xiaofeng Mei, Yi Jiang, Bingyue Peng, Zehuan Yuan

机构 * Bytedance Waver Team(字节跳动Waver团队)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16424 2025-08-25 eess.IV cs.CV 57%

Decoding MGMT Methylation: A Step Towards Precision Medicine in Glioblastoma

Hafeez Ur Rehman, Sumaiya Fazal, Moutaz Alazab, Ali Baydoun

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11025 2025-08-25 cs.LG cs.AI cs.CV 57%

When Algorithms Play Favorites: Lookism in the Generation and Perception of Faces

Miriam Doh, Aditya Gulati, Matei Mancas, Nuria Oliver

机构 * ISIA Lab - Université de Mons, IRIDIA Lab - Université Libre de Bruxelles(ISIA实验室-蒙斯大学,IRIDIA实验室-布鲁塞尔自由大学) ISIA Lab - Université de Mons(ISIA实验室-蒙斯大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted as an extended abstract at the Fourth European Workshop on Algorithmic Fairness (EWAF) (URL: https://2025.ewaf.org/home)

Journal ref Proceedings of Machine Learning Research 294 (2025) 474-480

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13503 2025-08-20 cs.CV eess.IV 57%

AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic Scenes

Tianyi Xu, Fan Zhang, Boxin Shi, Tianfan Xue, Yujin Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学多媒体信息处理国家重点实验室,计算机科学学院) National Engineering Research Center of Visual Technology, School of Computer Science, Peking University(北京大学视觉技术国家工程研究中心,计算机科学学院)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03341 2025-08-20 eess.IV cs.CV physics.med-ph 57%

UltraDfeGAN: Detail-Enhancing Generative Adversarial Networks for High-Fidelity Functional Ultrasound Synthesis

Zhuo Li, Xuhang Chen, Shuqiang Wang, Bin Yuan, Nou Sotheany, Ngeth Rithea

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) University of Chinese Academy of Sciences(中国科学院大学) Department of Information and Communication, Institute of Technology of Cambodia(柬埔寨技术学院信息与通信系) Department of Telecommunication and Network Engineering, Institute of Technology of Cambodia(柬埔寨技术学院电信与网络工程系)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏