arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86585 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2501.02146 2025-01-28 cs.CV cs.AI q-bio.NC 57%

Plasma-CycleGAN: Plasma Biomarker-Guided MRI to PET Cross-modality Translation Using Conditional CycleGAN

Yanxi Chen, Yi Su, Celine Dumitrascu, Kewei Chen, David Weidman, Richard J Caselli, Nicholas Ashton, Eric M Reiman, Yalin Wang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted by ISBI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12488 2025-01-23 eess.IV cs.CV q-bio.TO 57%

Bidirectional Brain Image Translation using Transfer Learning from Generic Pre-trained Models

Fatima Haimour, Rizik Al-Sayyed, Waleed Mahafza, Omar S. Al-Kadi

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 19 pages, 9 figures, 6 tables

Journal ref Computer Vision and Image Understanding, vol. 248, pp. 104100, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12327 2025-01-22 cs.CV 57%

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Xianwei Zhuang, Yuxin Xie, Yufan Deng, Liming Liang, Jinghan Ru, Yuguo Yin, Yuexian Zou

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04850 2025-01-14 cs.CV 57%

Robot Synesthesia: A Sound and Emotion Guided AI Painter

Vihaan Misra, Peter Schaldenbrand, Jean Oh

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04861 2025-01-14 cs.CV 57%

LayerMix: Enhanced Data Augmentation through Fractal Integration for Robust Deep Learning

Hafiz Mughees Ahmad, Dario Morle, Afshin Rahimi

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13553 2025-01-13 cs.CR cs.CV cs.LG 57%

AI-generated Image Detection: Passive or Watermark?

Moyang Guo, Yuepeng Hu, Zhengyuan Jiang, Zeyu Li, Amir Sadovnik, Arka Daw, Neil Gong

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02527 2025-01-07 cs.CV 57%

Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks

Leo Franklin, Apiradee Boonmee, Kritsada Wongsuwan

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06741 2025-01-07 cs.CV 57%

Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective

Ouxiang Li, Jiayin Cai, Yanbin Hao, Xiaolong Jiang, Yao Hu, Fuli Feng

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments KDD2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01125 2025-01-03 cs.CV 57%

DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

Feng Han, Kai Chen, Chao Gong, Zhipeng Wei, Jingjing Chen, Yu-Gang Jiang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments AAAI 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17628 2024-12-24 cs.CV 57%

Editing Implicit and Explicit Representations of Radiance Fields: A Survey

Arthur Hubert, Gamal Elghazaly, Raphael Frank

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10604 2024-12-19 cs.CV 57%

EvalGIM: A Library for Evaluating Generative Image Models

Melissa Hall, Oscar Mañas, Reyhane Askari-Hemmat, Mark Ibrahim, Candace Ross, Pietro Astolfi, Tariq Berrada Ifriqi, Marton Havasi, Yohann Benchetrit, Karen Ullrich, Carolina Braga, Abhishek Charnalia, Maeve Ryan, Mike Rabbat, Michal Drozdzal, Jakob Verbeek, Adriana Romero-Soriano

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments For code, see https://github.com/facebookresearch/EvalGIM/tree/main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07802 2024-12-12 cs.CV 57%

Language Model as Visual Explainer

Xingyi Yang, Xinchao Wang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04831 2024-12-09 cs.CV 57%

Customized Generation Reimagined: Fidelity and Editability Harmonized

Jian Jin, Yang Shen, Zhenyong Fu, Jian Yang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments 18 pages, 12 figures, ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03829 2024-12-06 cs.CV 57%

CLIP-FSAC++: Few-Shot Anomaly Classification with Anomaly Descriptor Based on CLIP

Zuo Zuo, Jiahao Dong, Yao Wu, Yanyun Qu, Zongze Wu

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00067 2024-12-03 cs.CV cs.LG 57%

Targeted Therapy in Data Removal: Object Unlearning Based on Scene Graphs

Chenhan Zhang, Benjamin Zi Hao Zhao, Hassan Asghar, Dali Kaafar

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17617 2024-11-27 eess.IV cs.CV 57%

An Ensemble Approach for Brain Tumor Segmentation and Synthesis

Juampablo E. Heras Rivera, Agamdeep S. Chopra, Tianyi Ren, Hitender Oswal, Yutong Pan, Zineb Sordo, Sophie Walters, William Henry, Hooman Mohammadi, Riley Olson, Fargol Rezayaraghi, Tyson Lam, Akshay Jaikanth, Pavan Kancharla, Jacob Ruzevick, Daniela Ushizima, Mehmet Kurt

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16087 2024-11-26 cs.CV 57%

AI-Generated Image Quality Assessment Based on Task-Specific Prompt and Multi-Granularity Similarity

Jili Xia, Lihuo He, Fei Gao, Kaifan Zhang, Leida Li, Xinbo Gao

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15735 2024-11-26 cs.CV 57%

Test-time Alignment-Enhanced Adapter for Vision-Language Models

Baoshun Tong, Kaiyu Song, Hanjiang Lai

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15453 2024-11-26 cs.CV cs.AI 57%

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

Te Yang, Jian Jia, Xiangyu Zhu, Weisong Zhao, Bo Wang, Yanhua Cheng, Yan Li, Shengyuan Liu, Quan Chen, Peng Jiang, Kun Gai, Zhen Lei

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14205 2024-11-22 cs.CV cs.AI 57%

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body

Zeqing Wang, Qingyang Ma, Wentao Wan, Haojie Li, Keze Wang, Yonghong Tian

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments 16 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10033 2024-11-18 cs.CV 57%

GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization

Yanhao Sun, RunZe Tian, Xiao Han, XinYao Liu, Yan Zhang, Kai Xu

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Pacific Graphics 2024

Journal ref Computer Graphics Forum (2024), 43: e15215

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07664 2024-11-13 cs.CV 57%

Evaluating the Generation of Spatial Relations in Text and Image Generative Models

Shang Hong Sim, Clarence Lee, Alvin Tan, Cheston Tan

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02592 2024-11-06 cs.CV cs.AI 57%

Decoupled Data Augmentation for Improving Image Classification

Ruoxin Chen, Zhe Wang, Ke-Yue Zhang, Shuang Wu, Jiamu Sun, Shouli Wang, Taiping Yao, Shouhong Ding

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02545 2024-11-06 cs.CV cs.CL 57%

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

Maitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim, Tejas Gokhale, Chitta Baral, Yezhou Yang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted at: NeurIPS 2024 | Project Page: https://tripletclip.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20340 2024-10-30 eess.IV cs.AI cs.CV cs.LG 57%

Enhancing GANs with Contrastive Learning-Based Multistage Progressive Finetuning SNN and RL-Based External Optimization

Osama Mustafa

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04306 2024-10-29 cs.CV cs.AI cs.LG 57%

Effectiveness Assessment of Recent Large Vision-Language Models

Yao Jiang, Xinyu Yan, Ge-Peng Ji, Keren Fu, Meijun Sun, Huan Xiong, Deng-Ping Fan, Fahad Shahbaz Khan

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted by Visual Intelligence

Journal ref Visual Intelligence, 2024, Vol. 2, article no. 17

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21001 2024-10-28 cs.CV cs.AI cs.LG 57%

GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models

Ali Abdollahi, Mahdi Ghaznavi, Mohammad Reza Karimi Nejad, Arash Mari Oriyad, Reza Abbasi, Ali Salesi, Melika Behjati, Mohammad Hossein Rohban, Mahdieh Soleymani Baghshah

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Journal ref Volume 392 of ECAI 2024, Pages 729 - 736

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06131 2024-10-28 cs.CV 57%

Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era

Chenghao Li, Chaoning Zhang, Joseph Cho, Atish Waghwase, Lik-Hang Lee, Francois Rameau, Yang Yang, Sung-Ho Bae, Choong Seon Hong

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17959 2024-10-24 eess.IV cs.CV cs.LG 57%

Medical Imaging Complexity and its Effects on GAN Performance

William Cagas, Chan Ko, Blake Hsiao, Shryuk Grandhi, Rishi Bhattacharya, Kevin Zhu, Michael Lam

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted to ACCV, Workshop on Generative AI for Synthetic Medical Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16820 2024-10-23 cs.CV 57%

AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models

Yongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yan Xu

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments This article has been accepted for publication in a future issue of IEEE Transactions on Medical Imaging (TMI), but has not been fully edited. Content may change prior to final publication. Citation information: DOI: https://doi.org/10.1109/TMI.2024.3473745 . Code: https://github.com/wuyongjianCODE/AttriPrompter

详情

展开后加载摘要…

URL PDF HTML 收藏