arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86585 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2504.11455 2025-04-16 cs.CV 57%

SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Junke Wang, Zhi Tian, Xun Wang, Xinyu Zhang, Weilin Huang, Zuxuan Wu, Yu-Gang Jiang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments technical report, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08586 2025-04-16 cs.CV 57%

PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm

Haoyi Zhu, Honghui Yang, Xiaoyang Wu, Di Huang, Sha Zhang, Xianglong He, Hengshuang Zhao, Chunhua Shen, Yu Qiao, Tong He, Wanli Ouyang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2301.00157

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19390 2025-04-15 eess.IV cs.AI cs.CV 57%

Multi-modal Contrastive Learning for Tumor-specific Missing Modality Synthesis

Minjoo Lim, Bogyeong Kang, Tae-Eui Kam

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05508 2025-04-09 cs.CV 57%

PartStickers: Generating Parts of Objects for Rapid Prototyping

Mo Zhou, Josh Myers-Dean, Danna Gurari

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted to CVPR CVEU workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05456 2025-04-09 cs.CV 57%

Generative Adversarial Networks with Limited Data: A Survey and Benchmarking

Omar De Mitri, Ruyu Wang, Marco F. Huber

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05712 2025-04-09 cs.CV cs.AI 57%

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

Jianwen Jiang, Gaojie Lin, Zhengkun Rong, Chao Liang, Yongming Zhu, Jiaqi Yang, Tianyun Zhong

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04510 2025-04-08 cs.CV 57%

Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification

Shijian Wang, Linxin Song, Ryotaro Shimizu, Masayuki Goto, Hanqian Wu

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04448 2025-04-08 cs.CV eess.IV 57%

Thermoxels: a voxel-based method to generate simulation-ready 3D thermal models

Etienne Chassaing, Florent Forest, Olga Fink, Malcolm Mielle

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03376 2025-04-07 cs.CV 57%

FLAIRBrainSeg: Fine-grained brain segmentation using FLAIR MRI only

Edern Le Bot, Rémi Giraud, Boris Mansencal, Thomas Tourdias, Josè V. Manjon, Pierrick Coupé

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23125 2025-04-01 cs.CV cs.AI 57%

Evaluating Compositional Scene Understanding in Multimodal Generative Models

Shuhao Fu, Andrew Jun Lee, Anna Wang, Ida Momennejad, Trevor Bihl, Hongjing Lu, Taylor W. Webb

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01167 2025-04-01 cs.CV 57%

Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data

Haoxin Li, Boyang Li

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06740 2025-04-01 cs.LG cs.AI cs.CV cs.IR 57%

Sustainable techniques to improve Data Quality for training image-based explanatory models for Recommender Systems

Jorge Paz-Ruza, David Esteban-Martínez, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12202 2025-04-01 cs.AI cs.CV 57%

Nepotistically Trained Generative-AI Models Collapse

Matyas Bohacek, Hany Farid

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Journal ref Published in ICLR DATA-FM Workshop, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20540 2025-03-27 cs.CV 57%

Beyond Intermediate States: Explaining Visual Redundancy through Language

Dingchen Yang, Bowen Cao, Anran Zhang, Weibo Gu, Winston Hu, Guang Chen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07767 2025-03-25 cs.CV 57%

Learning Visual Generative Priors without Text

Shuailei Ma, Kecheng Zheng, Ying Wei, Wei Wu, Fan Lu, Yifei Zhang, Chen-Wei Xie, Biao Gong, Jiapeng Zhu, Yujun Shen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Project Page: https://ant-research.github.io/lumos

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03261 2025-03-19 eess.IV cs.CV 57%

Is JPEG AI going to change image forensics?

Edoardo Daniele Cannas, Sara Mandelli, Nataša Popović, Ayman Alkhateeb, Alessandro Gnutti, Paolo Bestagini, Stefano Tubaro

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10779 2025-03-17 cs.CV 57%

The Power of One: A Single Example is All it Takes for Segmentation in VLMs

Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10166 2025-03-14 cs.IR cs.AI cs.MM 57%

ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning

Pengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia, Linli Xu, Enhong Chen

专题命中 文生图 :text-to-image(abstract);分类 cs.MM

Comments WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08977 2025-03-13 cs.CV 57%

Decoupled Doubly Contrastive Learning for Cross Domain Facial Action Unit Detection

Yong Li, Menglin Liu, Zhen Cui, Yi Ding, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing 2025. A novel and elegant feature decoupling method for cross-domain facial action unit detection

Journal ref IEEE Transactions on Image Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08014 2025-03-12 cs.CV cs.AI 57%

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents

Yun Xing, Nhat Chung, Jie Zhang, Yue Cao, Ivor Tsang, Yang Liu, Lei Ma, Qing Guo

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07491 2025-03-11 eess.IV cs.CV 57%

NeAS: 3D Reconstruction from X-ray Images using Neural Attenuation Surface

Chengrui Zhu, Ryoichi Ishikawa, Masataka Kagesawa, Tomohisa Yuzawa, Toru Watsuji, Takeshi Oishi

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06287 2025-03-11 cs.CV cs.AI 57%

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03651 2025-03-06 cs.CV 57%

DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

Rui Zhao, Weijia Mao, Mike Zheng Shou

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17710 2025-02-26 cs.AI cs.CL cs.CV cs.LG 57%

Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures

Akhila Yerukola, Saadia Gabriel, Nanyun Peng, Maarten Sap

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments 40 pages, 49 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16062 2025-02-25 cs.HC cs.GR 57%

Creative Blends of Visual Concepts

Zhida Sun, Zhenyao Zhang, Yue Zhang, Min Lu, Dani Lischinski, Daniel Cohen-Or, Hui Huang

专题命中 文生图 :text-to-image(abstract);分类 cs.GR

Comments Accepted to ACM CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12095 2025-02-18 cs.CV 57%

Descriminative-Generative Custom Tokens for Vision-Language Models

Pramuditha Perera, Matthew Trager, Luca Zancato, Alessandro Achille, Stefano Soatto

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02589 2025-02-05 cs.CV 57%

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

Xueqing Deng, Qihang Yu, Ali Athar, Chenglin Yang, Linjie Yang, Xiaojie Jin, Xiaohui Shen, Liang-Chieh Chen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments project website: https://xdeng7.github.io/coconut.github.io/coconut_pancap.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01972 2025-02-05 eess.IV cs.AI cs.CV cs.LG 57%

Layer Separation: Adjustable Joint Space Width Images Synthesis in Conventional Radiography

Haolin Wang, Yafei Ou, Prasoon Ambalathankandy, Gen Ota, Pengyu Dai, Masayuki Ikebe, Kenji Suzuki, Tamotsu Kamishima

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16971 2025-01-29 cs.CV cs.LG 57%

RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples

Hossein Mirzaei, Mohammad Jafari, Hamid Reza Dehbashi, Ali Ansari, Sepehr Ghobadi, Masoud Hadi, Arshia Soltani Moakhar, Mohammad Azizmalayeri, Mahdieh Soleymani Baghshah, Mohammad Hossein Rohban

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Accepted at the Forty-First International Conference on Machine Learning (ICML) 2024. The implementation of our work is available at: \url{https://github.com/rohban-lab/RODEO}

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16164 2025-01-28 cs.HC cs.AI cs.ET cs.MM 57%

MetaDecorator: Generating Immersive Virtual Tours through Multimodality

Shuang Xie, Yang Liu, Jeannie S. A. Lee, Haiwei Dong

专题命中 文生图 :image synthesis(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏