arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86539 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2108.01806 2022-07-26 cs.CV cs.GR 62%

Neural Scene Decoration from a Single Photograph

Hong-Wing Pang, Yingshu Chen, Phuoc-Hieu Le, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments ECCV 2022 paper. 14 pages of main content, 4 pages of references, and 11 pages of appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02308 2022-07-25 cs.CV cs.GR 62%

MoFaNeRF: Morphable Facial Neural Radiance Field

Yiyu Zhuang, Hao Zhu, Xusen Sun, Xun Cao

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments accepted to ECCV2022; code available at http://github.com/zhuhao-nju/mofanerf

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02546 2022-07-05 cs.CV cs.GR 62%

GANSpace: Discovering Interpretable GAN Controls

Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain Paris

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted to NeurIPS 2020

Journal ref Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 9841-9850

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.04455 2022-04-12 cs.GR cs.CV eess.IV 62%

Noise-based Enhancement for Foveated Rendering

Taimoor Tariq, Cara Tursun, Piotr Didyk

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments 14 pages including refences

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05349 2022-03-11 cs.MM cs.CV 62%

Two-stream Hierarchical Similarity Reasoning for Image-text Matching

Ran Chen, Hanli Wang, Lei Wang, Sam Kwong

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.14020 2022-03-01 cs.CV cs.GR cs.LG 62%

State-of-the-Art in the Architecture, Methods and Applications of StyleGAN

Amit H. Bermano, Rinon Gal, Yuval Alaluf, Ron Mokady, Yotam Nitzan, Omer Tov, Or Patashnik, Daniel Cohen-Or

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13162 2022-03-01 cs.CV cs.AI cs.GR cs.LG 62%

Pix2NeRF: Unsupervised Conditional $π$-GAN for Single Image to Neural Radiance Fields Translation

Shengqu Cai, Anton Obukhov, Dengxin Dai, Luc Van Gool

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.09855 2021-12-02 cs.CV cs.GR 62%

Infinite Nature: Perpetual View Generation of Natural Scenes from a Single Image

Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, Angjoo Kanazawa

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments ICCV 2021 (oral); Project page: https://infinite-nature.github.io/; Video: https://www.youtube.com/watch?v=oXUf6anNAtc

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06688 2021-10-14 cs.CV cs.GR 62%

DeepVecFont: Synthesizing High-quality Vector Fonts via Dual-modality Learning

Yizhi Wang, Zhouhui Lian

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments SIGGRAPH Asia 2021 Technical Paper. Code: https://github.com/yizhiwang96/deepvecfont ; Homepage: https://yizhiwang96.github.io/deepvecfont_homepage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04654 2021-09-13 cs.GR cs.CV 62%

Per Garment Capture and Synthesis for Real-time Virtual Try-on

Toby Chong, I-Chao Shen, Nobuyuki Umetani, Takeo Igarashi

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted to UIST2021. Project page: https://sites.google.com/view/deepmannequin/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.16011 2021-03-30 cs.CV cs.GR 62%

Intrinsic Autoencoders for Joint Neural Rendering and Intrinsic Image Decomposition

Hassan Abu Alhaija, Siva Karthik Mustikovela, Justus Thies, Varun Jampani, Matthias Nießner, Andreas Geiger, Carsten Rother

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12884 2020-12-24 cs.CV cs.GR 62%

Vid2Actor: Free-viewpoint Animatable Person Synthesis from Video in the Wild

Chung-Yi Weng, Brian Curless, Ira Kemelmacher-Shlizerman

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project Page: https://grail.cs.washington.edu/projects/vid2actor/ Supplementary Video: https://youtu.be/Zec8Us0v23o

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04639 2020-08-11 cs.CV cs.GR 62%

Visual Indeterminacy in GAN Art

Aaron Hertzmann

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Leonardo / SIGGRAPH 2020 Art Papers

Journal ref Leonardo, Volume 53, Issue 4, August 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02867 2020-04-07 cs.CV cs.GR 62%

Rethinking Spatially-Adaptive Normalization

Zhentao Tan, Dongdong Chen, Qi Chu, Menglei Chai, Jing Liao, Mingming He, Lu Yuan, Nenghai Yu

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11378 2019-11-27 cs.LG cs.CV cs.MM eess.IV stat.ML 62%

Text2FaceGAN: Face Generation from Fine Grained Textual Descriptions

Osaid Rehman Nasir, Shailesh Kumar Jha, Manraj Singh Grover, Yi Yu, Ajit Kumar, Rajiv Ratn Shah

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.02714 2019-08-12 cs.GR cs.CV 62%

Relighting Humans: Occlusion-Aware Inverse Rendering for Full-Body Human Images

Yoshihiro Kanamori, Yuki Endo

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Published at SIGGRAPH Asia 2018 (ACM Transactions on Graphics). Project page with codes, pretrained models, and human model lists is at http://kanamori.cs.tsukuba.ac.jp/projects/relighting_human/

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.11478 2019-06-28 cs.CV cs.GR 62%

A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Isaak Lim, Moritz Ibing, Leif Kobbelt

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Symposium on Geometry Processing 2019

Journal ref Computer Graphics Forum 38 (5), 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.12620 2019-04-30 cs.CV cs.CR cs.IR cs.MM 62%

AnonymousNet: Natural Face De-Identification with Measurable Privacy

Tao Li, Lei Lin

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM

Comments CVPR-19 Workshop on Computer Vision: Challenges and Opportunities for Privacy and Security (CV-COPS 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06601 2018-12-04 cs.CV cs.GR cs.LG 62%

Video-to-Video Synthesis

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, Bryan Catanzaro

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments In NeurIPS, 2018. Code, models, and more results are available at https://github.com/NVIDIA/vid2vid

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04139 2023-05-05 cs.CV cs.AI 61%

DALLE-URBAN: Capturing the urban design expertise of large text to image transformers

Sachith Seneviratne, Damith Senanayake, Sanka Rasnayaka, Rajith Vidanaarachchi, Jason Thompson

专题命中 文生图 :text-to-image(abstract);分类 cs.CV;diffusion(comments)

Comments Accepted to DICTA 2022, released 11000+ environmental scene images generated by Stable Diffusion and 1000+ images generated by DALLE-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17564 2026-08-19 cs.CV cs.AI 新提交 57%

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

新概念必须进入之处:统一多模态模型中的入口门跨任务可用性

Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Columbia University(哥伦比亚大学) CUHK(香港中文大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 该研究通过分离统一多模态模型的理解与生成任务方向,发现跨任务可用性取决于概念绑定的入口层,提出的对齐目标可在极低损失下实现概念跨任务迁移。

Comments 27 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15395 2026-08-18 cs.CV 新提交 57%

JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation

JoLT:用于上下文引导的高分辨率分块生成的联合潜在轨迹

Mathis Koroglu, Guillaume Jeanneret, Hugo Caselles-Dupré, Matthieu Cord, Arnaud Dapogny

机构 * Obvious Research(奥布弗西斯研究公司) Sorbonne Université(索邦大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 本文提出JoLT方法,通过联合去噪LR与HR潜在图像生成高分辨率图像,其生成的图像细节丰富、视觉效果佳,优于竞争基线,为艺术创作提供新方向。

Comments 25 pages, 10 figures, 7 tables. Accepted at the AI4VA Workshop at ECCV 2026. Project page: https://obvious-research.github.io/jolt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00410 2026-08-18 cs.AI cs.CL cs.CV 版本更新 57%

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

歧义去了哪里?探究多模态模型如何解释多义词

Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths

机构 * Princeton University(普林斯顿大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 该研究对比17个文本到图像模型和15个文本生成模型,发现多模态模型生成图像的词义多样性低于文本,揭示了基础模型在不同模态间意义表达的迁移 gap。

Comments Oral Presentation, Sci-FM Workshop @ COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13671 2026-08-17 cs.CV 新提交 57%

PROVE: Training-Free Prompt Recovery using Verifiable Evidence

PROVE:基于可验证证据的无训练提示词恢复

Rupayan Mallick, Mahsa Khoshnoodi, Sarah Adel Bargal

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 该研究提出无训练的黑盒提示词反转攻击PROVE,通过可验证场景描述重建提示词,在多数据集上优于基线方法,可用于版权保护相关研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10908 2026-08-12 cs.CV cs.CL 新提交 57%

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

顺序很重要:LVLMs作为图像序列时间推理的评判者

Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes

机构 * University of Bologna(博洛尼亚大学) NOVA School of Science and Technology(NOVA科技学院) NOVA Laboratory for Computer Science and Informatics(NOVA计算机科学与信息实验室)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 本研究发现大视觉语言模型(LVLMs)作为多模态评判者存在时间顺序判别缺陷,受首因、近因等结构性偏差影响,呼吁构建时间感知的视觉序列评估范式。

Comments 34 pages, camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06264 2026-08-07 cs.CV cs.LG eess.IV 新提交 57%

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

OTLesMix:用于生成形状与位置多样化合成病灶的Wasserstein重心与最优传输映射

Robin Trombetta, Carole Lartizien

机构 * Univ. Lyon(里昂大学) INSA Lyon(里昂国立应用科学学院) UCBL(里昂第一大学) CNRS(法国国家科学研究中心) Inserm(法国国家健康与医学研究院) CREATIS

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 本研究提出OTLesMix方法,利用Wasserstein重心与最优传输映射生成多样化合成病灶,在三项脑病灶分割任务上使Dice分数提升2.9至6.6个百分点,性能优于现有混合类方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05600 2026-08-07 cs.LG cs.AI cs.CV 新提交 57%

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

LC-GRPO:通过朗之万校正弥合基于流的GRPO的训练-推理差距

Yingqing Guo, Hui Yuan, Zijian He, Mengdi Wang, Zheng Ding

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 LC-GRPO 是带朗之万校正的基于流的 GRPO 框架,通过对齐推理的 ODE 欧拉步加朗之万校正,缩小流模型训练与推理的样本差距,在多任务上提升奖励优化并保留生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05233 2026-08-07 cs.AI cs.CV 新提交 57%

Coherence-Oriented Dream Scene Visualisation

面向连贯性的梦境场景可视化

Azra Açıl, Simon Colton

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦玛丽女王大学电子工程与计算机科学学院)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 该研究提出DSV系统,通过大型语言模型拆分梦境描述为四部分,结合文本到图像模型生成连贯的四面板图像序列,经DreamBank数据集50次可视化评估,用CLIP等模型指标验证了其质量、保真度与连贯性。

Comments short paper accepted at ICCC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00994 2026-08-04 cs.CV cs.CL 新提交 57%

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

零样本图像描述中合成监督的实体忠实修复

Zhiyue Liu, Wenkai Zhou, Jian Qin, Qipeng Jiang

机构 * Guangxi University(广西大学) School of Computer, Electronics and Information(计算机与电子信息学院)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 针对零样本图像描述中合成监督的实体级错位问题,提出即插即用框架ReCap,通过实体级重对齐与自适应加权策略优化合成数据,在两类基准上实现最优性能。

Comments Accepted to the 34th ACM International Conference on Multimedia (ACM MM 2026). 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00559 2026-08-04 cs.CV 新提交 57%

Test-Time Curriculum for Open-Set AIGC Detection

开放集AIGC检测的测试时课程

Yiqian Zhang, Zheyuan Gu, Xiangzhao Hao, Zefeng Zhang, Jingjia Mao, Jiahao Hu, Jiaxu Miao, Jun Yu, Zhenyu Zhang, Shuohuan Wang, Yu Sun

机构 * Baidu Inc.(百度公司)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 本研究针对开放集AIGC检测的分布偏移问题,提出Test-Time Curriculum框架,结合跨尺度伪标签细化技术,构建AIGCGuard基准,经实验验证可显著提升未见生成器偏移下的检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏