arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4201 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4201 篇

2503.18393 2025-03-25 cs.CV 79%

PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes

Xinhua Xu, Hong Liu, Jianbing Wu, Jinfu Liu

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12781 2025-03-25 cs.CV 79%

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace, Guocheng Qian, Michael Vasilkovsky, Hsin-Ying Lee, Chaoyang Wang, Jiaxu Zou, Andrea Tagliasacchi, David B. Lindell, Sergey Tulyakov

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments ICLR 2025; Project Page: https://snap-research.github.io/vd3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15617 2025-03-21 cs.CV cs.AI 79%

CAM-Seg: A Continuous-valued Embedding Approach for Semantic Image Generation

Masud Ahmed, Zahid Hasan, Syed Arefinul Haque, Abu Zaher Md Faridee, Sanjay Purushotham, Suya You, Nirmalya Roy

专题命中 可控生成 :image generation(title);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11236 2025-03-20 cs.CV 79%

Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling

Guiyu Zhang, Huan-ang Gao, Zijian Jiang, Hao Zhao, Zhedong Zheng

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02206 2025-03-05 cs.CV 79%

Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation

Zhichao Yang, Leida Li, Pengfei Chen, Jinjian Wu, Giuseppe Valenzise

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19894 2025-03-03 cs.CV 79%

High-Fidelity Relightable Monocular Portrait Animation with Lighting-Controllable Video Diffusion Model

Mingtao Guo, Guanyu Xing, Yanli Liu

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18364 2025-02-26 cs.CV 79%

ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation

Yifan Pu, Yiming Zhao, Zhicong Tang, Ruihong Yin, Haoxing Ye, Yuhui Yuan, Dong Chen, Jianmin Bao, Sirui Zhang, Yanbin Wang, Lin Liang, Lijuan Wang, Ji Li, Xiu Li, Zhouhui Lian, Gao Huang, Baining Guo

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Project page: https://art-msra.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08642 2025-02-13 cs.CV 79%

SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation

Ellie Arar, Yarden Frenkel, Daniel Cohen-Or, Ariel Shamir, Yael Vinker

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments https://swiftsketch.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01666 2025-02-05 cs.CV cs.LG 79%

Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding

Jingming Xia, Guanqun Cao, Guang Ma, Yiben Luo, Qinzhao Li, John Oyekan

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04052 2025-01-22 cs.CV 79%

Beyond Imperfections: A Conditional Inpainting Approach for End-to-End Artifact Removal in VTON and Pose Transfer

Aref Tabatabaei, Zahra Dehghanian, Maryam Amirmazlaghani

专题命中 可控生成 :inpainting(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06769 2025-01-14 cs.CV 79%

ODPG: Outfitting Diffusion with Pose Guided Condition

Seohyun Lee, Jintae Park, Sanghyeok Park

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments 11 pages, 5 figures. Preprint submitted to VISAPP 2025: the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03630 2025-01-13 cs.CV 79%

MC-VTON: Minimal Control Virtual Try-On Diffusion Transformer

Junsheng Luan, Guangyuan Li, Lei Zhao, Wei Xing

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02519 2025-01-07 cs.CV 79%

Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors

Minglin Chen, Longguang Wang, Sheng Ao, Ye Zhang, Kai Xu, Yulan Guo

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16839 2024-12-25 cs.CV cs.AI 79%

Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets

Changjian Chen, Fei Lv, Yalong Guan, Pengcheng Wang, Shengjie Yu, Yifan Zhang, Zhuo Tang

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Accepted by TVCG2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14168 2024-12-20 cs.CV 79%

FashionComposer: Compositional Fashion Image Generation

Sihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu, Hao Luo, Hengshuang Zhao

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments https://sihuiji.github.io/FashionComposer-Page

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11512 2024-12-17 cs.CV 79%

SpatialMe: Stereo Video Conversion Using Depth-Warping and Blend-Inpainting

Jiale Zhang, Qianxi Jia, Yang Liu, Wei Zhang, Wei Wei, Xin Tian

专题命中 可控生成 :inpainting(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11420 2024-12-17 cs.CV 79%

Category Level 6D Object Pose Estimation from a Single RGB Image using Diffusion

Adam Bethell, Ravi Garg, Ian Reid

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07774 2024-12-13 cs.CV 79%

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Xi Chen, Zhifei Zhang, He Zhang, Yuqian Zhou, Soo Ye Kim, Qing Liu, Yijun Li, Jianming Zhang, Nanxuan Zhao, Yilin Wang, Hui Ding, Zhe Lin, Hengshuang Zhao

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments webpage: https://xavierchen34.github.io/UniReal-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04424 2024-12-06 cs.CV cs.AI 79%

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

Jiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li, Jianfeng Gao, Tianyi Zhou, Bin Xiao

专题命中 可控生成 :generative vision(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01801 2024-12-04 cs.CV 79%

SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation

Alexey Bokhovkin, Quan Meng, Shubham Tulsiani, Angela Dai

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments 21 pages, 12 figures; https://alexeybokhovkin.github.io/scenefactor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00835 2024-12-03 cs.CV 79%

Particle-based 6D Object Pose Estimation from Point Clouds using Diffusion Models

Christian Möller, Niklas Funk, Jan Peters

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19005 2024-12-02 cs.CV 79%

Locally-Focused Face Representation for Sketch-to-Image Generation Using Noise-Induced Refinement

Muhammad Umer Ramzan, Ali Zia, Abdelwahed Khamis, yman Elgharabawy, Ahmad Liaqat, Usman Ali

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Paper accepted for publication in 25th International Conference on Digital Image Computing: Techniques & Applications (DICTA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15233 2024-11-26 cs.CV 79%

LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer

Yu Li, Yifan Chen, Gongye Liu, Fei Yin, Qingyan Bai, Jie Wu, Hongfa Wang, Ruihang Chu, Yujiu Yang

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14781 2024-11-25 cs.CV 79%

Reconciling Semantic Controllability and Diversity for Remote Sensing Image Synthesis with Hybrid Semantic Embedding

Junde Liu, Danpei Zhao, Bo Yuan, Wentao Li, Tian Li

专题命中 可控生成 :image synthesis(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12290 2024-11-20 cs.CV cs.AI 79%

SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model

Haowen Zheng, Yanyan Liang

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01956 2024-11-13 cs.CV 79%

Enhance Image-to-Image Generation with LLaVA-generated Prompts

Zhicheng Ding, Panfeng Li, Qikai Yang, Siyang Li

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Accepted by 2024 5th International Conference on Information Science, Parallel and Distributed Systems

Journal ref Proceedings of the 2024 5th International Conference on Information Science, Parallel and Distributed Systems (ISPDS), 2024, pp. 77-81

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02332 2024-11-12 cs.CV 79%

UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control

Tian Xia, Xuweiyi Chen, Sihan Xu

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to TMLR | Project Page: https://unified-attention-control.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17952 2024-11-07 cs.CV 79%

BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation

Xiang Zhang, Bingxin Ke, Hayko Riemenschneider, Nando Metzger, Anton Obukhov, Markus Gross, Konrad Schindler, Christopher Schroers

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01770 2024-10-31 cs.CV 79%

StyleAdapter: A Unified Stylized Image Generation Model

Zhouxia Wang, Xintao Wang, Liangbin Xie, Zhongang Qi, Ying Shan, Wenping Wang, Ping Luo

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Accepted by IJCV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14617 2024-10-28 cs.CV cs.AI cs.LG 79%

Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion

Xiang Fan, Anand Bhattad, Ranjay Krishna

专题命中 可控生成 :diffusion(title);inpainting(abstract);分类 cs.CV

Comments Accepted to ECCV 2024. Project page: https://videoshop-editing.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏