arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4201 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4201 篇

2510.15874 2025-10-21 cs.GR 79%

Sketch-based Fluid Video Generation Using Motion-Guided Diffusion Models in Still Landscape Images

Hao Jin, Haoran Xie

专题命中 可控生成 :diffusion(title,abstract);分类 cs.GR

Comments 2 pages, 5 figures. SIGGRAPH 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04126 2025-10-20 cs.CV cs.AI 79%

Multi-identity Human Image Animation with Structural Video Diffusion

Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Yuwei Guo, Dahua Lin, Tianfan Xue, Bo Dai

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments ICCV 2025 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11111 2025-10-20 cs.CV 79%

Diffusion Models are Efficient Data Generators for Human Mesh Recovery

Yongtao Ge, Wenjia Wang, Yongfan Chen, Fanzhou Wang, Lei Yang, Hao Chen, Chunhua Shen

机构 * The University of Adelaide(阿德莱德大学) Zhejiang University(浙江大学) Zhejiang University of Technology(浙江工业大学) The University of Hong Kong(香港大学) SenseTime(秒针科技)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by TPAMI, project page: https://yongtaoge.github.io/projects/humanwild

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14179 2025-10-17 cs.CV cs.AI 79%

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

Yuancheng Xu, Wenqi Xian, Li Ma, Julien Philip, Ahmet Levent Taşel, Yiwei Zhao, Ryan Burgert, Mingming He, Oliver Hermann, Oliver Pilarski, Rahul Garg, Paul Debevec, Ning Yu

机构 * Eyeline Labs(Eyeline实验室)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07203 2025-10-16 cs.CV 79%

Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion

Xingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang, Qiang Zhang, Shunsi Zhang

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Journal ref Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:41791-41806, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11346 2025-10-14 cs.CV cs.AI 79%

Uncertainty-Aware ControlNet: Bridging Domain Gaps with Synthetic Image Generation

Joshua Niemeijer, Jan Ehrhardt, Heinz Handels, Hristina Uzunova

机构 * German Aerospace Center (DLR)(德国航空航天中心) University of Lübeck(吕贝克大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 可控生成 :image generation(title);diffusion(abstract);分类 cs.CV

Comments Accepted for presentation at ICCV Workshops 2025, "The 4th Workshop on What is Next in Multimodal Foundation Models?" (MMFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15117 2025-10-13 cs.CV 79%

Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer

Minh Bui, Kostas Alexis

机构 * Norwegian University of Science and Technology (NTNU)(挪威科学技术大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21593 2025-10-13 cs.CV cs.AI 79%

Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model

Yang Yang, Siming Zheng, Qirui Yang, Jinwei Chen, Boxi Wu, Xiaofei He, Deng Cai, Bo Li, Peng-Tao Jiang

机构 * Zhejiang University(浙江大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments project page: https://vivocameraresearch.github.io/any2bokeh/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07340 2025-10-10 cs.GR cs.LG 79%

SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation

Yongzhi Li, Saining Zhang, Yibing Chen, Boying Li, Yanxin Zhang, Xiaoyu Du

专题命中 可控生成 :image generation(title,abstract);分类 cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04220 2025-10-07 cs.CV cs.AI cs.LG 79%

MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering

Lixuan He, Shikang Zheng, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22139 2025-09-29 cs.CV cs.AI 79%

REFINE-CONTROL: A Semi-supervised Distillation Method For Conditional Image Generation

Yicheng Jiang, Jin Yuan, Hua Yuan, Yao Zhang, Yong Rui

机构 * School of Computer Science and Engineering(计算机科学与工程学院) AI Lab(人工智能实验室) Lenovo Research(联想研究院)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments 5 pages,17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13065 2025-09-26 cs.CV 79%

Odo: Depth-Guided Diffusion for Identity-Preserving Body Reshaping

Siddharth Khandelwal, Sridhar Kamath, Arjun Jain

机构 * Fast Code AI Consult Pvt. Ltd.(Fast Code AI咨询私有有限公司)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23711 2025-09-23 cs.CV 79%

Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion

Haoyang Chen, Dongfang Sun, Caoyuan Ma, Shiqin Wang, Kewei Zhang, Zheng Wang, Zhixiang Wang

机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University(国家多媒体软件工程研究中心,人工智能研究院,计算机科学学院,武汉大学) Hubei Key Laboratory of Multimedia and Network Communication Engineering(湖北省多媒体与网络通信工程重点实验室) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国) CyberAgent AI Lab, Japan(CyberAgent AI实验室,日本)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15185 2025-09-19 cs.CV 79%

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

Xiaoyu Yue, Zidong Wang, Yuqing Wang, Wenlong Zhang, Xihui Liu, Wanli Ouyang, Lei Bai, Luping Zhou

机构 * Shanghai AI Laboratory(上海人工智能实验室) University of Sydney(悉尼大学) Chinese University of Hong Kong(香港中文大学) University of Hong Kong(香港大学)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13756 2025-09-18 cs.CV 79%

Controllable-Continuous Color Editing in Diffusion Model via Color Mapping

Yuqi Yang, Dongliang Chang, Yuanchen Fang, Yi-Zhe SonG, Zhanyu Ma, Jun Guo

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学) SketchX, CVSSP, University of Surrey(SketchX、CVSSP、 Surrey大学)

专题命中 可控生成 :diffusion(title);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07809 2025-09-10 cs.CV 79%

SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting

Mahtab Dahaghin, Milind G. Padalkar, Matteo Toso, Alessio Del Bue

机构 * Pattern Analysis and Computer Vision (PAVIS)(模式分析与计算机视觉(PAVIS)) Istituto Italiano di Tecnologia (IIT)(意大利技术研究所(IIT))

专题命中 可控生成 :inpainting(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02983 2025-09-04 cs.RO cs.CV 79%

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features

Jinghe Yang, Minh-Quan Le, Mingming Gong, Ye Pu

机构 * Department of Electrical and Electronic Engineering, The University of Melbourne, Australia(墨尔本大学电子与电气工程系) School of Mathematics and Statistics, The University of Melbourne, Australia(墨尔本大学数学与统计学学院)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01107 2025-09-03 cs.CV 79%

FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation

Wenzhuang Wang, Yifan Zhao, Mingcan Ma, Ming Liu, Zhonglin Jiang, Yong Chen, Jia Li

机构 * State Key Laboratory of Virtual Reality Technology and Systems, SCSE&QRI, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学) Geely Automobile Research Institute (Ningbo) Co., Ltd(吉利汽车研究院(宁波)有限公司)

专题命中 可控生成 :image generation(title);diffusion(abstract);分类 cs.CV

Comments 21 pages, 19 figures, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00428 2025-09-03 cs.CV 79%

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

Xuechao Zou, Shun Zhang, Xing Fu, Yue Li, Kai Li, Yushe Cao, Congyan Lang, Pin Tao, Junliang Xing

机构 * Beijing Jiaotong University(北京交通大学) Ant Group(蚂蚁集团) Qinghai University(青海大学) Tsinghua University(清华大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19575 2025-08-29 cs.CV cs.AI 79%

Interact-Custom: Customized Human Object Interaction Image Generation

Zhu Xu, Zhaowen Wang, Yuxin Peng, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) Adobe Research(Adobe研究)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18835 2025-08-27 quant-ph cs.CV 79%

Quantum-Circuit-Based Visual Fractal Image Generation in Qiskit and Analytics

Hillol Biswas

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06905 2025-08-27 cs.CV 79%

MultiRef: Controllable Image Generation with Multiple Visual References

Ruoxi Chen, Dongping Chen, Siyuan Wu, Sinan Wang, Shiyun Lang, Petr Sushko, Gaoyang Jiang, Yao Wan, Ranjay Krishna

机构 * Zhejiang Wanli University(浙江万里大学) University of Washington(华盛顿大学) Huazhong University of Science and Technology(华中科技大学) Allen Institute for AI(人工智能研究院)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Accepted to ACM MM 2025 Datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13188 2025-08-27 cs.CV 79%

StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

Yunzhi Yan, Zhen Xu, Haotong Lin, Haian Jin, Haoyu Guo, Yida Wang, Kun Zhan, Xianpeng Lang, Hujun Bao, Xiaowei Zhou, Sida Peng

机构 * Zhejiang University(浙江大学) Li Auto Inc.(Li汽车公司) Cornell University(康奈尔大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Project page: https://zju3dv.github.io/street_crafter

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08736 2025-08-26 cs.CV 79%

GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation

Tianwei Xiong, Jun Hao Liew, Zilong Huang, Jiashi Feng, Xihui Liu

机构 * The University of Hong Kong(香港大学) ByteDance Seed Project(字节跳动种子项目)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://silentview.github.io/GigaTok

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03144 2025-08-22 cs.CV 79%

LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing

Liangyang Ouyang, Jiafeng Mao

专题命中 可控生成 :image editing(title,abstract);分类 cs.CV

Comments Our implementation is available at https://github.com/oyly16/LORE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11153 2025-08-18 cs.CV 79%

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

Maoquan Zhang, Bisser Raytchev, Xiujuan Sun

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(Hiroshima大学研究生院) Department of Computer Science, Weifang University of Science and Technology(潍坊科技大学计算机科学系)

专题命中 可控生成 :image generation(title);diffusion(abstract);分类 cs.CV

Comments The International Conference on Neural Information Processing (ICONIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10280 2025-08-15 cs.CV 79%

High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance

Danyi Gao

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08949 2025-08-13 cs.CV 79%

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

Ao Ma, Jiasong Feng, Ke Cao, Jing Wang, Yun Wang, Quanwei Zhang, Zhanjie Zhang

机构 * JD.com, Inc.(京东公司)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07246 2025-08-12 cs.CV 79%

Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers

Xin Ma, Yaohui Wang, Genyun Jia, Xinyuan Chen, Tien-Tsin Wong, Cunjian Chen

机构 * Department of Data Science & AI, Faculty of Information Technology, Monash University(数据科学与人工智能系,信息科技学院,墨尔本大学) Nanjing University of Posts and Telecommunications(南京邮电大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Project Page: https://maxin-cn.github.io/miramo_project

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23186 2025-08-12 cs.CV 79%

HiGarment: Cross-modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment Image

Junyi Guo, Jingxuan Zhang, Fangyu Wu, Huanda Lu, Qiufeng Wang, Wenmian Yang, Eng Gee Lim, Dongming Lu

机构 * Xi’an Jiaotong Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Beijing Normal University(北京师范大学) Zhejiang University(浙江大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏