arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4201 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4201 篇

1912.05237 2020-03-25 cs.CV 83%

Towards Unsupervised Learning of Generative Models for 3D Controllable Image Synthesis

Yiyi Liao, Katja Schwarz, Lars Mescheder, Andreas Geiger

专题命中 可控生成 :image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.06221 2020-03-17 cs.CV cs.LG 83%

Semantic Pyramid for Image Generation

Assaf Shocher, Yossi Gandelsman, Inbar Mosseri, Michal Yarom, Michal Irani, William T. Freeman, Tali Dekel

专题命中 可控生成 :image generation(title,abstract);inpainting(abstract);分类 cs.CV

Journal ref IEEE Conference on Computer Vision and Pattern Recognition, 2020. CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.01908 2017-05-09 cs.CV 83%

Auto-painter: Cartoon Image Generation from Sketch by Using Conditional Generative Adversarial Networks

Yifan Liu, Zengchang Qin, Zhenbo Luo, Hua Wang

专题命中 可控生成 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16842 2026-05-19 cs.AI 82%

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models

草图然后绘画:用于扩散多模态大语言模型的分层强化学习

Siqi Luo, Jianghan Shen, Yi Xin, Huayu Zheng, Haoxing Chen, Yan Tai, Yue Li, Junjun He, Yihao Liu, Guangtao Zhai, Yuewen Cao, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院) Peking University(北京大学)

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出了一种分层强化学习方法HT-GRPO,通过Sketch-Then-Paint训练方案和分层信用分配机制,解决扩散多模态大语言模型在强化学习优化中的关键问题,提升图像质量和审美效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21666 2025-05-29 cs.LG cs.AI 82%

Efficient Controllable Diffusion via Optimal Classifier Guidance

Owen Oertell, Shikun Sun, Yiding Chen, Jin Peng Zhou, Zhiyong Wang, Wen Sun

机构 * Cornell University(康奈尔大学) CUHK(香港中文大学)

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract)

Comments 28 pages, 9 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09391 2025-01-17 cs.NI 82%

Contract-Inspired Contest Theory for Controllable Image Generation in Mobile Edge Metaverse

Guangyuan Liu, Hongyang Du, Jiacheng Wang, Dusit Niyato, Dong In Kim

专题命中 可控生成 :image generation(title,abstract);diffusion(abstract)

Comments 16 pages, 10figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13549 2026-04-16 cs.CV 81%

Reconstruction of a 3D wireframe from a single line drawing via generative depth estimation

通过生成深度估计从单张线稿中重建3D线框

Elton Cao, Hod Lipson

机构 * Creative Machines Lab, Columbia University(创意机器实验室,哥伦比亚大学)

专题命中 可控生成 :diffusion(summary_cn,abstract);分类 cs.CV

AI总结 本文提出一种生成方法,将线稿重建视为条件密集深度估计任务,利用Latent Diffusion Model和ControlNet风格的条件框架解决正投影的固有歧义,并通过图遍历的BFS掩码策略支持迭代的线稿-重建-线稿工作流,实现从稀疏2D线稿到密集3D表示的转换。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14741 2026-03-17 cs.CV 81%

PHAC: Promptable Human Amodal Completion

PHAC: 可提示的人形不可见区域补全

Seung Young Noh, Ju Yong Chang

专题命中 可控生成 :image generation(abstract);diffusion(abstract);inpainting(abstract);image synthesis(abstract)

AI总结 本文提出PHAC任务,通过满足可见外观约束和多用户提示完成遮挡人像,使用ControlNet模块编码提示信号并微调交叉注意力块以提升提示对齐性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03257 2025-08-05 cs.CV 81%

LACONIC: A 3D Layout Adapter for Controllable Image Creation

Léopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks Ovsjanikov

机构 * LIX, École Polytechnique, IP Paris(巴黎高等理工学院LIX研究所) Dassault Systèmes(达索系统)

专题命中 可控生成 :text-to-image(abstract);diffusion(abstract);image editing(abstract);image synthesis(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21416 2025-06-27 cs.CV 81%

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

Bowen Chen, Mengyi Zhao, Haomiao Sun, Li Chen, Xu Wang, Kang Du, Xinglong Wu

机构 * Intelligent Creation Team, ByteDance(字节跳动智能创造团队)

专题命中 可控生成 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments Project Page: https://bytedance.github.io/XVerse Github Link: https://github.com/bytedance/XVerse

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08111 2025-04-14 cs.CV 81%

POEM: Precise Object-level Editing via MLLM control

Marco Schouten, Mehmet Onurcan Kaya, Serge Belongie, Dim P. Papadopoulos

专题命中 可控生成 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

Comments Accepted to SCIA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04699 2025-01-09 cs.CV 81%

EditAR: Unified Conditional Generation with Autoregressive Models

Jiteng Mu, Nuno Vasconcelos, Xiaolong Wang

专题命中 可控生成 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

Comments Project page: https://jitengmu.github.io/EditAR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21541 2026-08-04 cs.GR cs.CV 版本更新 81%

ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering

ControlHair:协同物理模拟器与视频扩散实现可控动态毛发渲染

Weikai Lin, Haoxiang Li, Yuhao Zhu

机构 * University of Rochester(罗切斯特大学) Pixocial Technology(Pixocial技术)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 研究针对毛发模拟渲染难题,提出ControlHair框架,融合物理模拟器与视频扩散,通过三阶段管道实现可控动态毛发渲染,性能优于基线并展示多种应用。

Comments Accepted to ECCV 2026. 21 pages. Project page: https://linwk20.github.io/controlhair-web

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23679 2026-06-23 cs.CV cs.AI cs.GR cs.LG 新提交 81%

Semantic Browsing: Controllable Diversity for Image Generation

语义浏览:图像生成的可控多样性

Sara Dorfman, Maya Vishnevsky, Omer Dahary, Or Patashnik, Daniel Cohen-Or

机构 * Tel Aviv University(特拉维夫大学)

专题命中 可控生成 :image generation(title);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 针对文本到图像模型生成多样性不足且缺乏语义结构的问题,提出一种在文本层面引入多样性的方法,通过视觉语言模型和代理工作流实现用户可导航的语义变化空间。

Comments ECCV 2026. Project page: https://saradorfman1.github.io/SemanticBrowsing-webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26006 2026-06-03 cs.CV cs.GR cs.RO 81%

MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control

MIND: 多尺度意图扩散用于文本驱动的基于物理的人形控制

Bin Li, Ruichi Zhang, Han Liang, Jingyan Zhang, Juze Zhang, Xin Chen, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) University of Pennsylvania(宾夕法尼亚大学) Bytedance Seed(字节跳动种子) Stanford University(斯坦福大学) InstAdapt

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 提出MIND框架,通过多尺度意图扩散机制将文本命令与低级动作语义对齐,实现基于物理的人形机器人行为生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26796 2026-03-13 cs.CV cs.GR 81%

See4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting

See4D: 通过自回归视频修复实现无姿态4D生成

Dongyue Lu, Ao Liang, Tianxin Huang, Xiao Fu, Yuyang Zhao, Baorui Ma, Liang Pan, Wei Yin, Lingdong Kong, Wei Tsang Ooi, Ziwei Liu

机构 * NUS(新加坡国立大学) HKU(香港大学) CUHK(香港中文大学) THU(清华大学) Shanghai AI Lab(上海人工智能实验室) Horizon Robotics(地平线机器人) NTU(国立科技大学)

专题命中 可控生成 :inpainting(title,abstract);分类 cs.CV、cs.GR

AI总结 See4D通过自回归视频修复实现无姿态4D生成,提升从随意视频到4D世界建模的实用性。

Comments Eurographics2026; 26 pages; 21 figures; 3 tables; project page: https://see-4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05285 2026-03-03 cs.GR cs.CV 81%

Improved 3D Scene Stylization via Text-Guided Generative Image Editing with Region-Based Control

通过文本引导的生成图像编辑与基于区域的控制改进3D场景风格化

Haruo Fujiwara, Yusuke Mukuta, Tatsuya Harada

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所)

专题命中 可控生成 :image editing(title);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出了一种改进的3D场景风格化方法,通过文本引导和基于区域的控制,提升风格一致性和视图一致性。

Comments Project Page: https://haruolabs.github.io/improved-gs-style-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19381 2026-02-27 cs.CV cs.GR 81%

Enhancing Sketch Animation: Text-to-Video Diffusion Models with Temporal Consistency and Rigidity Constraints

增强草图动画:带有时间一致性和刚性约束的文本到视频扩散模型

Gaurav Rai, Ojaswa Sharma

机构 * Graphics Research Group, Indraprastha Institute of Information Technology Delhi(印度普拉斯塔信息技术研究所图形研究组)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 本文提出了一种基于文本到视频扩散模型的草图动画方法,通过引入时间一致性和刚性约束,提升了动画的平滑度和形状保持性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10287 2025-12-11 cs.SD cs.CV cs.GR eess.AS 81%

MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment

MACS:基于上下文意义和语义对齐的多源音频到图像生成

Hao Zhou, Xiaobao Guo, Yuzhe Zhu, Adams Wai-Kin Kong

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV、cs.GR

AI总结 MACS通过分离多源音频并利用语义对齐提升音频到图像生成的质量和表现。

Comments Accepted at AAAI 2026. Code available at https://github.com/alxzzhou/MACS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21129 2025-11-27 cs.CV cs.GR 81%

CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion

通过统一多模态视频扩散实现可控视频生成

Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu, Jialun Liu, Hao Pan, Yuchi Huo, Rui Wang, Haibin Huang, Chi Zhang, Xuelong Li

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院) Tsinghua University(清华大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 CtrlVDiff通过统一多模态视频扩散模型,实现可控视频生成,支持多模态输入并提升生成的可控性和保真度。

Comments 27 pages, 18 figures, 9 tables. Project page: https://tele-ai.github.io/CtrlVDiff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18597 2025-09-09 cs.GR cs.CV 81%

SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis

Xiaohao Sun, Divyam Goel, Angel X. Chang

机构 * Simon Fraser University(西蒙弗雷泽大学) CMU(卡内基梅隆大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

Comments Project page: https://3dlg-hcvc.github.io/SemLayoutDiff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05621 2025-07-09 cs.CV cs.MM 81%

AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework

Suoxiang Zhang, Xiaxi Li, Hongrui Chang, Zhuoyan Hou, Guoxin Wu, Ronghua Ji

机构 * China Agricultural University(中国农业大学)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14018 2025-06-24 cs.CV cs.AI cs.MM cs.RO 81%

SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation

Tong Chen, Shuya Yang, Junyi Wang, Long Bai, Hongliang Ren, Luping Zhou

机构 * The University of Sydney, Sydney, Australia(悉尼大学) The University of Hong Kong, Hong Kong SAR, China(香港大学) The Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.MM

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15312 2025-06-19 cs.GR cs.CR cs.CV cs.CY 81%

One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning

Han Wu, Junyao Li, Kangbo Zhao, Sen Zhang, Yukai Shi, Liang Lin

机构 * Guangdong University of Technology(广东工业大学) Sun Yat-sen University(中山大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

Comments We propose a novel framework for face sketch synthesis, where merely a single pair of samples suffices to enable in-the-wild face sketch synthesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03150 2025-06-04 cs.CV cs.AI cs.LG cs.MM 81%

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark, Ming-Hsuan Yang

机构 * University of Oxford(牛津大学) UC Merced(加州大学默塞德分校) NEC Labs America(NEC美国实验室) Atmanity Inc.(Atmanity公司) Google DeepMind Project(谷歌DeepMind项目)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.MM

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21146 2025-05-28 cs.GR cs.CV 81%

IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model

Yang Zhao, Yan Zhang, Xubo Yang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14728 2025-02-19 cs.CV cs.MM 81%

Semantically Consistent Person Image Generation

Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal, Michael Blumenstein

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV、cs.MM

Comments Accepted in The International Conference on Pattern Recognition (ICPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02717 2025-02-19 cs.CV cs.MM 81%

Scene Aware Person Image Generation through Global Contextual Conditioning

Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal, Michael Blumenstein

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV、cs.MM

Comments Accepted in The International Conference on Pattern Recognition (ICPR) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05413 2025-01-10 cs.SD cs.CV cs.GR eess.AS 81%

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation

Darius Petermann, Mahdi M. Kalayeh

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13609 2024-12-20 cs.CV cs.MM 81%

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

Shengeng Tang, Jiayi He, Dan Guo, Yanyan Wei, Feng Li, Richang Hong

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏