arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4201 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4201 篇

2512.18772 2025-12-23 cs.CV 79%

In-Context Audio Control of Video Diffusion Transformers

上下文音频控制视频扩散变换器

Wenze Liu, Weicai Ye, Minghong Cai, Quande Liu, Xintao Wang, Xiangyu Yue

机构 * MMLab, The Chinese University of Hong Kong(香港中文大学MMLab) Kling Team, Kuaishou Technology(快手科技Kling团队)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出ICAC框架,通过三维注意力机制实现音频驱动的视频生成,解决时间同步信号在视频扩散模型中的整合问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06833 2025-12-19 cs.CV 79%

ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search

ConsistTalk: 可控强度的时序一致说话头生成与扩散噪声搜索

Zhenjie Liu, Jianzhang Lu, Renjie Lu, Cong Liang, Shangfei Wang

机构 * The corresponding author.(通讯作者)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 ConsistTalk通过引入光流引导时间模块、音频到强度模型和扩散噪声初始化策略,实现了可控强度和时序一致的说话头生成,有效减少闪烁并提升音频视频同步质量。

Comments AAAI26 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14061 2025-12-17 cs.CV 79%

Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution

弥合保真度与现实的差距:可控制的一步扩散用于图像超分辨率

Hao Chen, Junyang Chen, Jinshan Pan, Jiangxin Dong

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 CODSR通过LQ引导特征调制、区域自适应生成先验激活和文本匹配指导策略,提升图像超分辨率的保真度与感知质量。

Comments Project page: https://github.com/Chanson94/CODSR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09561 2025-12-16 cs.CV 79%

TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control

TC-LoRA: 基于时间调节的条件LoRA用于自适应扩散控制

Minkyoung Cho, Ruben Ohana, Christian Jacobsen, Adityan Jothi, Min-Hung Chen, Z. Morley Mao, Ethem Can

机构 * University of Michigan(密歇根大学) NVIDIA(英伟达)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 TC-LoRA通过动态调整模型权重实现自适应扩散控制,提升生成保真度和空间条件符合性。

Comments Project Page: https://minkyoungcho.github.io/tc-lora/; NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE); 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17364 2025-12-16 cs.CV cs.AI 79%

Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation

条件编织与专家调节:迈向通用且可控的图像生成

Guoqing Zhang, Xingtong Ge, Lu Shi, Xin Zhang, Muqing Xue, Wanru Xu, Yigang Cen, Yidong Li

机构 * State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) School of Computer Science and Technology(计算机科学与技术学院) Visual Intellgence +X International Cooperation Joint Laboratory of MOE(教育部视觉智能+X国际合作联合实验室) Hong Kong University of Science and Technology(香港科技大学) SenseTime Research(商汤科技研究院) Beijing Jiaotong University(北京交通大学) SenseTime Research Institute(时光机器研究院)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

AI总结 提出UniGen框架,通过CoMoE模块和WeaveNet机制实现通用且可控的图像生成,提升效率和表达性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00227 2025-12-16 cs.CV cs.AI cs.RO 79%

Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes

Ctrl-Crash: 可控扩散用于逼真汽车碰撞

Anthony Gosselin, Ge Ya Luo, Luis Lara, Florian Golemo, Derek Nowrouzezahrai, Liam Paull, Alexia Jolicoeur-Martineau, Christopher Pal

机构 * McGill University(麦吉尔大学) CIFAR AI Chair(CIFAR人工智能 chair)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 Ctrl-Crash通过可控扩散生成逼真汽车碰撞视频,提升交通安全模拟的可控性和真实性。

Comments Under review at Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00996 2025-12-15 cs.CV 79%

Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models

具有时间推理的上下文微调用于视频扩散模型的多功能控制

Kinam Kim, Junha Hyung, Jaegul Choo

机构 * KAIST AI(韩国科学技术院人工智能研究中心)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 TIC-FT通过时间推理实现视频扩散模型的多功能控制,无需架构修改,仅需少量样本即可实现高质量视频生成。

Comments project page: https://kinam0252.github.io/TIC-FT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09095 2025-12-11 cs.CV 79%

Food Image Generation on Multi-Noun Categories

多名词类别的食物图像生成

Xinyue Pan, Yuhao Chen, Jiangpeng He, Fengqing Zhu

机构 * Purdue University(普渡大学) University of Waterloo(滑铁卢大学)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

AI总结 本文提出FoCULR方法,通过整合食品领域知识和早期引入核心概念,解决多名词类别食物图像生成中的语义误解和布局错误问题。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07170 2025-12-09 cs.CV cs.AI 79%

Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

迈向统一的语义和可控图像融合:一种扩散变换器方法

Jiayang Li, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma, Liqiang Nie

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Electronic Information School, Wuhan University(武汉大学电子信息学院)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 DiTFuse通过融合图像与自然语言指令,实现端到端、语义感知的图像融合,统一了多种融合任务并在多个基准测试中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04248 2025-12-05 cs.CV cs.AI 79%

MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models

MVRoom: 基于多视角扩散模型的可控3D室内场景生成

Shaoheng Fang, Chaohui Yu, Fan Wang, Qixing Huang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) DAMO Academy, Alibaba Group(阿里巴巴达摩院) Hupan Lab(虎斑实验室)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 MVRoom通过多视角扩散模型实现可控的3D室内场景生成,结合布局感知的极线注意力机制和迭代框架,提升多视角一致性和生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22864 2025-12-01 cs.CV 79%

ControlEvents: Controllable Synthesis of Event Camera Datawith Foundational Prior from Image Diffusion Models

ControlEvents: 基于图像扩散模型的基础先验的可控事件相机数据合成

Yixuan Hu, Yuxuan Xue, Simon Klenk, Daniel Cremers, Gerard Pons-Moll

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出ControlEvents,利用图像扩散模型生成高质量事件数据,通过文本标签、2D骨骼和3D姿态控制信号,减少标注数据成本并提升视觉任务性能。

Comments Accepted to WACV2026. Project website:https://yuxuan-xue.com/controlevents/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02465 2025-11-27 cs.CV 79%

Towards Consistent and Controllable Image Synthesis for Face Editing

面向面部编辑的图像合成一致性与可控性研究

Mengting Wei, Tuomas Varanka, Yante Li, Xingxun Jiang, Huai-Qian Khor, Guoying Zhao

机构 * Center for Machine Vision and Signal Analysis, Faculty of Information Technology and Electrical Engineering, University of Oulu(机器视觉与信号分析中心,信息科技与电气工程学院,奥卢大学) Key Laboratory of Child Development and Learning Science of Ministry of Education, School of Biological Sciences and Medical Engineering, Southeast University(教育部儿童发展与学习科学重点实验室,生物科学与医学工程学院,东南大学)

专题命中 可控生成 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 RigFace通过结合SD模型和3D面部模型,实现面部图像的光照、表情和姿态可控,提升身份保持和图像真实度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20218 2025-11-26 cs.CV 79%

Text-guided Controllable Diffusion for Realistic Camouflage Images Generation

基于文本引导的可控扩散生成逼真伪装图像

Yuhang Qian, Haiyan Chen, Wentong Li, Ningzhong Liu, Jie Qin

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出CT-CIG方法,通过文本引导和可控扩散生成逼真且逻辑合理的伪装图像,利用VLM和FIRM模块提升伪装图像质量。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16249 2025-11-26 cs.GR 79%

Controllable Layer Decomposition for Reversible Multi-Layer Image Generation

可控制的分层分解用于可逆的多层图像生成

Zihao Liu, Zunnan Xu, Shi Shu, Jun Zhou, Ruicheng Zhang, Zhenchao Tang, Xiu Li

专题命中 可控生成 :image generation(title);inpainting(abstract);分类 cs.GR

AI总结 本文提出CLD方法,通过可控分层分解实现多层图像的细粒度控制与精确生成,提升图像编辑的可控性和实用性。

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03596 2025-11-25 cs.CV 79%

ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning

ControlThinker: 通过视觉推理揭示潜在语义以实现可控图像生成

Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu, Jingjing Chen, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center on Intelligent Visual Computing(上海智能视觉计算协同创新中心) MiniMax

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

AI总结 ControlThinker通过视觉推理挖掘潜在语义,提升可控图像生成的语义一致性和视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16666 2025-11-21 cs.CV 79%

SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation

SceneDesigner: 9自由度姿态操控的可控多物体图像生成

Zhenyuan Qin, Xincheng Shuai, Henghui Ding

机构 * Fudan University(复旦大学)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

AI总结 SceneDesigner通过9自由度姿态操控实现多物体图像生成,采用CNOCS图和两阶段训练策略提升可控性与质量。

Comments NeurIPS 2025 (Spotlight), Project Page: https://henghuiding.com/SceneDesigner/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10825 2025-11-18 cs.CV 79%

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang, Chi Zhang, Xuelong Li

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2026. Our project page: https://tele-ai.github.io/OmniVDiff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12032 2025-11-18 cs.CV 79%

Improved Masked Image Generation with Knowledge-Augmented Token Representations

Guotao Liang, Baoquan Zhang, Zhiyuan Wen, Zihao Han, Yunming Ye

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10180 2025-11-13 cs.CV cs.LG 79%

CART: Compositional Auto-Regressive Transformer for Image Generation

Siddharth Roheda, Rohit Chowdhury, Aniruddha Bala, Rohan Jaiswal

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments figures compressed to meet arxiv size limit

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07987 2025-11-12 cs.CV 79%

CSF-Net: Context-Semantic Fusion Network for Large Mask Inpainting

Chae-Yeon Heo, Yeong-Jun Cho

机构 * Department of Artificial Intelligence Convergence(人工智能融合系)

专题命中 可控生成 :inpainting(title,abstract);分类 cs.CV

Comments 8 pages, 5 figures, Accepted to WACV 2026 (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07091 2025-11-11 cs.CV cs.AI 79%

How Bias Binds: Measuring Hidden Associations for Bias Control in Text-to-Image Compositions

Jeng-Lin Li, Ming-Ching Chang, Wei-Chao Chen

专题命中 可控生成 :text-to-image(title,abstract);分类 cs.CV

Comments Accepted for publication at the Alignment Track of The 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15530 2025-11-04 cs.RO cs.CV cs.LG 79%

VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation

Zehao Ni, Yonghao He, Lingfeng Qian, Jilei Mao, Fa Fu, Wei Sui, Hu Su, Junran Peng, Zhipeng Wang, Bin He

机构 * D-Robotics(D机器人) National Key Laboratory of Autonomous Intelligent Unmanned Systems(国家级自主智能无人机系统实验室) University of Science and Technology Beijing(北京科技大学) State Key Laboratory of Multimodal Artificial Intelligence System (MAIS) Institute of Automation of Chinese Academy of Sciences(多模态人工智能系统(MAIS)实验室,中国科学院自动化研究所) Frontiers Science Center for Intelligent Autonomous Systems(智能自主系统前沿科学中心) Shanghai Institute of Intelligent Science and Technology, Tongji University(上海智能科学技术研究院,同济大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00103 2025-11-04 cs.CV cs.AI 79%

FreeSliders: Training-Free, Modality-Agnostic Concept Sliders for Fine-Grained Diffusion Control in Images, Audio, and Video

Rotem Ezra, Hedi Zisling, Nimrod Berman, Ilan Naiman, Alexey Gorkor, Liran Nochumsohn, Eliya Nachmani, Omri Azencot

机构 * Faculty of Computer and Information Science, Ben-Gurion University of the Negev(计算机与信息科学学院,贝内-约尔大学纳格尔分校) Lightricks(Lightricks公司) School of Electrical and Computer Engineering, Ben-Gurion University of the Negev(电气与计算机工程学院,贝内-约尔大学纳格尔分校)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27169 2025-11-03 cs.CV 79%

DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model

Yucheng Xing, Jinxing Yin, Xiaodong Liu

机构 * Stony Brook University(石溪大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07316 2025-10-30 cs.CV 79%

Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers

Gangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang, Jingfeng Yao, Lianghui Zhu, Yuechuan Pu, Cheng Chi, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Sida Peng, Xin Yang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车) Zhejiang University(浙江大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments NeurIPS 2025. Project page: https://pixel-perfect-depth.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20803 2025-10-24 cs.CV 79%

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou

机构 * Ant Group(蚂蚁集团)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19272 2025-10-23 cs.CV 79%

SCEESR: Semantic-Control Edge Enhancement for Diffusion-Based Super-Resolution

Yun Kai Zhuang

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

Comments 10 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18716 2025-10-22 cs.CV 79%

SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation

Siyong Jian, Huan Wang

机构 * Westlake University(西湖大学)

专题命中 可控生成 :image generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17105 2025-10-21 cs.CV 79%

Boosting Fidelity for Pre-Trained-Diffusion-Based Low-Light Image Enhancement via Condition Refinement

Xiaogang Xu, Jian Wang, Yunfan Lu, Ruihang Chu, Ruixing Wang, Jiafei Wu, Bei Yu, Liang Lin

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学) Snap Research HKUST (GZ)(香港科技大学(广州)) Alibaba Tongyi Lab(阿里云联合实验室) camera group of DJI(大疆创新相机团队) The University of Hong Kong(香港大学) Sun Yat-Sen University(中山大学)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16772 2025-10-21 cs.CV cs.AI 79%

Region in Context: Text-condition Image editing with Human-like semantic reasoning

Thuy Phuong Vu, Dinh-Cuong Hoang, Minhhuy Le, Phan Xuan Tan

机构 * Greenwich Vietnam FPT University(越南格林威治FPT大学)

专题命中 可控生成 :image editing(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏