arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 4201 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 4201 篇

2602.21824 2026-02-26 cs.LG 78%

DocDjinn: Controllable Synthetic Document Generation with VLMs and Handwriting Diffusion

DocDjinn: 基于VLMs和手写扩散的可控合成文档生成

Marcel Lamott, Saifullah Saifullah, Nauman Riaz, Yves-Noel Weweler, Tobias Alt-Veit, Ahmad Sarmad Ali, Muhammad Armaghan Shakir, Adrian Kalwa, Momina Moetesum, Andreas Dengel, Sheraz Ahmed, Faisal Shafait, Ulrich Schwanecke, Adrian Ulges

机构 * RheinMain University of Applied Sciences(莱茵-美因应用科学大学) German Research Center for Artificial Intelligence(德国人工智能研究中心) DeepReader GmbH(DeepReader公司) Insiders Technologies GmbH(Insiders Technologies公司) National University of Sciences and Technology(国家科学与技术大学)

专题命中 可控生成 :diffusion(title,abstract)

AI总结 DocDjinn利用VLMs和手写扩散生成可控合成文档,通过聚类和参数化采样从未标注种子生成高质量标注文档,实现隐私保护和高效数据生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13786 2026-02-10 cs.SD cs.AI 78%

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

DegDiT:基于动态事件图引导的扩散变换器的可控音频生成

Yisu Liu, Chenxing Li, Wanqian Zhang, Wenfu Wang, Meng Yu, Ruibo Fu, Zheng Lin, Weiping Wang, Dong Yu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Tencent, AI Lab(腾讯AI实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 可控生成 :diffusion(title,abstract)

AI总结 DegDiT通过动态事件图引导的扩散变换器框架,实现了开放式词汇可控音频生成,提升了生成音频的准确性和多样性。

Journal ref IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06823 2026-01-13 cs.HC 78%

Generative Modeling of Human-Computer Interfaces with Diffusion Processes and Conditional Control

基于扩散过程和条件控制的人机界面生成模型

Rui Liu, Liuqingqing Yang, Runsheng Zhang, Shixiao Wang

专题命中 可控生成 :diffusion(title,abstract)

AI总结 本文提出基于扩散过程和条件控制的界面生成模型,通过整合用户意图和任务约束,提升界面生成的多样性、合理性和智能化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01045 2026-01-07 cs.LG 78%

Coarse-Grained Kullback--Leibler Control of Diffusion-Based Generative AI

粗粒度Kullback-Leibler控制的扩散基生成AI

Tatsuaki Tsuruyama

机构 * Department of Physics, Tohoku University(东大理学部) Department of Drug Discovery Medicine, Kyoto University(京都大学药学部)

专题命中 可控生成 :diffusion(title,abstract)

AI总结 本文提出了一种基于V-delta势能的反向扩散方案,通过控制粗粒度量来提升生成图像的质量和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21070 2025-12-11 math.NA cs.NA 78%

Numerical analysis of a stabilized scheme for an optimal control problem governed by a parabolic convection--diffusion equation

对由抛物型对流-扩散方程所支配的最优控制问题的稳定化方案的数值分析

Christos Pervolianakis

专题命中 可控生成 :diffusion(title,abstract)

AI总结 本文研究了由抛物型对流-扩散方程支配的最优控制问题,提出稳定化方案并推导了误差估计,通过数值实验验证了其收敛性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12175 2025-12-09 cs.SD eess.AS 78%

Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis

音频调色盘:一种多信号条件的扩散变压器用于可控的 Foley 合成

Junnuo Wang

机构 * New York University(纽约大学)

专题命中 可控生成 :diffusion(title,abstract)

AI总结 Audio Palette通过多信号条件和LoRA技术实现可控Foley合成,提供高效、可解释的音频生成与控制方法。

Comments Accepted for publication in the Artificial Intelligence Technology Research (AITR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16900 2025-11-24 eess.SY cs.SY 78%

When Motion Learns to Listen: Diffusion-Prior Lyapunov Actor-Critic Framework with LLM Guidance for Stable and Robust AUV Control in Underwater Tasks

当运动学会倾听:带有LLM引导的扩散先验Lyapunov动作-批评者框架用于水下任务中稳定和鲁棒的AUV控制

Jingzehua Xu, Weiyi Liu, Weihang Zhang, Zhuofan Xi, Guanwen Xie, Shuai Zhang, Yi Li

专题命中 可控生成 :diffusion(title,abstract)

AI总结 本文提出一种结合扩散模型、Lyapunov批评者和LLM的框架,用于提升水下机器人控制的稳定性与鲁棒性,通过生成-过滤-优化机制实现高效探索和多目标优化。

Comments This paper is currently under review and does not represent the final version. Jingzehua Xu, Weiyi Liu and Weihang Zhang are co-first authors of this paper, with Zhuofan Xi as the second author

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27555 2025-11-03 math.AP 78%

On the global existence and uniform-in-time bounds for three-component reaction-diffusion systems with mass control and polynomial growth

Redouane Douaifia, Salem Abdelmalek, Mokhtar Kirane

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24591 2025-11-03 cs.RO cs.AI 78%

PoseDiff: A Unified Diffusion Model Bridging Robot Pose Estimation and Video-to-Action Control

Haozhuo Zhang, Michele Caprio, Jing Shao, Qiang Zhang, Jian Tang, Shanghang Zhang, Wei Pan

专题命中 可控生成 :diffusion(title,abstract)

Comments The experimental setup and metrics lacks rigor, affecting the fairness of the comparisons

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14455 2025-10-23 cs.CL cs.AI 78%

CtrlDiff: Boosting Large Diffusion Language Models with Dynamic Block Prediction and Controllable Generation

Chihan Huang, Hao Tang

机构 * Nanjing University of Science and Technology(南京理工大学) Peking University(北京大学)

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17122 2025-10-21 cs.LG math.OC 78%

Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control

Chengxiu Hua, Jiawen Gu, Yushun Tang

机构 * Southern University of Science and Technology(南方科技大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07135 2025-08-12 cs.HC 78%

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas

Runlin Duan, Yuzhao Chen, Rahul Jain, Yichen Hu, Jingyu Shi, Karthik Ramani

专题命中 可控生成 :image generation(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21816 2025-07-30 eess.IV 78%

Control Copy-Paste: Controllable Diffusion-Based Augmentation Method for Remote Sensing Few-Shot Object Detection

Yanxing Liu, Jiancheng Pan, Bingchen Zhang

专题命中 可控生成 :diffusion(title,abstract)

Comments 5 Pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17428 2025-07-08 eess.IV 78%

Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications

Chengyang Liang, Dong Li

专题命中 可控生成 :image generation(title);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00971 2025-06-23 physics.med-ph 78%

Deep learning-based Fast Volumetric Image Generation for Image-guided Proton FLASH Radiotherapy

Chih-Wei Chang, Yang Lei, Tonghe Wang, Sibo Tian, Justin Roper, Liyong Lin, Jeffrey Bradley, Tian Liu, Jun Zhou, Xiaofeng Yang

专题命中 可控生成 :image generation(title,abstract)

Journal ref IEEE Transactions on Radiation and Plasma Medical Sciences, vol. 8, no. 8, pp. 973-983, Nov. 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08278 2025-06-19 eess.SY cs.SY 78%

Toward Near-Globally Optimal Nonlinear Model Predictive Control via Diffusion Models

Tzu-Yuan Huang, Armin Lederer, Nicolas Hoischen, Jan Brüdigam, Xuehua Xiao, Stefan Sosnowski, Sandra Hirche

专题命中 可控生成 :diffusion(title,abstract)

Comments This paper has been accepted by the 2025 7th Annual Learning for Dynamics & Control Conference (L4DC) as an oral presentation and has been nominated for the best paper award

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07879 2025-06-05 cs.HC 78%

Towards Sustainable Creativity Support: An Exploratory Study on Prompt Based Image Generation

Daniel Hove Paludan, Julie Fredsgård, Kasper Patrick Bährentz, Ilhan Aslan, Niels van Berkel

专题命中 可控生成 :image generation(title,abstract)

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03025 2025-05-30 math.OC math.AP 78%

Optimal control of the fidelity coefficient in a Cahn-Hilliard image inpainting model

Elena Beretta, Cecilia Cavaterra, Matteo Fornoni, Maurizio Grasselli

专题命中 可控生成 :inpainting(title,abstract)

Comments 40 pages, revised version, to appear in ESAIM: Control Optim. Calc. Var

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09283 2025-05-15 cs.HC 78%

A Note on Semantic Diffusion

Alexander P. Ryjov, Alina A. Egorova

专题命中 可控生成 :diffusion(title,abstract)

Comments 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00965 2025-05-15 cs.RO 78%

SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation

Cheng-Chun Hsu, Bowen Wen, Jie Xu, Yashraj Narang, Xiaolong Wang, Yuke Zhu, Joydeep Biswas, Stan Birchfield

机构 * NVIDIA University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California San Diego(加州大学圣地亚哥分校)

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13196 2025-05-15 quant-ph eess.IV 78%

Quantum Generative Learning for High-Resolution Medical Image Generation

Amena Khatun, Kübra Yeter Aydeniz, Yaakov S. Weinstein, Muhammad Usman

专题命中 可控生成 :image generation(title,abstract)

Journal ref Mach. Learn.: Sci. Technol. 6 025032, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07980 2025-05-14 cs.CL 78%

Task-Adaptive Semantic Communications with Controllable Diffusion-based Data Regeneration

Fupei Guo, Achintha Wijesinghe, Songyang Zhang, Zhi Ding

机构 * University of Louisiana at Lafayette(路易斯安那大学拉斐特分校) University of California at Davis(加州大学戴维斯分校)

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08739 2025-04-15 cs.IR cs.HC 78%

Enhancing Product Search Interfaces with Sketch-Guided Diffusion and Language Agents

Edward Sun

专题命中 可控生成 :diffusion(title,abstract)

Comments Companion Proceedings of the ACM Web Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18057 2025-01-31 math.AP math.OC 78%

Stochastic scattering control of spider diffusion governed by an optimal diffraction probability measure selected from its own local-time

Isaac Ohavi

专题命中 可控生成 :diffusion(title,abstract)

Comments 53 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10869 2025-01-22 cs.LG cs.RO 78%

Diffusion-Based Imitation Learning for Social Pose Generation

Antonio Lech Martin-Ozimek, Isuru Jayarathne, Su Larb Mon, Jouh Yeong Chew

专题命中 可控生成 :diffusion(title,abstract)

Comments This paper was submitted as an LBR to HRI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09279 2025-01-17 cs.AI 78%

Text Semantics to Flexible Design: A Residential Layout Generation Method Based on Stable Diffusion Model

Zijin Qiu, Jiepeng Liu, Yi Xia, Hongtuo Qi, Pengkun Liu

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05151 2025-01-17 eess.AS cs.SD 78%

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue, Ying Shan, Mangsuo Zhao, Chao Zhang

专题命中 可控生成 :diffusion(title,abstract)

Comments Accepted for publication at ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06552 2024-11-12 eess.IV 78%

CASC: Condition-Aware Semantic Communication with Latent Diffusion Models

Weixuan Chen, Qianqian Yang

专题命中 可控生成 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01349 2024-11-05 cs.RO cs.LG 78%

The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control

Oleg Kaidanov, Firas Al-Hafez, Yusuf Suvari, Boris Belousov, Jan Peters

专题命中 可控生成 :diffusion(title,abstract)

Comments Conference on Robot Learning, Workshop on Whole-Body Control and Bimanual Manipulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02121 2024-10-04 eess.IV cs.LG cs.NI 78%

SC-CDM: Enhancing Quality of Image Semantic Communication with a Compact Diffusion Model

Kexin Zhang, Lixin Li, Wensheng Lin, Yuna Yan, Wenchi Cheng, Zhu Han

专题命中 可控生成 :diffusion(title,abstract)

Comments arXiv admin note: text overlap with arXiv:2408.05112

详情

展开后加载摘要…

URL PDF HTML 收藏