arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-04-28 至 2026-04-28 共收录 96 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 96 篇

2604.24351 2026-04-28 cs.LG cs.AI cs.CV cs.SE 90%

Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion

扩散模板:一种统一的插件框架,用于可控扩散

Zhongjie Duan, Hong Zhang, Yingda Chen

机构 * Alibaba Group(阿里巴巴集团)

专题命中 扩散模型 :diffusion(title,summary_cn);image editing(abstract);inpainting(abstract);分类 cs.CV

AI总结 本文提出Diffusion Templates框架,通过解耦基础模型推理与可控能力注入,统一了扩散模型的可控能力,支持多种任务和不同backbone的兼容性,提升了模块化和可扩展性。

Comments 21 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23609 2026-04-28 cs.RO 87%

Tube Diffusion Policy: Reactive Visual-Tactile Policy Learning for Contact-rich Manipulation

管扩散策略:面向富接触操作的反应式视觉-触觉策略学习

Teng Xue, Alberto Rigo, Bingjian Huang, Jiayi Shen, Zhengtong Xu, Nick Colonnese, Amirhossein H. Memar

机构 * Meta Reality Labs Research(Meta现实实验室) Idiap Research Institute(Idiap研究机构) École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) University of Toronto(多伦多大学) Purdue University(普渡大学)

专题命中 扩散模型 :diffusion(title,summary_cn)

AI总结 本文提出Tube Diffusion Policy,通过结合扩散模型与管反馈控制,解决富接触操作中对不确定性和高频触觉反馈的快速反应需求,实验验证其在多种任务中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24575 2026-04-28 cs.CV 83%

Diffusion Model as a Generalist Segmentation Learner

扩散模型作为通用分割学习者

Haoxiao Wang, Antao Xiang, Haiyang Sun, Peilin Sun, Changhao Pan, Yifu Chen, Minjie Hong, Weijie Wang, Shuang Chen, Yue Chen, Zhou Zhao

机构 * Zhejiang University(浙江大学) South China University of Technology(华南理工大学) Nanjing University(南京大学) Peking University(北京大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文提出DiGSeg,利用预训练扩散模型实现通用分割框架,通过编码图像和掩码到潜在空间,并结合CLIP对齐的文本路径,实现基于外观和任意文本提示的结构化分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24493 2026-04-28 cs.CV 83%

CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping

CA-IDD:跨注意力引导的身份条件扩散用于身份一致的面部交换

Md Shohel Rana, Tanoy Debnath

机构 * School of Computing(计算学院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出CA-IDD,一种基于扩散的面部交换方法,通过多尺度交叉注意力整合 gaze、身份和面部解析,提升身份一致性和视觉真实性,实现稳定训练和精细区域控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23858 2026-04-28 cs.CV 83%

Latent Inter-Frame Pruning: A Training-Free Method Bridging Traditional Video Compression and Modern Diffusion Transformers for Efficient Generation

潜在帧剪枝:一种无训练方法,连接传统视频压缩与现代扩散变换器以实现高效生成

Dennis Menn, Chih-Hsien Chou

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Futurewei Technologies, Inc.(未来科技公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种无训练的潜在帧剪枝方法,通过剪枝重复的潜在块来减少计算负担并提高生成速度,同时引入注意力恢复机制以解决训练与推理间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11789 2026-04-28 stat.ML cs.CV cs.LG 83%

Statistical Test for Diffusion-Based Anomaly Localization via Selective Inference

基于选择性推断的扩散模型异常定位统计检验

Teruyuki Katsuoka, Tomohiro Shiraishi, Daiki Miwa, Vo Nguyen Le Duy, Ichiro Takeuchi

机构 * Nagoya University(名古屋大学) University of Information Technology(信息技术大学) Vietnam National University, Ho Chi Minh City(越南国家大学,胡志明市) RIKEN(理化学研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出基于选择性推断的统计框架,用于量化检测到的异常区域显著性,通过提供p值评估假阳性率,提升异常定位的可靠性。

Comments 35 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23536 2026-04-28 cs.CV 83%

$Z^2$-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models

$Z^2$-采样:用于扩散模型语义对齐的零成本zigzag轨迹

Haosen Li, Wenshuo Chen, Shaofeng Liang, Lei Wang, Kaishen Yuan, Yutao Yue

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Griffith University & Data61/CSIRO(格里菲斯大学及Data61/CSIRO)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出$Z^2$-采样,通过隐式代数坍缩与动态缓存的时间语义代理,实现零成本zigzag轨迹,提升采样效率并保持语义探索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14897 2026-04-28 cs.CV 83%

Seer: Language Instructed Video Prediction with Latent Diffusion Models

Seer:基于潜在扩散模型的语言指导视频预测

Xianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song, Yang Gao

机构 * Shanghai Qi Zhi Institute(上海启智研究院) IIIS, Tsinghua University(清华大学人工智能研究院) Shanghai AI Lab(上海人工智能实验室) NVIDIA

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 Seer通过扩展预训练文本到图像扩散模型的时间轴,提出高效模型以实现文本条件视频预测,提升机器人未来预测能力,实验表明其在多个数据集上性能优越。

Comments 31 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24487 2026-04-28 cs.RO 82%

Guiding Vector Field Generation via Score-based Diffusion Model

通过基于分数的扩散模型引导向量场生成

Zirui Chen, Shiliang Guo, Shiyu Zhao

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) WINDY Lab, Department of Artificial Intelligence, Westlake University(西湖大学人工智能系WINDY实验室)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出基于分数的引导向量场(SGVF),利用生成模型直接从数据分布构建向量场,解决传统方法在复杂路径中的不足,实验表明其在机器人导航中表现优异。

Comments 8 pages, 6 figrues, ICRA2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13312 2026-04-28 cs.RO cs.AI cs.LG 82%

EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation

EL3DD:扩展的潜在3D扩散用于语言条件的多任务操作

Jonas Bode, Raphael Memmesheimer, Sven Behnke

机构 * Autonomous Intelligent Systems, University of Bonn(博恩大学自主智能系统)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 本文提出EL3DD模型,通过结合视觉和文本输入,利用扩散模型生成精确的机器人轨迹,提升多任务操作的性能和长周期成功率。

Comments 10 pages; 2 figures; 1 table

Journal ref European Robotics Forum 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24719 2026-04-28 cs.CV 79%

DiffuSAM: Diffusion-Based Prompt-Free SAM2 for Few-Shot and Source-Free Medical Image Segmentation

DiffuSAM: 基于扩散的无提示SAM2用于少样本和源无关医学图像分割

Tal Grossman, Noa Cahan, Lev Ayzenberg, Hayit Greenspan

机构 * School of Electrical Engineering, Tel Aviv University, Israel(特拉维夫大学电气工程学院) School of Biomedical Engineering, Tel Aviv University, Israel(特拉维夫大学生物医学工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DiffuSAM通过轻量级扩散先验生成SAM2兼容的分割掩码嵌入,实现无提示医学图像分割,无需用户输入,在BTCV和CHAOS数据集上表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24407 2026-04-28 cs.CV 79%

AD-Relight: Training-Free Banner Relighting via Illumination Translation with Diffusion Priors

AD-Relight:通过扩散先验进行无训练横幅照明转换的免训练横幅重照明

Rameshwar Mishra, A V Subramanyam

机构 * Indraprastha Institute of Information Technology(印度理工学院信息技术研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出AD-Relight,一种无训练的多阶段框架,利用扩散模型在测试时适应重照明新插入的Photoshop横幅,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06166 2026-04-28 cs.CV 79%

B-FIRE: Binning-Free Diffusion Implicit Neural Representation for Hyper-Accelerated Motion-Resolved MRI

B-FIRE:无分组扩散隐式神经表示用于超加速运动分辨MRI

Di Xu, Hengjie Liu, Yang Yang, Mary Feng, Jin Ning, Xin Miao, Jessica E. Scholey, Alexandra E. Hotca-cho, William C. Chen, Michael Ohliger, Martina Descovich, Huiming Dong, Wensha Yang, Ke Sheng

机构 * Radiation Oncology, University of California, San Francisco, California(加州大学旧金山分校放射肿瘤学系) Radiology and Biomedical Imaging, University of California, San Francisco, California(加州大学旧金山分校放射学与生物医学成像系) Siemens Medical Solutions USA, Inc., Cleveland, Ohio(西门子医疗解决方案美国公司,克利夫兰,俄亥俄) Radiology at Children’s Hospital Los Angeles, Keck School of Medicine, University of Southern California, Los Angeles, California(洛杉矶儿童医院放射学,美国南加州大学凯克医学院,洛杉矶,加利福尼亚)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 B-FIRE通过扩散隐式神经表示框架实现超加速MRI重建,解决运动分辨信息模糊问题,提升3D腹部解剖的即时重建精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23709 2026-04-28 cs.CV eess.IV 79%

ZID-Net: Zero-Inference Diffusion Prior Decoupling Network for Single Image Dehazing

ZID-Net:零推理扩散先验解耦网络用于单图像去雾

Xinheng Li, Minghao Chen, Mengqing Wu, Yan Liu, Guanying Huo

机构 * College of Information Science and Engineering, Hohai University(河海大学信息科学与工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ZID-Net通过解耦扩散监督与前馈推理,提出一种高效去雾框架,结合频率-空间解耦前馈骨干网络和物理先验,实现高PSNR和低延迟的去雾效果。

Comments Submitted to Neurocomputing. Includes 12 figures and 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23651 2026-04-28 cs.CV 79%

Geometry-Conditioned Diffusion for Occlusion-Robust In-Bed Pose Estimation

基于几何条件的扩散模型用于抗遮挡的卧床姿态估计

Navid Aslankhani Khameneh, Marco Carletti, Cigdem Beyan

机构 * Department of Computer Science, University of Verona(威尼斯大学计算机科学系) EVS - Embedded Vision Systems Srl(嵌入式视觉系统股份有限公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于几何条件的扩散模型,解决卧床姿态估计中遮挡问题,通过生成模型直接从骨骼关键点生成遮挡图像,提升遮挡鲁棒性。

Comments This is the preprint version of the paper. The final version has been accepted for publication in the Proceedings of the 20th IEEE International Conference on Automatic Face and Gesture Recognition (IEEE FG 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23636 2026-04-28 cs.CV 79%

Discriminator-Guided Adaptive Diffusion for Source-Free Test-Time Adaptation under Image Corruptions

判别器引导的自适应扩散用于无源测试时间适应下的图像损坏

Francesco Olivato, Cigdem Beyan, Vittorio Murino

机构 * Department of Computer Science, University of Verona(威尼斯大学计算机科学系) AI for Good (AIGO), Istituto Italiano di Tecnologia(意大利技术研究院人工智能与善部(AIGO))

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种基于扩散的输入级适应框架,在测试时保持所有源训练模型冻结,通过判别器引导的自适应扩散策略动态控制每个测试样本的扰动量,以抑制特定域的损坏,提升鲁棒性。

Comments This is the preprint (submitted version) of the paper. The final version has been accepted for publication in the Proceedings of the 28th International Conference on Pattern Recognition (ICPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04749 2026-04-28 cs.CV 79%

Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation

通过提示控制的扩散增强缓解长尾偏差

Buddhi Wijenayake, Nichula Wasalathilake, Roshan Godaliyadda, Vijitha Herath, Parakrama Ekanayake, Vishal M. Patel

机构 * University of Peradeniya(珀德尼亚大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于提示控制的扩散增强框架,通过生成带标签的图像样本来增强少数类,提升遥感图像分割中长尾不平衡问题的处理效果。

Comments Accepted to Publication at 2026 IEEE International Geoscience and Remote Sensing Symposium (IGARSS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06904 2026-04-28 cs.CV 79%

BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration

BIR-Adapter:一种参数高效的扩散适配器用于盲图像恢复

Cem Eteke, Alexander Griessel, Wolfgang Kellerer, Eckehard Steinbach

机构 * Chair of Media Technology, Munich Institute of Robotics and Machine Intelligence(媒体技术教授会,慕尼黑机器人与机器智能研究所) School of Computation, Information and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出BIR-Adapter,一种参数高效的扩散适配器,用于盲图像恢复。通过减少训练参数数量和引入采样引导机制,提升恢复可靠性,实验表明其在多个设置中性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22808 2026-04-28 cs.CV cs.AI eess.IV 79%

FreqFormer: Hierarchical Frequency-Domain Attention with Adaptive Spectral Routing for Long-Sequence Video Diffusion Transformers

FreqFormer:具有自适应频谱路由的分层频域注意力机制用于长序列视频扩散变换器

Haopeng Jin

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FreqFormer通过分层频域注意力机制,利用视频特征的频谱结构,采用不同操作符处理不同频段,降低长序列视频扩散变换器的计算与内存开销。

Comments 24 pages, 17 figures, 14 tables, Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21066 2026-04-28 cs.CV cs.LG stat.ME 79%

Optimizing Diffusion Priors in Image Reconstruction from a Single Observation

从单个观测重建图像中优化扩散先验

Frederic Wang, Katherine L. Bouman

机构 * Caltech(加州理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出通过结合现有扩散先验生成单专家先验并优化指数,以提升单观测图像重建的可靠性,验证了在黑洞成像和文本条件去模糊中效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18836 2026-04-28 eess.IV cs.AI cs.CV 79%

Dual-domain Multi-path Self-supervised Diffusion Model for Accelerated MRI Reconstruction

双域多路径自监督扩散模型用于加速MRI重建

Yuxuan Zhang, Jinkui Hao, Bo Zhou

机构 * Department of Radiology, Northwestern University(放射科,西北大学) Department of Biomedical Engineering, Huazhong University of Science and Technology(生物医学工程系,华中科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出双域多路径自监督扩散模型,通过自监督训练方案、轻量混合注意力网络和多路径推理策略,提升MRI重建的准确性、效率和可解释性,克服传统模型依赖全采样数据的局限。

Comments Accepted at IEEE-TNNLS, 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24692 2026-04-28 cs.LG 78%

Diffusion-Guided Feature Selection via Nishimori Temperature: Noise-Based Spectral Embedding

基于尼希莫里温度的扩散引导特征选择:基于噪声的谱嵌入

Vasiliy S. Usatyuk, Denis A. Sapozhnikov, Sergey I. Egorov

机构 * Department of Computer Science(计算机科学系) South-West State University(西南州大学) T8 LLC Moscow, Russia(T8 LLC莫斯科俄罗斯) South-West State University Kursk, Russia(西南州大学库尔斯克俄罗斯)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出基于噪声的谱嵌入方法,通过构建稀疏相似图并识别尼希莫里温度,实现高维数据中信息特征的选择,实验表明其在压缩下保持分类精度优于传统方法。

Comments 8 pages, 3 figures, extended version (with noise shift proof) of DSPA2026 article

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24683 2026-04-28 cond-mat.mtrl-sci physics.chem-ph 78%

Improved Electrochemical Performance and Diffusion kinetics by Boron-doping in Na$_{0.66}$Mn$_{0.8}$Fe$_{0.2}$O$_{2}$ Layered Cathodes for Sodium-Ion Batteries

通过硼掺杂提升Na₀.₆₆Mn₀.₈Fe₀.₂O₂层状正极材料的电化学性能与扩散动力学性能用于钠离子电池

Jayashree Pati, P. Senthilkumar, Deepak Seth, Riya Gulati, Manish Kr. Singh, Madhav Sharma, Anita Dhaka, M. Ali Haider, Rajendra S. Dhaka

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 研究通过硼掺杂改善钠离子电池正极材料的比容量和循环稳定性,利用电化学测试和理论计算分析其扩散动力学及结构稳定性。

Comments submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24640 2026-04-28 quant-ph 78%

DiffQEC: A versatile diffusion model for quantum error correction

DiffQEC:一种用于量子纠错的通用扩散模型

Tianyi Xu, Qinglong Liu, Maolin Wang, Fei Zhang, Zhe Zhao, Yang Wang, Ye Wei

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出DiffQEC,一种基于离散去噪扩散的生成解码器,通过结合综合征处理器和特征调制,提升量子纠错的解码效率与准确性。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24416 2026-04-28 cs.CL cs.AI cs.LG 78%

Scaling Properties of Continuous Diffusion Spoken Language Models

连续扩散口语语言模型的可扩展性特性

Jason Ramapuram, Eeshan Gunesh Dhekane, Amitis Shidani, Dan Busbridge, Bogdan Mazoure, Zijin Gu, Russ Webb, Tatiana Likhomanenko, Navdeep Jaitly

机构 * Apple(苹果公司)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究连续扩散口语语言模型在性能上的可扩展性,提出基于音素Jensen-Shannon散度的评估指标,发现其在参数规模增加时生成质量提升,但长文本连贯性仍面临挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01338 2026-04-28 cs.LG math.ST stat.ML stat.TH 78%

High-accuracy sampling for diffusion models and log-concave distributions

扩散模型与对数凹分布的高精度采样

Fan Chen, Sinho Chewi, Constantinos Daskalakis, Alexander Rakhlin

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出扩散模型采样算法,实现δ误差在polylog(1/δ)步内完成,改进了现有结果。针对数据内在维度d*和非均匀L-Lipschitz条件,提出复杂度为polylog(1/δ)的采样方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19110 2026-04-28 cond-mat.stat-mech 78%

Fate of diffusion under integrability breaking of classical integrable magnets

经典可积磁体在非可积性破坏下的扩散命运

Jiaozi Wang, Sourav Nandy, Markus Kraft, Tomaž Prosen, Robin Steinigeweg

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 研究经典可积磁体在非可积性破坏下的扩散行为,发现扩散常数随扰动强度变化及磁化转移统计分布的非高斯到高斯转变。

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17381 2026-04-28 cs.LG 78%

Beyond Binary Out-of-Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories

超越二元分布外检测:利用多统计扩散轨迹表征分布偏移

Achref Jaziri, Martin Rogmann, Martin Mundt, Visvanathan Ramesh

机构 * Goethe University(弗赖堡大学) University of Bremen(不莱梅大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出DISC方法,通过扩散模型的迭代去噪过程提取多维特征向量,表征多噪声水平下的统计差异,实现对分布外数据类型的分类,突破传统二元检测的局限。

Comments Accepted at AISTATS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06027 2026-04-28 cs.SD cs.AI eess.AS 78%

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

DreamAudio: 基于扩散模型的定制化文本到音频生成

Yi Yuan, Xubo Liu, Haohe Liu, Xiyuan Kang, Zhuo Chen, Yuxuan Wang, Mark D. Plumbley, Wenwu Wang

机构 * School of Computer Science and Electronic Engineering, University of Surrey(Surrey大学计算机科学与电子工程学院) Department of Informatics, King’s College London(伦敦国王学院信息学院) Seed Group, ByteDance Inc.(字节跳动Seed团队)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出DreamAudio,通过参考音频样本生成定制化音频,提升细粒度音频控制能力,实验显示其在定制生成任务中表现优异。

Comments Lastest arxiv version. Accepted by IEEE/ACM Transactions on Audio, Speech, and Language Processing. Demos are available at https://yyua8222.github.io/DreamAudio_demopage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24238 2026-04-28 cs.LG 78%

GeoEdit: Local Frames for Fast, Training-Free On-Manifold Editing in Diffusion Models

GeoEdit:用于扩散模型中无训练快速在流形上编辑的局部框架

Yiming Zhang, Sitong Liu, Ke Li, Zhihong Wu, Alex Cloninger, Melvin Leok

机构 * University of California San Diego, La Jolla, CA, USA(加州大学圣迭戈分校) University of Washington, Seattle, WA, USA(华盛顿大学) Xidian University, Xi'an, China(西电大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出GeoEdit,通过局部流形切线空间实现无需训练的快速在流形上编辑,利用小扰动构建切线框架,实现精细编辑。

详情

展开后加载摘要…

URL PDF HTML 收藏