arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-03-10 至 2026-03-10 共收录 105 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 105 篇

2503.00897 2026-03-10 cs.LG cs.AI cs.CV 86%

A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning

一种简单而有效的文本到图像扩散微调强化学习方法

Shashank Gupta, Chaitanya Ahuja, Tsung-Yu Lin, Sreya Dutta Roy, Harrie Oosterhuis, Maarten de Rijke, Satya Narayan Shukla

机构 * University of Amsterdam(阿姆斯特丹大学) Meta Radboud University(拉德堡德大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(title);分类 cs.CV

AI总结 本文提出LOOP方法,结合REINFORCE的方差减少技术和PPO的鲁棒性,提升扩散模型在黑盒目标上的样本效率和性能。

Comments Published at Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08271 2026-03-10 cs.CV 83%

Prototype-Guided Concept Erasure in Diffusion Models

扩散模型中的原型引导概念擦除

Yuze Cai, Jiahao Lu, Hongxiang Shi, Yichao Zhou, Hong Lu

机构 * Fudan University(复旦大学) National University of Singapore(国立新加坡大学)

专题命中 扩散模型 :diffusion(title);image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出通过原型引导的方法在扩散模型中实现更可靠的概念擦除,尤其针对广义概念如性与暴力,提升图像生成的安全性和可控性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24758 2026-03-10 cs.CV 83%

ExGS: Extreme 3D Gaussian Compression with Diffusion Priors

ExGS:基于扩散先验的极端3D高斯压缩

Jiaqi Chen, Xinhao Ji, Yuanyuan Gao, Hao Li, Yuning Gong, Yifei Liu, Dan Xu, Zhihang Zhong, Dingwen Zhang, Xiao Sun

机构 * Northwestern Polytechnical University(北western工业大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 ExGS通过结合UGC和GaussPainter,利用扩散先验实现高效且高质量的极端3DGS压缩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07815 2026-03-10 cs.CV cs.AI 83%

HybridStitch: Pixel and Timestep Level Model Stitching for Diffusion Acceleration

HybridStitch: 像素和时间步级模型拼接用于扩散加速

Desen Sun, Jason Hon, Jintao Zhang, Sihang Liu

机构 * University of Waterloo(多伦多大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 HybridStitch通过结合大模型和小模型,实现T2I生成的高效加速,提升Stable Diffusion 3的生成速度1.83倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07700 2026-03-10 cs.CV cs.AI 83%

TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward

TDM-R1: 通过非可微奖励强化少步扩散模型

Yihong Luo, Tianyang Hu, Weijian Luo, Jing Tang

机构 * Hong Kong University of Science and Technology(香港科技大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) hi-Lab, Xiaohongshu Inc(小红书实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 TDM-R1通过非可微奖励强化少步扩散模型,提升文本到图像生成性能

Comments https://luo-yihong.github.io/TDM-R1-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15128 2026-03-10 eess.IV cs.CV 83%

MAP-based Problem-Agnostic diffusion model for Inverse Problems

基于最大后验概率的逆问题问题无关扩散模型

Pingping Tao, Haixia Liu, Jing Su

机构 * organization= Library, Shandong University , addressline= No. 180 West Culture Road , city= Weihai , postcode= 264209 , state= Shandong , country= China organization= School of Mathematics Statistics \& Hubei Key Laboratory of Engineering Modeling Scientific Computing \& Institute of Interdisciplinary Research for Mathematics Applied Science, Huazhong University of Science organization= Department of Basic Sciences, Dalian University of Science

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 本文提出了一种基于最大后验概率的逆问题扩散模型,通过引入高斯型先验和引导项估计,提升图像处理任务中的内容保留效果。

Comments 26 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15199 2026-03-10 cs.CV cs.AI cs.LG 83%

Input-Adaptive Generative Dynamics in Diffusion Models

扩散模型中的输入自适应生成动态

Yucheng Xing, Xiaodong Liu, Xin Wang

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本研究提出了一种输入自适应的扩散模型生成方法,通过调整生成动态以适应不同输入需求,从而提升生成质量和效率。

Comments 14 pages, 6 figures. Updated version with revised title

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08023 2026-03-10 cs.CV cs.AI cs.GR cs.SD 81%

Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model

不同于Transformer:用基于Mamba的扩散模型生成舞蹈,摒弃节奏表示

Sangjune Park, Inhyeok Choi, Donghyeon Soon, Youngwoo Jeon, Kyungdon Joo

机构 * Ulsan National Institute of Science and Technology, South Korea(乌山国立科学技术研究院,韩国) Daegu Gyeongbuk Institute of Science and Technology, South Korea(大邱庆北科学技术研究院,韩国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 本文提出基于Mamba的扩散模型,用于生成舞蹈,通过替代Transformer并引入基于高斯的节奏表示,有效生成合理舞蹈动作,保持从短到长舞蹈的一致性。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08709 2026-03-10 cs.CV cs.AI 80%

Scale Space Diffusion

尺度空间扩散

Soumik Mukhopadhyay, Prateksha Udhayanan, Abhinav Shrivastava

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出尺度空间扩散模型,通过融合尺度空间理论与扩散过程,改进图像生成与去噪效果。

Comments Project website: https://prateksha.github.io/projects/scale-space-diffusion/ . The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08364 2026-03-10 cs.CV 79%

Diffusion-Based Data Augmentation for Image Recognition: A Systematic Analysis and Evaluation

基于扩散的数据增强用于图像识别:系统分析与评估

Zekun Li, Yinghuan Shi, Yang Gao, Dong Xu

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出UniDiffDA框架,系统分析DiffDA方法的三个核心组件,通过全面评估揭示不同策略的优劣,提供可复现的代码和配置以促进研究。

Journal ref Int J Comput Vis 134, 126 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02919 2026-03-10 cs.CV cs.AI cs.LG 79%

Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers

可解释的运动注意力图:在视频扩散变换器中进行时空定位概念

Youngjun Jun, Seil Kang, Woojung Han, Seong Jae Hwang

机构 * Yonsei University(延世大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种可解释的运动注意力图方法,通过时空定位运动概念,提升视频扩散变换器的可解释性与定位能力。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14980 2026-03-10 cs.RO cs.AI cs.CV 79%

M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation

M4Diffuser:多视角扩散策略与具有操纵性意识的控制的多自由度扩散策略用于鲁棒移动操纵

Ju Dong, Lei Zhang, Liding Zhang, Yao Ling, Yu Fu, Kaixin Bai, Zoltán-Csaba Márton, Zhenshan Bing, Zhaopeng Chen, Alois Christian Knoll, Jianwei Zhang

机构 * TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg(汉堡大学信息学院技术多模态系统部门) Technical University of Munich(慕尼黑技术大学) Agile Robots SE(敏捷机器人有限公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 M4Diffuser通过多视角扩散策略与操纵性意识控制器实现鲁棒移动操纵,提升任务成功率与环境适应性。

Comments Project page: https://sites.google.com/view/m4diffuser, 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08135 2026-03-10 cs.CV 79%

VesselFusion: Diffusion Models for Vessel Centerline Extraction from 3D CT Images

VesselFusion:基于扩散模型的3D CT图像血管中心线提取

Soichi Mita, Shumpei Takezaki, Ryoma Bise

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 VesselFusion通过扩散模型实现3D CT图像中血管中心线的提取,采用粗到细表示和投票聚合方法,提升提取精度和结果自然度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08020 2026-03-10 cs.CV 79%

VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion

VSDiffusion:通过可见性约束扩散镇定模糊的阴影生成

Jing Li, Jing Zhang

机构 * East China University of Science and Technology(东华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 VSDiffusion通过可见性约束扩散框架,结合多尺度结构指导和软先验图,有效生成逼真阴影,提升复杂场景下的几何一致性。

Comments 12 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07889 2026-03-10 cs.CV 79%

Structure and Progress Aware Diffusion for Medical Image Segmentation

结构与进展感知扩散用于医学图像分割

Siyuan Song, Guyue Hu, Chenglong Li, Dengdi Sun, Zhe Jin, Jin Tang

机构 * School of Artificial Intelligence(人工智能学院) School of Computer Science and Technology(计算机科学与技术学院) State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) Anhui Provincial Key Laboratory of Security Artificial Intelligence(安徽省安全人工智能重点实验室) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出SPAD方法,结合语义集中扩散和边界集中扩散,通过进展感知调度器实现医学图像分割的粗到细的扩散过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07865 2026-03-10 cs.SD cs.CV eess.AS 79%

SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving

SoundWeaver: 语义预热启动用于文本到音频扩散服务

Ayush Barik, Sofia Stoica, Nikhil Sarda, Arnav Kethana, Abhinav Khanduja, Muchen Xu, Fan Lai

机构 * University of Illinois Urbana-Champaign, USA(伊利诺伊大学厄巴纳-香槟分校) Assured Intelligence, USA(Assured Intelligence)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SoundWeaver通过语义预热启动技术,提升文本到音频扩散服务的效率,减少延迟并保持音频质量。

Comments Submitted to INTERSPEECH 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07759 2026-03-10 cs.CV cs.AI 79%

DECADE: A Temporally-Consistent Unsupervised Diffusion Model for Enhanced Rb-82 Dynamic Cardiac PET Image Denoising

DECADE:一种用于增强Rb-82动态心脏PET图像去噪的时序一致无监督扩散模型

Yinchi Zhou, Liang Guo, Huidong Xie, Yuexi Du, Ashley Wang, Menghua Xia, Tian Yu, Ramesh Fazzone-Chettiar, Christopher Weyman, Bruce Spottiswoode, Vladimir Panin, Kuangyu Shi, Edward J. Miller, Attila Feher, Albert J. Sinusas, Nicha C. Dvornek, Chi Liu

机构 * Yale University(耶鲁大学) Yale School of Medicine(耶鲁医学院) Siemens Medical Solutions USA, Inc.(西门子医疗解决方案美国公司) Inselspital, Bern University Hospital(伯尔尼大学医院Inselspital) University of Bern(伯尔尼大学) Department of Biomedical Engineering(生物医学工程系) Department of Radiology and Biomedical Imaging(放射科和生物医学成像系) Department of Medicine (Cardiology)(医学系(心脏病学))

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DECADE是一种用于增强Rb-82动态心脏PET图像去噪的无监督扩散模型,通过引入时间一致性提升去噪效果,实现高质量动态和参数图像重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07697 2026-03-10 cs.CV 79%

Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation

学习上下文自适应的运动先验以实现遮蔽运动扩散模型的高效运动学注意力聚合

Junkun Jiang, Jie Chen, Ho Yin Au, Jingyu Xiang

机构 * Department of Computer Science, Hong Kong Baptist University(计算机科学系,香港 Baptist 大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MMDM通过高效运动学注意力聚合机制,学习上下文自适应的运动先验,实现遮蔽运动数据的高效重建与生成。

Comments Accepted by IEEE Transactions on Multimedia. Supplementary material is included

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07430 2026-03-10 cs.CV 79%

Disentangled Textual Priors for Diffusion-based Image Super-Resolution

解耦文本先验用于基于扩散的图像超分辨率

Lei Jiang, Xin Liu, Xinze Tong, Zhiliang Li, Jie Liu, Jie Tang, Gangshan Wu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210023, China(新型软件技术国家重点实验室,南京大学,南京)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DTPSR通过解耦文本先验提升基于扩散的图像超分辨率性能,引入空间层次和频率语义两个维度,实现高感知质量和强泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07361 2026-03-10 cs.LG cs.CV 79%

N-Tree Diffusion for Long-Horizon Wildfire Risk Forecasting

N-Tree Diffusion用于长周期野火风险预测

Yucheng Xing, Xin Wang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 N-Tree Diffusion通过共享早期去噪阶段和后期分支,提高了长周期野火风险预测的效率和准确性。

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08131 2026-03-10 cs.CV 79%

Real-Time Motion-Controllable Autoregressive Video Diffusion

实时可控的自回归视频扩散

Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) Xmax.AI Ltd(Xmax.AI有限公司) Zhejiang University(浙江大学) Singapore Management University(新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 AR-Drag是一种基于强化学习的实时图像到视频生成模型,通过自回归机制和轨迹奖励模型实现高效运动控制,提升视频生成质量和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14357 2026-03-10 cs.CV cs.LG 79%

Vid2World: Crafting Video Diffusion Models to Interactive World Models

Vid2World: 构建视频扩散模型以交互式世界模型

Siqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao, Mingsheng Long

机构 * Tsinghua University(清华大学) Chongqing University(重庆大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Vid2World通过因果化视频扩散模型,提升其在交互式世界建模中的可控性和生成能力,适用于机器人操作、游戏模拟和开放世界导航等多种领域。

Comments Project page: http://knightnemo.github.io/vid2world/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06432 2026-03-10 cs.CV cs.AI 79%

Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single-Image Denoising

Prompt-SID: 通过潜在扩散学习结构表示提示进行单图像去噪

Huaqiu Li, Wang Zhang, Xiaowan Hu, Tao Jiang, Zikang Chen, Haoqian Wang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 Prompt-SID通过潜在扩散学习结构表示提示,提升单图像去噪效果,有效保留结构细节。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07169 2026-03-10 cs.CV 79%

RDM: Recurrent Diffusion Model for Human Motion Generation

RDM:用于人类动作生成的递归扩散模型

Mirgahney Mohamed, Harry Jake Cunningham, Marc P. Deisenroth, Lourdes Agapito

机构 * Department of Computer Science, University College London(计算机科学系,伦敦大学学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 RDM通过递归扩散模型生成长序列人类动作,利用归一化流建模循环连接,提升效率并保持概率性质。

Comments v2: Major revision with extensive text polishing and structural updates. Added new experiments on the rollout effect, specifically analyzing the trade-offs between compute time and sequence length. Includes several new visualizations (Figures 6, 9, 10) and an expanded discussion in Section 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06696 2026-03-10 cs.CV 79%

HARP: HARmonizing in-vivo diffusion MRI using Phantom-only training

HARP: 仅使用假体数据进行体内扩散磁共振成像的协调

Hwihun Jeong, Qiang Liu, Kathryn E. Keenan, Elisabeth A. Wilde, Walter Schneider, Sudhir Pathak, Anthony Zuccolotto, Lauren J. O'Donnell, Lipeng Ning, Yogesh Rathi

机构 * Department of Psychiatry(精神医学系) Brigham and Women's Hospital(布里奇沃特医院) Harvard Medical School(哈佛医学院) College of Engineering(工程学院) Northeastern University(东北大学) National Institute of Standards and Technology(国家标准技术研究院) University of Utah School of Medicine(犹他大学医学院) George E. Wahlen Veterans Affairs Medical Center(乔治·E·瓦伦的退伍军人事务医疗中心) University of Pittsburgh(匹兹堡大学) Department of Radiology(放射医学系) Harvard-MIT Health Sciences and Technology(哈佛-麻省理工健康科学与技术)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 HARP通过仅使用假体数据训练深度学习模型,实现了无需多站点活体数据的扩散磁共振成像协调,有效降低了扫描仪间变异性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06640 2026-03-10 cs.CV cs.LG 79%

Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion Models

剪枝之下:揭示基于剪枝的去学习中概念复兴的风险

Ci Zhang, Zhaojun Ding, Chence Yang, Jun Liu, Xiaoming Zhai, Shaoyi Huang, Beiwen Li, Xiaolong Ma, Jin Lu, Geng Yuan

机构 * University of Georgia(佐治亚大学) Carnegie Mellon University(卡内基梅隆大学) Northeastern University(东北大学) Stevens Institute of Technology(史蒂文斯理工学院) University of Arizona(亚利桑那大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文揭示基于剪枝的去学习中概念复兴的风险,提出攻击框架可无数据恢复被擦除概念,并探讨安全剪枝机制。

Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12719 2026-03-10 cs.CV 79%

S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation

S2DiT:用于移动流媒体视频生成的 Sandwich Diffusion Transformer

Lin Zhao, Yushu Wu, Aleksei Lebedev, Dishani Lahiri, Meng Dong, Arpit Sahni, Michael Vasilkovsky, Hao Chen, Ju Hu, Aliaksandr Siarohin, Sergey Tulyakov, Yanzhi Wang, Anil Kag, Yanyu Li

机构 * Snap Inc. Northeastern University(东北大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 S2DiT 通过高效注意力机制和 2-in-1 深度学习框架,在移动设备上实现高质量、高速度的流式视频生成。

Comments https://snap-research.github.io/S2DiT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00790 2026-03-10 cs.CV cs.AI 79%

LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling

LD-RPS:通过潜势扩散递归后验采样实现零样本统一图像修复

Huaqiu Li, Yong Wang, Tongwen Huang, Hailang Huang, Haoqian Wang, Xiangxiang Chu

机构 * Tsinghua University(清华大学) AMAP, Alibaba Group(阿里云(AMAP))

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 LD-RPS通过潜势扩散递归后验采样实现零样本统一图像修复,结合多模态模型和轻量模块提升修复效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17788 2026-03-10 cs.CV cs.AI 79%

From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction

从二维对齐到三维合理性:统一异构二维先验和无穿透扩散以实现抗遮挡的双手重建

Gaoge Han, Yongkang Cheng, Zhe Chen, Shaoli Huang, Tongliang Liu

机构 * AgiBot Mohamed bin Zayed University of Artificial Intelligence The University of Sydney La Trobe University

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出统一异构二维先验和无穿透扩散模型,以实现抗遮挡的双手重建,提升交互对齐和穿透抑制性能。

Comments Accepted by CVPR 2026 Main, Project: https://gaogehan.github.io/A2P/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09396 2026-03-10 cs.SD cs.CV eess.AS 79%

ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance

ExpGest:利用扩散模型和混合音频-文本引导的表达性说话生成

Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Gaoge Han, Jifeng Ning, Wei Liu

机构 * Northwest A&F University(西北农林科技大学) University of Technology Sydney(悉尼大学) Tencent AILab(腾讯AI实验室) City University of Hong Kong(香港城市大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ExpGest通过结合文本和音频信息,利用扩散模型生成更自然、可控的全身表达性手势。

Comments Accepted by ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏