arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 2537 信号源:cs.CV, cs.GR, cs.MM

1. 效率与蒸馏 2537 篇

2207.10776 2022-07-25 cs.CV 83%

Auto-regressive Image Synthesis with Integrated Quantization

Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Kaiwen Cui, Changgong Zhang, Shijian Lu

专题命中 效率与蒸馏 :image synthesis(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted to ECCV 2022 as Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06999 2025-05-29 cs.LG 83%

Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models

Siddarth Venkatraman, Mohsin Hasan, Minsu Kim, Luca Scimeca, Marcin Sendera, Yoshua Bengio, Glen Berseth, Nikolay Malkin

机构 * Mila -- Qu\'ebec AI Institute Jagiellonian University CIFAR University of Edinburgh

专题命中 效率与蒸馏 :diffusion(title,abstract);image generation(abstract)

Comments ICML 2025; code: https://github.com/HyperPotatoNeo/Outsourced_Diffusion_Sampling

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17546 2026-05-19 astro-ph.IM astro-ph.GA cs.LG 82%

Accelerating Redshift-Conditioned Galaxy Image Synthesis with One-step Generative Modeling

通过一步生成建模加速红移条件下的星系图像合成

Tianyue Yang, Sandro Tacchella, Xiao Xue

机构 * The Center for Computational Science(计算科学中心) University College London(伦敦大学学院) Cavendish Laboratory(卡文迪许实验室) Kavli Institute for Cosmology University of Cambridge(剑桥大学卡文迪许宇宙研究所)

专题命中 效率与蒸馏 :image synthesis(title,abstract);diffusion(abstract)

AI总结 本文研究了利用扩散模型和像素MeanFlow实现高效红移条件下的星系图像生成,通过对比不同模型在GalaxiesML-64数据集上的表现,发现一步生成模型在计算成本大幅降低的情况下能有效恢复星系形态统计信息,为大规模宇宙巡天和基于模拟的科学推断提供了新路径。

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19962 2025-09-25 cs.LG stat.ML 82%

Learnable Sampler Distillation for Discrete Diffusion Models

Feiyang Fu, Tongxian Guo, Zhaoqiang Liu

机构 * University of Electronic Science and Technology of China(电子科学与技术大学)

专题命中 效率与蒸馏 :diffusion(title,abstract);image generation(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13162 2025-06-18 physics.ins-det hep-ex hep-ph 82%

Choose Your Diffusion: Efficient and flexible ways to accelerate the diffusion model in fast high energy physics simulation

Cheng Jiang, Sitian Qian, Huilin Qu

专题命中 效率与蒸馏 :diffusion(title,abstract);image generation(abstract)

Comments Address comments from SciPost

Journal ref SciPost Phys. 18, 195 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08398 2025-04-14 cs.AR cs.LG 82%

MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization

Daeun Kim, Jinwoo Hwang, Changhun Oh, Jongse Park

专题命中 效率与蒸馏 :diffusion(title,abstract);image generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16852 2024-12-09 cs.LG cs.AI stat.ML 82%

EM Distillation for One-step Diffusion Models

Sirui Xie, Zhisheng Xiao, Diederik P Kingma, Tingbo Hou, Ying Nian Wu, Kevin Patrick Murphy, Tim Salimans, Ben Poole, Ruiqi Gao

专题命中 效率与蒸馏 :diffusion(title,abstract);text-to-image(abstract)

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20289 2024-05-31 cs.SD cs.AI cs.LG 82%

DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation

Zachary Novack, Julian McAuley, Taylor Berg-Kirkpatrick, Nicholas Bryan

专题命中 效率与蒸馏 :diffusion(title,abstract);inpainting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01454 2023-09-06 hep-ph stat.ML 82%

Accelerating Markov Chain Monte Carlo sampling with diffusion models

N. T. Hunt-Smith, W. Melnitchouk, F. Ringer, N. Sato, A. W Thomas, M. J. White

专题命中 效率与蒸馏 :diffusion(title,abstract);image synthesis(abstract)

Comments 21 pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00512 2022-06-08 cs.LG cs.AI stat.ML 82%

Progressive Distillation for Fast Sampling of Diffusion Models

Tim Salimans, Jonathan Ho

专题命中 效率与蒸馏 :diffusion(title,abstract);image generation(abstract)

Comments Published as a conference paper at ICLR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18824 2025-09-24 cs.CV 81%

Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation

Yanzuo Lu, Xin Xia, Manlin Zhang, Huafeng Kuang, Jianbin Zheng, Yuxi Ren, Xuefeng Xiao

机构 * ByteDance Seed(字节跳动种子)

专题命中 效率与蒸馏 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16720 2025-04-28 cs.CV 81%

Importance-Based Token Merging for Efficient Image and Video Generation

Haoyu Wu, Jingyi Xu, Hieu Le, Dimitris Samaras

机构 * Stony Brook University(石溪大学) EPFL(苏黎世联邦理工学院)

专题命中 效率与蒸馏 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11158 2025-02-19 cs.CV 81%

AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks

Ming Xie, Chenjie Cao, Yunuo Cai, Xiangyang Xue, Yu-Gang Jiang, Yanwei Fu

专题命中 效率与蒸馏 :text-to-image(abstract);diffusion(abstract);image editing(abstract);inpainting(abstract)

Comments 19 pages, submitted to TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02712 2024-03-22 cs.CV cs.AI cs.LG stat.ML 81%

ED-NeRF: Efficient Text-Guided Editing of 3D Scene with Latent Space NeRF

Jangho Park, Gihyun Kwon, Jong Chul Ye

专题命中 效率与蒸馏 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image editing(abstract)

Comments ICLR 2024; Project Page: https://jhq1234.github.io/ed-nerf.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12691 2026-02-27 cs.GR cs.CV cs.LG 81%

Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration

自适应混合缓存用于高效文本到视频扩散模型加速

Yuanxin Wei, Lansong Diao, Bujiao Chen, Shenggan Cheng, Zhengping Qian, Wenyuan Yu, Nong Xiao, Wei Lin, Jiangsu Du

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团) National University of Singapore(新加坡国立大学)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 本文提出MixCache,一种基于缓存的训练自由框架,通过自适应混合缓存策略提升视频DiT模型的生成效率和质量。

Comments 9 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01204 2025-04-03 cs.GR cs.CV 81%

Articulated Kinematics Distillation from Video Diffusion Models

Xuan Li, Qianli Ma, Tsung-Yi Lin, Yongxin Chen, Chenfanfu Jiang, Ming-Yu Liu, Donglai Xiang

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06897 2025-02-12 cs.GR cs.AI cs.CV 81%

PyPotteryInk: One-Step Diffusion Model for Sketch to Publication-ready Archaeological Drawings

Lorenzo Cardarelli

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15052 2025-01-28 cs.CV cs.AI cs.MM 81%

Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval

Bingjun Luo, Jinpeng Wang, Wang Zewen, Junjie Zhu, Xibin Zhao

专题命中 效率与蒸馏 :text-to-image(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10126 2023-05-18 cs.CV cs.MM 81%

Fusion-S2iGan: An Efficient and Effective Single-Stage Framework for Speech-to-Image Generation

Zhenxing Zhang, Lambert Schomaker

专题命中 效率与蒸馏 :image generation(title);text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10363 2022-10-17 cs.CV cs.GR eess.IV 81%

Towards Device Efficient Conditional Image Generation

Nisarg A. Shah, Gaurav Bharaj

专题命中 效率与蒸馏 :image generation(title,abstract);分类 cs.CV、cs.GR

Comments British Machine Vision Conference 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09180 2026-07-10 cs.LG cs.CV 80%

FSampler: Training Free Acceleration of Diffusion Sampling via Epsilon Extrapolation

FSampler: 通过epsilon外推实现无需训练的扩散采样加速

Michael A. Vladimir

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 FSampler通过epsilon外推减少模型调用和采样时间,提升扩散采样效率。

Comments 10 pages; diffusion models; accelerated sampling; ODE solvers; epsilon extrapolation; training free inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09430 2026-05-13 cs.CV 80%

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

FlashAR: 用于自回归图像生成的高效后训练加速

Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang

机构 * Zhejiang University(浙江大学) University of Adelaide(阿德莱德大学)

专题命中 效率与蒸馏 :image generation(title,abstract);分类 cs.CV

AI总结 本文提出FlashAR框架,通过双向往前预测将预训练自回归模型转化为高效并行生成器,实现512x512图像生成速度提升22.9倍。

Comments Post-training acceleration for autoregressive image generation, code is available at https://lxazjk.github.io/FlashAR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01035 2026-08-17 cs.RO cs.AI cs.CV 版本更新 79%

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

WAM-Diff2:面向高效自动驾驶视觉-语言-动作(VLA)的分层自回归到扩散蒸馏

Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 WAM-Diff2通过三阶段分层蒸馏策略,将预训练自回归VLA模型转化为高效扩散模型,缓解暴露偏差、实现与基线相当性能,解码速度提升2.8倍,结合系统级优化后达15.1倍加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07327 2026-08-10 cs.CV 版本更新 79%

Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations

教师特征漂移:基于预训练扩散表示的一步扩散蒸馏

Yuan Zhang, Chenyi Li, Haodong Yu, Guoqing Ma, Jiajun Zha, Yuanming Yang, Bo Wang, Wei Tang, Wenbo Li, Haoyang Huang, Nan Duan

机构 * JD Explore Academy(京东探索研究院) Peking University(北京大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学) Beijing Institute of Technology(北京理工大学)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出通过单步扩散蒸馏简化模型,利用预训练扩散教师自身的特征空间,无需额外网络即可实现语义特征几何的漂移,同时引入轻量模式覆盖损失以提升生成质量和多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03811 2026-08-03 cs.CV 版本更新 79%

Progressive Checkerboards for Autoregressive Multiscale Image Generation

渐进棋盘用于自回归多尺度图像生成

David Eigen

专题命中 效率与蒸馏 :image generation(title,abstract);分类 cs.CV

AI总结 本研究提出了一种基于渐进棋盘的多尺度自回归图像生成方法,通过并行采样和保持平衡结构,实现高效且有效的条件建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08293 2026-07-30 cs.CV 版本更新 79%

Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion

Distill, Diffuse, and Semanticize (DDS): 基于多粒度蒸馏和图扩散的无标注3D场景理解

Yijing Wang, Ruonan Li, Qilin Wang, Rongqiang Zhao, Jie Liu

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Pengcheng Laboratory(鹏城实验室)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 DDS通过多粒度蒸馏和图扩散实现轻量级无标注3D场景理解,提升区域一致性和语义识别,实验显示在多个数据集上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15355 2026-07-29 cs.CV 版本更新 79%

DAV-GSWT: Diffusion-Active-View Sampling for Data-Efficient Gaussian Splatting Wang Tiles

DAV-GSWT:基于扩散优先的主动视角采样用于数据高效的高斯溅射瓦片

Rong Fu, Jiekai Wu, Yee Tan Jia, Yang Li, Xiaowen Ma, Wangyu Wu, Simon Fong

机构 * University of Macau(澳门大学) Juntendo University(立命馆大学) Tongji University(同济大学) Renmin University of China(中国人民大学) University of Chinese Academy of Sciences(中国科学院大学) Zhejiang University(浙江大学) University of Liverpool(利物浦大学)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出DAV-GSWT框架,利用扩散先验和主动视角采样,从少量输入观测合成高保真高斯溅射瓦片,减少数据需求并保持视觉质量和交互性能。

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21446 2026-07-24 cs.LG cs.CV 新提交 79%

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

KroQuant:用于扩散Transformer高效训练后量化的克罗内克结构块变换

Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann

机构 * EPFL(洛桑联邦理工学院) Advanced Micro Devices, Inc.(超威半导体公司)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 研究针对扩散Transformer训练后量化到W4A4质量严重下降问题,提出KroQuant方法,对激活值3个2元素块应用克罗内克结构可逆变换,存储参数少,运行快,结合离线LoRaQ权重校准,输出更接近FP参考且保持或提高图像质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20628 2026-07-24 cs.CV cs.AI 新提交 79%

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring

RealVDeblur:用于通用真实世界视频去模糊的一步扩散法

Renbiao Jin, Mingxin Yang, Yutian Chen, Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) CUHK MMLab(香港中文大学多媒体实验室) CPII under InnoHK(创新香港研究院下的CPII) Tsinghua University(清华大学)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 针对真实世界视频去模糊难题,提出RealVDeblur框架。构建模糊合成管道提供数据,利用视频扩散先验恢复,采用逐帧编码并简化采样为一步生成器,还有时间窗口掩码稳定推理,在多方面表现出色且提升3D重建鲁棒性。

Comments Project page with code: https://rbjin.github.io/RealVDeblur/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30116 2026-07-23 cs.CV cs.LG 版本更新 79%

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

SGMD: 得分梯度匹配蒸馏用于少步视频扩散蒸馏

Zhuguanyu Wu, Ruihao Gong, Yang Yong, Yushi Huang, Xiangyu Fan, Lei Yang, Dahua Lin, Xianglong Liu

机构 * Beihang University(北京航空航天大学) SenseTime Research(商汤科技研究院) Hong Kong University of Science and Technology(香港科技大学)

专题命中 效率与蒸馏 :diffusion(title,abstract);分类 cs.CV

AI总结 针对分布匹配蒸馏在少步视频扩散中训练昂贵且运动动态保守的问题,提出得分梯度匹配蒸馏(SGMD),通过直接优化假得分朝向教师并使用教师停止梯度Fisher作为稳定目标,实现约3倍训练加速并显著提升运动动态。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏