arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 2537 信号源:cs.CV, cs.GR, cs.MM

1. 效率与蒸馏 2537 篇

2607.04923 2026-07-07 cs.CV 新提交 74%

UniSpine-GS: An Efficient Physics-Aware Gaussian Framework for Cross-Modality Multi-view Spine Image Synthesis

UniSpine-GS:一种用于跨模态多视图脊柱图像合成的高效物理感知高斯框架

Qiuhua Chen, Changning Yu, Na Huang, Chao Sun, Bo Du

机构 * School of Computer Science, Wuhan University(武汉大学计算机科学学院) The Department of Ultrasound of Renmin Hospital East Branch of Wuhan University(武汉大学人民医院东院区超声科) Institute of Artificial Intelligence, Wuhan University(武汉大学人工智能研究院) National Engineering Research Center for Multimedia Software, Wuhan University(武汉大学多媒体软件国家工程研究中心) Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(武汉大学多媒体网络通信工程湖北省重点实验室)

专题命中 效率与蒸馏 :image synthesis(title);分类 cs.CV

AI总结 为解决脊柱疾病诊断中3D成像成本高、模态差异挑战等问题,提出UniSpine-GS框架,通过3D感知表示用于多视图脊柱成像新视图投影渲染,引入策略提升边界保真度和局部细节,性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19703 2026-07-03 cs.CV eess.IV 74%

High-Quality Spatial Reconstruction and Orthoimage Generation Using Efficient 2D Gaussian Splatting

利用高效2D高斯散射实现高质量空间重建和正射影像生成

Qian Wang, Zhihao Zhan, Jialei He, Zhituo Tu, Jie Yuan

机构 * TopXGun Robotics(TopXGun机器人)

专题命中 效率与蒸馏 :image generation(title);分类 cs.CV

AI总结 本文提出基于2D高斯散射的高效方法,实现高质量空间重建和正射影像生成,无需传统DSM和遮挡检测,提升复杂地形和细长结构的渲染质量与效率。

Journal ref Signal, Image and Video Processing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27678 2026-06-29 cs.CV 新提交 74%

Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation

两阶段跨域宫颈异常筛查:细胞病理图像合成与知识蒸馏

Jincheng Li, Yuzhi He, Yihui Zhan, Xinmei Zhang, Yifei Sun, Zelin Liu, Lichi Zhang, Minye Shao, Lili Zhao

机构 * School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) School of Telecommunications Engineering, Xidian University(西安电子科技大学通信工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院) Department of Computer Science, Durham University(杜伦大学计算机科学系)

专题命中 效率与蒸馏 :image synthesis(title);分类 cs.CV

AI总结 提出两阶段框架,先用SC-UNSB合成中间域缓解域偏移,再用知识蒸馏对齐特征,提升跨域宫颈细胞检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15178 2026-05-15 cs.CV 74%

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

SANA-WM:高效分钟级世界建模的混合线性扩散变换器

Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie

机构 * NVIDIA(英伟达)

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

AI总结 SANA-WM通过混合线性注意力、双分支相机控制等设计,实现高效分钟级视频生成,提升效率与视觉质量。

Comments https://nvlabs.github.io/Sana/WM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09227 2026-04-13 eess.IV cs.CV 74%

Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models

无需训练的感知一致低分辨率预览与高分辨率图像的高效扩散模型工作流程

Wongi Jeong, Hoigi Seo, Se Young Chun

机构 * Seoul National University(首尔国立大学)

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

AI总结 本文提出无需训练的低分辨率预览生成方法,通过选择下采样矩阵和交换子零条件,实现与高分辨率图像的感知一致性,减少计算量并提升工作流程效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17652 2026-03-19 cs.RO cs.CV 74%

VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs

VectorWorld: 通过向量图上的扩散流实现高效的流式世界模型

Chaokang Jiang, Desen Zhou, Jiuming Liu, Kevin Li Sun

机构 * University of Cambridge, Cambridge, United Kingdom(剑桥大学)

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

AI总结 VectorWorld通过向量图上的扩散流实现高效的流式世界模型,解决了自动驾驶政策闭环评估中的初始化不匹配、采样延迟和运动可行性问题,提升了地图结构精度和闭环运行稳定性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03208 2026-02-04 cs.LG cs.CV 74%

Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation

谱演化搜索:用于奖励对齐图像生成的高效推理时缩放

Jinyan Ye, Zhongjie Duan, Zhiwen Li, Cen Chen, Daoyuan Chen, Yaliang Li, Yingda Chen

机构 * School of Data Science and Engineering, East China Normal University, Shanghai, China(数据科学与工程学院,东华大学,上海,中国) Alibaba Group, Hangzhou, China(阿里巴巴集团,杭州,中国)

专题命中 效率与蒸馏 :image generation(title);分类 cs.CV

AI总结 本文提出谱演化搜索(SES),通过在低频子空间内执行梯度自由演化搜索,提升奖励对齐图像生成的效率和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00107 2026-02-03 cs.CV cs.RO eess.IV 74%

Efficient UAV trajectory prediction: A multi-modal deep diffusion framework

高效无人机轨迹预测:一种多模态深度扩散框架

Yuan Gao, Xinyu Guo, Wenjing Xie, Zifan Wang, Hongwen Yu, Gongyang Li, Shugong Xu

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

AI总结 本文提出一种多模态深度融合框架,通过融合激光雷达和毫米波雷达数据提升无人机轨迹预测精度,实验显示其比基线模型提升40%。

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25161 2025-09-30 cs.CV 74%

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time

Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, Shijian Lu

机构 * Nanyang Technological University(南洋理工大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

Comments Project page: https://kunhao-liu.github.io/Rolling_Forcing_Webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14177 2025-09-01 stat.ML cs.CV cs.LG 74%

From stability of Langevin diffusion to convergence of proximal MCMC for non-log-concave sampling

Marien Renaud, Valentin De Bortoli, Arthur Leclaire, Nicolas Papadakis

机构 * Univ. Bordeaux, CNRS, Bordeaux INP, IMB, UMR 5251(波尔多大学、法国国家科学研究中心、波尔多INP、IMB、UMR 5251) ENS, CNRS, PSL University Paris(巴黎高等师范学院、法国国家科学研究中心、巴黎PSL大学) LTCI, Télécom Paris, IP Paris, France(LTCI、巴黎电信学院、IP巴黎、法国) Univ. Bordeaux, CNRS, INRIA, Bordeaux INP, IMB, UMR 5251(波尔多大学、法国国家科学研究中心、INRIA、波尔多INP、IMB、UMR 5251)

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03254 2025-08-06 cs.CV cs.AI 74%

V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models

Jisoo Kim, Wooseok Seo, Junwan Kim, Seungho Park, Sooyeon Park, Youngjae Yu

机构 * Yonsei University(延世大学)

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

Comments ICCV2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19951 2025-07-23 cs.CV cs.CL cs.LG 74%

Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation

Shukang Yin, Chaoyou Fu, Sirui Zhao, Chunjiang Ge, Yan Yang, Yuhan Dai, Yongdong Luo, Tong Xu, Caifeng Shan, Enhong Chen

机构 * USTC(中国科学技术大学) NJU(南京大学) THU(清华大学) XMU(厦门大学)

专题命中 效率与蒸馏 :text-to-image(title);分类 cs.CV

Comments Project page: https://github.com/VITA-MLLM/Sparrow

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15822 2025-05-23 eess.IV cs.CV cs.LG 74%

MambaStyle: Efficient StyleGAN Inversion for Real Image Editing with State-Space Models

Jhon Lopez, Carlos Hinojosa, Henry Arguello, Bernard Ghanem

机构 * Universidad Industrial de Santander(圣安德烈大学) KAUST(科威特科学与技术研究局)

专题命中 效率与蒸馏 :image editing(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19369 2025-03-27 cs.CV 74%

EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models

Yufei Cai, Hu Han, Yuxiang Wei, Shiguang Shan, Xilin Chen

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04668 2024-12-09 cs.CV cs.AI 74%

Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation

Ali Abbasi, Shima Imani, Chenyang An, Gayathri Mahalingam, Harsh Shrivastava, Maurice Diesendruck, Hamed Pirsiavash, Pramod Sharma, Soheil Kolouri

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03512 2024-12-05 cs.CV 74%

Distillation of Diffusion Features for Semantic Correspondence

Frank Fundel, Johannes Schusterbauer, Vincent Tao Hu, Björn Ommer

专题命中 效率与蒸馏 :diffusion(title);分类 cs.CV

Comments WACV 2025, Page: https://compvis.github.io/distilldift

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11718 2023-05-22 cs.CV 74%

Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector Quantization

Mengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong Zhang

专题命中 效率与蒸馏 :image generation(title);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14635 2022-11-02 cs.CV cs.AI cs.LG 74%

Compressed Gastric Image Generation Based on Soft-Label Dataset Distillation for Medical Data Sharing

Guang Li, Ren Togo, Takahiro Ogawa, Miki Haseyama

专题命中 效率与蒸馏 :image generation(title);分类 cs.CV

Comments Published as a journal paper at Elsevier CMPB

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.02596 2020-08-07 cs.CV cs.RO 74%

Image Generation for Efficient Neural Network Training in Autonomous Drone Racing

Theo Morales, Andriy Sarabakha, Erdal Kayacan

专题命中 效率与蒸馏 :image generation(title);分类 cs.CV

Comments 2020 International Joint Conference on Neural Networks (IJCNN 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.03611 2019-09-10 eess.IV cs.CV 74%

An Acceleration Framework for High Resolution Image Synthesis

Jinlin Liu, Yuan Yao, Jianqiang Ren

专题命中 效率与蒸馏 :image synthesis(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.01121 2018-08-14 cs.CV 74%

Diverse Conditional Image Generation by Stochastic Regression with Latent Drop-Out Codes

Yang He, Bernt Schiele, Mario Fritz

专题命中 效率与蒸馏 :image generation(title);分类 cs.CV

Comments This version withdrawn by arXiv administrators because the submitter did not have the right to agree to our license at the time of submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.01058 2018-08-06 q-bio.NC cs.CV 74%

Cortical Microcircuits from a Generative Vision Model

Dileep George, Alexander Lavin, J. Swaroop Guntupalli, David Mely, Nick Hay, Miguel Lazaro-Gredilla

专题命中 效率与蒸馏 :generative vision(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09433 2025-01-17 cs.CV cs.GR 73%

CaPa: Carve-n-Paint Synthesis for Efficient 4K Textured Mesh Generation

Hwan Heo, Jangyeong Kim, Seongyeong Lee, Jeong A Wi, Junyoung Choi, Sangjun Ahn

专题命中 效率与蒸馏 :diffusion(abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments project page: https://ncsoft.github.io/CaPa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19310 2024-10-28 cs.CV cs.AI cs.LG cs.MM 73%

Flow Generator Matching

Zemin Huang, Zhengyang Geng, Weijian Luo, Guo-jun Qi

专题命中 效率与蒸馏 :text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13073 2023-12-21 cs.CV cs.LG cs.MM 73%

FusionFrames: Efficient Architectural Aspects for Text-to-Video Generation Pipeline

Vladimir Arkhipkin, Zein Shaheen, Viacheslav Vasilev, Elizaveta Dakhova, Andrey Kuznetsov, Denis Dimitrov

专题命中 效率与蒸馏 :text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.MM

Comments Project page: https://ai-forever.github.io/kandinsky-video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02537 2026-04-10 cs.CV cs.AI 72%

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

RectifiedHR: 通过能量 rectification 实现高效高分辨率合成

Zhen Yang, Guibao Shen, Minyang Li, Liang Hou, Mushui Liu, Luozhou Wang, Xin Tao, Ying-Cong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) Kuaishou Technology(快手科技) HKUST(香港科技大学) Zhejiang University(浙江大学)

专题命中 效率与蒸馏 :diffusion(abstract,comments);image editing(abstract);分类 cs.CV

AI总结 本文提出RectifiedHR,一种无需训练的高效高分辨率合成方法,通过噪声刷新策略提升效率,并首次发现能量衰减现象,通过调整引导超参数改善生成质量,兼容多种扩散模型技术。

Comments Project Page: https://zhenyangcs.github.io/RectifiedHR-Diffusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14878 2023-11-02 cs.LG cs.CV stat.CO stat.ML 72%

Restart Sampling for Improving Generative Processes

Yilun Xu, Mingyang Deng, Xiang Cheng, Yonglong Tian, Ziming Liu, Tommi Jaakkola

专题命中 效率与蒸馏 :diffusion(abstract,comments);text-to-image(abstract);分类 cs.CV

Comments Code is available at https://github.com/Newbeeer/diffusion_restart_sampling

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22060 2026-05-22 cs.CR cs.AI 71%

Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation

防范未经授权的知识蒸馏的文本到图像生成模型

Yilan Gao, Sida Huang, Hongyuan Zhang, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院、光学电子学院(iOPEN)、西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI)、中国电信) The University of Hong Kong(香港大学)

专题命中 效率与蒸馏 :text-to-image(title)

AI总结 本文提出WaveGuard,一种单次生成器基保护框架,通过在用户指定的扰动预算下保护发布的合成图像,以防止未经授权的知识蒸馏和能力复制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09232 2026-04-23 cs.LG math.ST stat.TH 71%

Diffusion Approximations for Thompson Sampling in the Small Gap Regime

差分方程近似在小间隙情形下的汤普森采样

Lin Fan, Peter W. Glynn

机构 * Kellogg School of Management, Northwestern University(西北大学凯洛格管理学院) Department of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程系)

专题命中 效率与蒸馏 :diffusion(title)

AI总结 研究小间隙情形下汤普森采样及相关采样算法的过程动态,展示其弱收敛于随机微分方程解,发现算法不变原理及对模型误设定的不敏感性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06605 2026-02-09 physics.optics 71%

Diffusion Schrödinger Bridges with enhanced posterior sampling for metasurface inverse design

基于增强后验采样的扩散薛定谔桥用于超材料反向设计

Mathys Le Grand, Pascal Urard, Denis Rideau, Loumi Trémas, Damien Maitre, Adam Fuchs, Louis-Henri Fernandez-Mouron, Régis Orobtchouk

专题命中 效率与蒸馏 :diffusion(title)

AI总结 本文提出基于增强后验采样的扩散薛定谔桥方法,用于高效高精度的超材料反向设计,解决了传统方法在复杂结构中的计算瓶颈和局部极小值问题。

Comments 47 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏