arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 1299 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 1299 篇

2603.16271 2026-07-01 cs.CV 版本更新 57%

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

VIGOR: 面向视频几何的时序生成对齐奖励

Tengjiao Yin, Jinglei Shi, Heng Guo, Xi Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出基于几何的奖励模型,利用预训练几何基础模型通过跨帧重投影误差评估多视图一致性,以点方式计算误差,并引入几何感知采样策略,通过后训练和推理时优化对齐视频扩散模型,提升生成视频的几何一致性。

Comments Project Page: https://vigor-geometry-reward.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05851 2026-07-01 cs.CV 版本更新 57%

VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction

VS3R:基于深度三维重建的鲁棒全帧视频稳定

Muhua Zhu, Xinhao Jin, Xinping Wang, Yu Zhang, Yifei Xue, Tie Ji, Yizhen Lao

机构 * Hunan University(湖南大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出VS3R框架,结合前馈三维重建与生成式视频扩散,通过联合估计相机参数、深度和掩码,并引入混合稳定渲染模块和稳定驱动扩散模型,实现高保真全帧视频稳定,在鲁棒性和视觉质量上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03710 2026-06-29 cs.CV cs.AI 版本更新 57%

MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction

MPFlow: 多模态后验引导的流匹配用于零样本MRI重建

Seunghoi Kim, Chen Jin, Henry F. J. Tregidgo, Matteo Figini, Daniel C. Alexander

机构 * 1 Hawkes Institute, UCL \, 2 Dept. of Medical Physics Biomedical Engineering, UCL 3 Dept. of Computer Science, UCL 4 Centre for AI, DS\&AI, AstraZeneca, UK

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出MPFlow框架,利用自监督预训练PAMRI学习跨模态共享表示,在推理时通过数据一致性和特征对齐引导采样,减少幻觉并提升零样本MRI重建效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09241 2026-06-29 cs.CV cs.RO 版本更新 57%

RAE-NWM: Navigation World Model in Dense Visual Representation Space

RAE-NWM:密集视觉表示空间中的导航世界模型

Mingkun Zhang, Wangtian Shen, Fan Zhang, Haijian Qin, Zihao Pei, Ziyang Meng

机构 * Department of Precision Instrument, Tsinghua University(清华大学精密仪器系) University of Rochester(罗切斯特大学) Beijing Information Science and Technology University(北京信息科技大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出RAE-NWM,在密集视觉表示空间中使用条件扩散Transformer建模导航动态,提升结构稳定性和动作精度,从而改善下游规划与导航。

Comments Code is available at: https://github.com/20robo/raenwm

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12066 2026-06-29 cs.CV 版本更新 57%

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

通过视频到视频翻译学习随机桥接用于视频对象移除

Zijie Lou, Xiangwei Feng, Jiaxin Wang, Jiangtao Yao, Fei Che, Tianbao Liu, Chengjing Wu, Xiaochao Qu, Luoqi Liu, Ting Liu

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出基于随机桥接模型的视频到视频翻译框架,利用输入视频的结构先验引导对象移除,并通过自适应掩码调制策略平衡背景保真度与生成灵活性,显著提升视觉质量和时间一致性。

Comments Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16893 2026-06-29 cs.CV 版本更新 57%

Instant Expressive Gaussian Head Avatars at Over 100 FPS

即时表达性高斯头部虚拟化身,帧率超过100 FPS

Kaiwen Jiang, Xueting Li, Seonwook Park, Ravi Ramamoorthi, Shalini De Mello, Koki Nagano

机构 * University of California, San Diego(加州大学圣地亚哥分校) NVIDIA(英伟达)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出一种前馈编码器流水线,将单张野外图像转换为3D一致、快速且富有表现力的可动画化表示,通过轻量级局部融合策略实现高动画表现力,运行速度达107.31 FPS。

Comments Project website is https://research.nvidia.com/labs/amri/projects/instant4d

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23814 2026-06-26 cs.CV cs.AI 版本更新 57%

Mapping License Plate Recoverability Under Extreme Viewing Angles for Opportunistic Urban Sensing

在极端视角下映射车牌可恢复性以实现机会性城市感知

Igor Adamenko, Orpaz Ben Aharon, Yehudit Aperstein, Alexander Apartsin

机构 * Afeka Academic College of Engineering(阿法卡工程学院) Holon Institute of Technology(霍隆理工学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文研究了在极端视角下通过机会性城市感知实现车牌识别的可恢复性问题,提出了一种任务无关的方法来量化恢复边界,结合合成退化参数密集扫描与两个总结指标,评估了不同恢复模型的性能。

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12062 2026-06-26 cs.CV 版本更新 57%

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

基于语言驱动的序列级模态不变表示学习用于视频可见光-红外行人重识别

Xiaomei Yang, Antai Liu, Xizhan Gao, Fa Zhu, Sijie Niu, Giancarlo Fortino

机构 * Shandong Key Laboratory of Ubiquitous Intelligent Computing, School of Information Science and Engineering, University of Jinan(山东省 Ubiquitous Intelligent Computing 重点实验室,济南大学信息科学与工程学院) College of Information Science and Technology & Artificial Intelligence, Nanjing Forestry University(信息科学与技术及人工智能学院,南京林业大学) Department of Informatics, Modeling, Electronics, and Systems, University of Calabria(信息学、建模、电子与系统系,卡利博大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出LSMRL方法,通过CLIP的时空特征学习、语义扩散和跨模态交互模块,结合模态级损失,学习序列级模态不变表示,在VVI-ReID任务上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22121 2026-06-25 cs.CV 版本更新 57%

MotionDPS: Motion-Compensated 3D Brain MRI Reconstruction

MotionDPS: 3D脑部MRI重建中的运动补偿

Antonio Ortiz-Gonzalez, Erich Kobler, Lukas Schletter, Alexander Effland

机构 * Life and Medical Sciences Institute, University of Bonn(波恩大学生命与医学科学研究所) Institute for Machine Learning, LIT AI Lab, Department of Virtual Morphology, Clinical Research Institute Medical AI, Johannes Kepler University Linz(林茨约翰尼斯·凯撒大学机器学习研究所、LIT AI实验室、虚拟形态部门、医学人工智能临床研究机构) German Center for Neurodegenerative Diseases (DZNE)(德国神经退行性疾病研究中心(DZNE)) Institute for Applied Mathematics, University of Bonn(波恩大学应用数学研究所)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出了一种统一的贝叶斯框架,用于运动补偿的3D MRI重建,通过直接从运动损坏的k空间数据中联合估计解剖图像、刚体运动参数和线圈灵敏度图,实现了无需配对无运动训练数据的完全无监督重建。

Comments This work has been accepted for publication in IEEE Transactions on Medical Imaging (TMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10702 2026-06-25 cs.CV cs.AI 版本更新 57%

Backbone-Conditional Behavior of Modality Gating in Multi-Modal Prostate MRI Segmentation: A 5-Fold Cross-Validation and Gate Mechanism Analysis

模态隔离门控融合:用于稳健多模态前列腺MRI分割的架构无关方法

Yongbo Shu, Wenzhao Xie, Shanhu Yao, Zirui Xin, Luo Lei, Kewen Chen, Aijing Luo

机构 * The Second Xiangya Hospital of Central South University(中南大学湘雅医学院第二医院) The Third Xiangya Hospital of Central South University(中南大学湘雅医学院第三医院) School of Life Sciences, Central South University(中南大学生命科学学院) Hunan Provincial Key Laboratory of Medical Information Research (Central South University)(湖南省医学信息研究重点实验室(中南大学)) Hunan Provincial Clinical Medical Research Center for Cardiovascular Intelligent Medicine(湖南省心血管智能医学临床医学研究中心)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出模态隔离门控融合(MIGF)方法,通过独立编码流和门控阶段提升多模态前列腺MRI分割的鲁棒性,实验表明其在不同模态缺失和伪影情况下均有效。

Comments Major revision. Single-fold analysis replaced by 5-fold cross-validation (180 trained models) plus a direct gate-mechanism analysis; conclusions updated to show that modality gating is backbone-conditional. Supersedes v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17773 2026-06-25 cs.CV 版本更新 57%

LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation

LoT-Pass: 面向图像到视频生成的长期鲁棒图像水印

Guanjie Wang, Zehua Ma, Han Fang, Weiming Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 针对图像引导视频生成中源图像追踪缺失的问题,提出跨模态水印框架I2VWM,通过视频模拟噪声层和光流对齐模块增强水印的时间鲁棒性,在开源和商业模型上验证了效果。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15490 2026-06-25 cs.CV cs.LG eess.IV 版本更新 57%

Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation

通过配对图像域检索和文本域增强提高3D脑MRI报告生成的事实准确性

Junhyeok Lee, Yujin Oh, Dahyoun Lee, Hyon Keun Joh, Minchul Kim, Chul-Ho Sohn, Sung Hyun Baik, Cheol Kyu Jung, Jung Hyun Park, Kyu Sung Choi, Byung-Hoon Kim, Jong Chul Ye

机构 * Cancer Biology, Seoul National University College of Medicine, Korea(首尔国立大学医学院癌症生物学系,韩国) Radiology, Massachusetts General Hospital(麻省总医院放射科) Harvard Medical School(哈佛医学院) Biomedical Systems Informatics, Yonsei University College of Medicine, Korea(延世大学医学院生物医学系统信息学系,韩国) Graduate School, Yonsei University, Korea(延世大学研究生院,韩国) Radiology, Seoul National University College of Medicine, Korea(首尔国立大学医学院放射科,韩国) Radiology, Seoul National University Hospital, Korea(首尔国立大学医院放射科,韩国) Radiology, Seoul National University Bundang Hospital, Korea(首尔国立大学 Bundang 医院放射科,韩国) Radiology, SMG-SNU Boramae Medical Center, Korea(SMG-SNU Boramae 医疗中心放射科,韩国) Psychiatry, Yonsei University College of Medicine, Korea(延世大学医学院精神病学系,韩国) Behavioral Sciences in Medicine, Yonsei University College of Medicine, Korea(延世大学医学院医学行为科学系,韩国) Yonsei Institute for Digital Health, Korea(延世大学数字健康研究院,韩国) Kim Jaechul Graduate School of AI, KAIST, Korea(金 Jaechul人工智能研究生院,韩国)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出PIRTA框架,通过检索相似3D DWI/ADC图像并利用其配对报告指导LLM生成,避免显式跨模态对齐,显著提高缺血区域准确性。

Comments MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05896 2026-06-24 cs.CV 版本更新 57%

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

共鸣心智:具备心智理论的闭环社交虚拟人

Jianxu Shangguan, Jing Xu, Hang Ye, Xiaoxuan Ma, Yizhou Wang, Jenq-Neng Hwang, Wentao Zhu

机构 * University of Washington(华盛顿大学) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学) Eastern Institute of Technology, Ningbo(宁波工程技术学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出一个闭环双智能体框架,通过整合感知、社会推理(基于心智理论)和多模态生成,实现具备社交智能的虚拟人,并在信息不对称数据集上取得优于全信息脚本模式的对话质量。

Comments Project page: https://resonantminds.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12683 2026-06-23 cs.CV q-bio.NC 版本更新 57%

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

Brain-DiT:一种基于元数据条件预训练的通用多状态fMRI基础模型

Junfeng Xia, Wenhao Ye, Xuanye Pan, Xinke Shen, Mo Wang, Quanying Liu

机构 * Department of Biomedical Engineering, Southern University of Science and Technology, China(南方科技大学生物医学工程系,中国) School of Biomedical Engineering, Shenzhen University, China(深圳大学生物医学工程学院,中国)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出Brain-DiT,一种基于扩散Transformer和元数据条件预训练的通用多状态fMRI基础模型,在24个数据集上预训练,通过生成式预训练学习多尺度表示,在7个下游任务中优于重建和对齐方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16192 2026-06-23 cs.CV 版本更新 57%

360Anything: Geometry-Free Lifting of Images and Videos to 360°

360Anything: 图像和视频的无几何约束360°生成

Ziyi Wu, Daniel Watson, Andrea Tagliasacchi, David J. Fleet, Marcus A. Brubaker, Saurabh Saxena

机构 * Google DeepMind(谷歌DeepMind) Simon Fraser University(西蒙弗雷泽大学) University of Toronto(多伦多大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出无几何约束框架360Anything,基于预训练扩散Transformer,以数据驱动方式学习透视到等距柱状投影映射,无需相机元数据,实现图像和视频到360°全景的高质量生成,并引入循环潜在编码消除拼接伪影。

Comments ECCV 2026. Project page: https://360anything.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22973 2026-06-23 cs.CV 版本更新 57%

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation

BIFE:更好的交互,更少的错误,用于分钟级视频生成

Zeyu Zhang, Jinyuan Mao, Shuning Chang, Yuanyu He, Yizeng Han, Jiasheng Tang, Fan Wang, Bohan Zhuang

机构 * DAMO Academy, Alibaba Group(阿里巴巴集团 DAMO 院) Zhejiang University(浙江大学) Hupan Lab(虎扑实验室)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出BIFE框架,通过语义稀疏KV缓存和块强制训练策略,解决长视频生成中的长程交互保持和误差累积问题,实现稳定连贯的分钟级视频生成,在VDE指标上提升约20%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00366 2026-06-23 physics.med-ph cs.AI cs.CV 版本更新 57%

AI-Augmented Thyroid Scintigraphy for Robust Classification of Disease

AI增强甲状腺闪烁显像用于疾病的鲁棒分类

Maziar Sabouri, Ghasem Hajianfar, Alireza Rafiei Sardouei, Milad Yazdani, Azin Asadzadeh, Soroush Bagheri, Mohsen Arabi, Seyed Rasoul Zakavi, Emran Askari, Atena Aghaee, Sam Wiseman, Dena Shahriari, Habib Zaidi, Arman Rahmim

机构 * organization= Department of Physics \& Astronomy, University of British Columbia , city= Vancouver , country= Canada organization= Department of Basic Translational Research, BC Cancer Research Institute , city= Vancouver , country= Canada Molecular Imaging, Department of Medical Imaging, Geneva University Hospital , city= Geneva , country= Switzerland organization= Department of Electrical Computer Engineering, University of British Columbia , city= Vancouver , country= Canada organization= Department of Nuclear Medicine, 5Azar Hospital, Golestan University of Medical Sciences , city= Gorgan , country= Iran organization= Department of Medical Physics, Kashan University of Medical Sciences , city= Kashan , country= Iran organization= Department of Pathology Radiology, School of Medicine, Alborz University of Medical Sciences , city= Karaj , country= Iran organization= Nuclear Medicine Research Center, Mashhad University of Medical Sciences , city= Mashhad , country= Iran organization= Department of Surgery, St. Paul's Hospital \& University of British Columbia , city= Vancouver , country= Canada organization= School of Biomedical Engineering, University of British Columbia , city= Vancouver , country= Canada organization= Department of Orthopaedics, Faculty of Medicine, University of British Columbia , city= Vancouver , country= Canada organization= Department of Radiology, University of British Columbia , city= Vancouver , country= Canada

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本研究探讨了稳定扩散、流匹配和常规增强三种数据增强策略对基于深度学习的甲状腺闪烁显像分类的影响,发现流匹配方法在性能上最优,结合原始数据实现了最高的F1分数和AUC值。

Journal ref Physica Medica 148 (2026) 105854

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17006 2026-06-19 cs.CV cs.RO 版本更新 57%

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

CoMo: 从互联网视频中学习连续潜在运动以实现可扩展的机器人学习

Jiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu, Kaijing Ma, Yating Wang, Gangshan Wu, Tong He, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Lab(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学) Fudan University(复旦大学) Tongji University(同济大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出CoMo方法,通过早期时间差分和时序对比学习从互联网视频中学习连续潜在运动,避免离散化信息损失,实现零样本泛化生成伪动作标签,联合训练策略在仿真和真实实验中表现优异。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15228 2026-06-19 cs.CV 版本更新 57%

Collaborative Multi-Modal Coding for High-Quality 3D Generation

协作多模态编码用于高质量3D生成

Ziang Cao, Zhaoxi Chen, Liang Pan, Ziwei Liu

机构 * S-Lab, Nanyang Technological University, Singapore(南洋理工大学S实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出TriMM,首个前馈式3D原生生成模型,通过协作多模态编码融合RGB、RGBD和点云特征,结合辅助2D/3D监督和三平面潜在扩散模型,实现高质量3D资产生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06361 2026-06-18 cs.CV 版本更新 57%

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

两步物理:在视觉细化之前锁定运动先验会擦除它们

Woojung Han, Seil Kang, Youngjun Jun, Min-Hung Chen, Fu-En Yang, Seong Jae Hwang

机构 * National Institute of Standards and Technology(国家标准与技术研究院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文发现图像到视频扩散模型在两步生成中比多步生成具有更好的物理一致性,通过频谱分析将原因归结为去噪过程中的相位侵蚀,并提出无需训练的PhaseLock框架,通过从两步推理中提取运动先验并利用潜在增量引导强制到高保真生成中,有效缓解相位退化,提升物理一致性平均6.2点,同时保持视觉保真度且开销极小。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21605 2026-06-18 cs.CV 版本更新 57%

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

S3OD:基于合成数据的通用显著目标检测

Orest Kupyn, Hirokatsu Kataoka, Christian Rupprecht

机构 * University of Oxford, VGG(牛津大学,视觉信息集团)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出S3OD方法,通过大规模合成数据生成和歧义感知架构,显著提升显著目标检测的跨数据集泛化能力,仅用合成数据训练即可降低20-50%误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09185 2026-06-18 cs.CV cs.AI 版本更新 57%

Learning Patient-Specific Disease Dynamics with Latent Flow Matching for Longitudinal Imaging Generation

学习患者特异性疾病动态:基于潜在流匹配的纵向影像生成

Hao Chen, Rui Yin, Yifan Chen, Qi Chen, Chao Li

机构 * University of Cambridge(剑桥大学) Nanjing First Hospital(南京第一医院) Nanjing Medical University(南京医科大学) Johns Hopkins University(约翰霍普金斯大学) University of Dundee(邓迪大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出Δ-LFM框架,利用流匹配对齐患者潜在轨迹,通过患者特异性潜在对齐实现单调疾病进展建模,在三个纵向MRI基准上验证了可解释性和性能。

Comments ICLR 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09439 2026-06-18 cs.CV 版本更新 57%

SuperCarver: Texture-Consistent 3D Geometry Super-Resolution for High-Fidelity Surface Detail Generation

SuperCarver: 纹理一致的3D几何超分辨率用于高保真表面细节生成

Qijian Zhang, Xiaozheng Jian, Xuan Zhang, Wenping Wang, Junhui Hou

机构 * Tencent Games, China(腾讯游戏,中国) Department of Computer Science & Engineering, Texas A & M University(电子与计算机工程系,德克萨斯A&M大学) Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出SuperCarver,一种3D几何超分辨率管线,通过先验引导的法线扩散模型和噪声鲁棒的逆渲染,为粗糙网格补充纹理一致的表面细节,实现高保真细节生成。

Comments Accepted in IEEE TVCG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09631 2026-06-18 cs.SD cs.CL cs.CV 版本更新 57%

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

DiFlow-TTS: 基于离散流匹配的紧凑低延迟零样本文本转语音

Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen

机构 * FPT Software AI Center(FPT软件AI中心) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出DiFlow-TTS框架,通过离散流匹配和分解离散流去噪器,在零样本TTS中实现高质量与低延迟的平衡。

Comments Accepted at Interspeech 2026 (Long Paper Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21615 2026-06-18 cs.CV 版本更新 57%

Epipolar Geometry Improves Video Generation Models

极线几何改进视频生成模型

Orest Kupyn, Théo Uscidda, Marta Tintore Gazulla, Fabian Manhardt, Federico Tombari, Christian Rupprecht

机构 * University of Oxford(牛津大学) Google Research(谷歌研究院) CREST-ENSAE, Institut Polytechnique de Paris(巴黎理工学院CREST-ENSAE研究中心) Technical University of Munich(慕尼黑技术大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 针对视频生成模型几何不一致和运动伪影问题,提出基于极线几何约束的偏好优化方法,在保持视觉质量的同时将极线误差降低31%,人类评分一致性从54%提升至72%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18245 2026-06-18 cs.CV cs.LG 版本更新 57%

VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset

VGGHeads: 基于大规模合成数据集的3D多头部对齐

Orest Kupyn, Eugene Khvedchenia, Christian Rupprecht

机构 * University of Oxford(牛津大学) Piñata Farms Ukrainian Catholic University(乌克兰天主大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出VGGHeads,一个由扩散模型生成的大规模合成数据集,用于单步同时进行头部检测和3D网格重建,在真实图像上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21583 2026-06-17 cs.CV cs.AI 版本更新 57%

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

基于流匹配的原理化强化学习从片段级策略优化中涌现

Yifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du, Keyu Fan, Bo Li, Sinan Du, Xu Wan, Zhiyu Chen, Bo Xia, Yongzhe Chang, Changqian Yu, Kun Gai, Tiantian Zhang, Xueqian Wang

专题命中 扩散模型 :text-to-image(abstract);分类 cs.CV

AI总结 本文提出了一种基于片段级策略优化的流匹配强化学习方法GCPO,通过将连续步骤聚合为相干片段并改变策略优化层级,有效缓解了优势归因不准确的问题,实验表明其在文本到图像生成任务中表现优于现有方法。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04704 2026-06-17 q-bio.QM cs.AI cs.CV 版本更新 57%

SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes

SPATIA: 空间细胞表型的多模态生成与预测

Zhenglun Kong, Mufan Qiu, John Boesen, Xiang Lin, Sukwon Yun, Tianlong Chen, Manolis Kellis, Marinka Zitnik

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学学校生物医学信息学系) Department of Computer Science, University of North Carolina(北卡罗来纳大学计算机科学系) Department of Computer Science, Massachusetts Institute of Technology(麻省理工学院计算机科学系)

专题命中 扩散模型 :image generation(abstract);分类 cs.CV

AI总结 提出SPATIA模型,融合细胞形态、基因表达和空间上下文,通过置信感知流匹配和形态-谱对齐实现多尺度生成与预测,在12项任务中优于18个基线模型。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05212 2026-06-17 cs.CV 版本更新 57%

FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching

FlowLet: 基于小波流匹配的条件性3D脑MRI合成

Danilo Danese, Angela Lombardi, Matteo Attimonelli, Giuseppe Fasano, Tommaso Di Noia

机构 * Politecnico di Bari(巴里理工学院) Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出FlowLet框架,利用可逆3D小波域中的流匹配生成年龄条件化的3D脑MRI,避免重建伪影并降低计算需求,实验证明其生成高保真体积且提升脑年龄预测模型对低代表性年龄组的性能。

Comments Accepted at Medical Image Analysis (Elsevier)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05163 2026-06-17 cs.CV 版本更新 57%

4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture

4DSloMo: 基于异步捕获的高速场景4D重建

Yutian Chen, Shi Guo, Tianshuo Yang, Lihe Ding, Xiuyuan Yu, Jinwei Gu, Tianfan Xue

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) NVIDIA

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 提出一种仅使用低帧率相机的高速4D捕获系统,通过异步捕获方案将等效帧率提升至100-200 FPS,并利用视频扩散模型修复稀疏视图伪影,实现高质量高速4D重建。

Comments Webpage: https://openimaginglab.github.io/4DSloMo/

详情

展开后加载摘要…

URL PDF HTML 收藏