arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-03-03 至 2026-03-03 共收录 161 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 117 篇

2202.00729 2026-03-03 econ.TH 78%

The Impact of Connectivity on the Production and Diffusion of Knowledge

连接性对知识生产与扩散的影响

Gustavo Manso, Farzad Pourbabaee

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究了连接性对知识生产与扩散的影响,发现高连接性可能抑制知识生产并降低社会福利。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18942 2026-03-03 cs.CV 77%

VeCoR -- Velocity Contrastive Regularization for Flow Matching

VeCoR -- 速度对比正则化用于流匹配

Zong-Wei Hong, Jing-lun Li, Lin-Ze Li, Shen Zhang, Yao Tang

机构 * JIIOV Technology(JIIOV技术公司)

专题命中 扩散模型 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 VeCoR通过引入速度对比正则化,提升流匹配模型的稳定性与生成质量,尤其在低步数和轻量级设置中表现优异。

Comments Accepted to Findings of CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02253 2026-03-03 cs.CV cs.AI cs.LG 70%

DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing

DragFlow: 通过基于区域的监督释放DiT先验以实现拖拽编辑

Zihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang, Zhuming Lian, Xinlei Yu, Adams Wai-Kin Kong

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(国立新加坡大学)

专题命中 扩散模型 :diffusion(abstract);image editing(abstract);分类 cs.CV

AI总结 DragFlow通过基于区域的监督利用FLUX先验,改进基于拖拽的图像编辑效果,实现对点式和区域式基线的超越。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00978 2026-03-03 cs.CV cs.AI 70%

EraseAnything++: Enabling Concept Erasure in Rectified Flow Transformers Leveraging Multi-Object Optimization

EraseAnything++: 通过多对象优化实现Rectified Flow Transformers中的概念擦除

Zhaoxin Fan, Nanxiang Jiang, Daiheng Gao, Shiji Zhou, Wenjun Wu

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University(北京未来区块链与隐私计算先进创新中心,人工智能学院,北航) University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 EraseAnything++通过多目标优化实现图像和视频扩散模型中的概念擦除,提升生成质量与时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02049 2026-03-03 cs.CV 57%

WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories

WorldStereo: 通过3D几何记忆桥接相机引导的视频生成与场景重建

Yisu Zhang, Chenjie Cao, Tengfei Wang, Xuhui Zuo, Junta Wu, Jianke Zhu, Chunchao Guo

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯文心)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 WorldStereo通过3D几何记忆模块实现相机引导视频生成与3D场景重建的高效融合,提升多视角一致性与重建质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01890 2026-03-03 cs.CV 57%

Resolving Blind Inverse Problems under Dynamic Range Compression via Structured Forward Operator Modeling

通过结构化前向算子建模解决动态范围压缩下的盲逆问题

Muyu Liu, Xuanyu Tian, Chenhe Du, Qing Wu, Hongjiang Wei, Yuyao Zhang

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家) ShanghaiTech University(上海科技大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出CaMB-Diff,通过结构化前向算子建模解决动态范围压缩下的盲逆问题,提升信号保真度和物理一致性。

Comments 16 pages, 10 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01594 2026-03-03 cs.CV 57%

Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference

偏好分数蒸馏:利用2D奖励对齐文本到3D生成

Jiaqi Leng, Shuyuan Tu, Haidong Cao, Sicheng Xie, Daoguo Dong, Zuxuan Wu, Yu-Gang Jiang

机构 * Fudan University(复旦大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出偏好分数蒸馏方法,利用2D奖励模型实现文本到3D生成的人类偏好对齐,无需3D训练数据,提升生成质量与扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01509 2026-03-03 cs.CV cs.AI 57%

Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling

通过提示优化和测试时扩展实现文本到视频生成的检索、细化与排序

Zillur Rahman, Alex Sheng, Cristian Meo

机构 * Algoverse AI

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出3R框架,通过提示优化和测试时扩展提升文本到视频生成的准确性与效率。

Comments 2026 ICLR TTU Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01328 2026-03-03 cs.CV cs.AI 57%

You Only Need One Stage: Novel-View Synthesis From A Single Blind Face Image

你只需一个阶段:从单张盲脸图像生成新颖视角

Taoyue Wang, Xiang Zhang, Xiaotian Li, Huiyuan Yang, Lijun Yin

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出一种单阶段方法NVB-Face,通过直接提取盲脸图像特征并利用扩散模型生成高质量一致的新视角人脸图像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01096 2026-03-03 cs.CV cs.AI cs.CL cs.LG 57%

Unified Vision-Language Modeling via Concept Space Alignment

通过概念空间对齐实现统一的视觉-语言建模

Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk

机构 * University of Edinburgh(爱丁堡大学) FAIR at Meta(Meta公司FAIR团队)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 通过概念空间对齐实现统一的视觉-语言建模,V-SONAR在多语言和多模态任务中超越现有模型。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26819 2026-03-03 eess.AS cs.AI cs.CV cs.SD 57%

See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement

看清说话者:通过语音优先引导和区域细化生成高分辨率的说话面孔

Jinting Wang, Jun Wang, Hei Victor Cheng, Li Liu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Speech and Acoustic Laboratory, Joy Future Academy, Jingdong Corporation(语音与声学实验室、Joy Future Academy、京东公司) Aarhus University(阿arhus大学)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本研究提出一种通过语音直接生成高分辨率说话面孔的方法,结合扩散模型与统计先验,实现高质量视频生成。

Comments 16 pages,15 figures, accepted by TASLP

Journal ref EEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 4267-4281, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12586 2026-03-03 cs.CV 57%

There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training

不存在VAE:通过自监督预训练实现端到端像素空间生成建模

Jiachen Lei, Keli Liu, Julius Berner, Haiming Yu, Hongkai Zheng, Jiahong Wu, Xiangxiang Chu

机构 * AMAP, Alibaba Group(阿里集团AMAP) Caltech(加州理工学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出了一种端到端像素空间生成建模框架,通过自监督预训练实现扩散和一致性模型的高性能,超越了传统VAE和扩散模型的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23566 2026-03-03 cs.CV 57%

Towards Interpretable Visual Decoding with Attention to Brain Representations

向基于脑表示的可解释视觉解码迈进

Pinyuan Feng, Hossein Adeli, Wenxuan Guo, Fan Cheng, Ethan Hwang, Nikolaus Kriegeskorte

机构 * Zuckerman Mind Brain Behavior Institute, Columbia University, USA(祖克曼心智大脑行为研究所,哥伦比亚大学,美国)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出NeuroAdapter框架,通过直接条件化潜在扩散模型于脑表示,实现更透明的脑-图像解码与重建。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12734 2026-03-03 cs.SD cs.AI cs.GR cs.HC eess.AS 57%

SounDiT: Geo-Contextual Soundscape-to-Landscape Generation

SounDiT:基于地理情境的声音景观到景观生成

Junbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang, Teng Fei, Qixing Huang, Bing Zhou, Zhengzhong Tu, Shan Ye, Yuhao Kang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Tennessee, Knoxville(田纳西大学基洛纳分校) University of South Carolina(南卡罗来纳大学) Arizona State University(亚利桑那州立大学) University of Canterbury(坎特伯雷大学) Texas A&M University(德克萨斯A&M大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 扩散模型 :diffusion(abstract);分类 cs.GR

AI总结 SounDiT通过结合环境声音景观和地理情境条件,生成地理上一致的景观图像,并引入Place Similarity Score评估生成一致性。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00144 2026-03-03 cs.CV cs.AI 57%

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

解耦层次变分自编码器用于3D人-人交互生成

Zichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu, Wei Liu, Ajmal Mian

机构 * Department of CSSE, The University of Western Australia(西澳大学计算机科学与工程系) Data61, CSIRO(澳大利亚联邦科学工业研究组织Data61部门) Australian Institute for Machine Learning, The University of Adelaide(澳大利亚阿德莱德大学人工智能研究所) NERC-RVC, Hunan University(湖南大学NERC-RVC部门)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出DHVAE,通过解耦层次变分自编码器生成结构化且可控的3D人-人交互,提升运动保真度和物理合理性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00116 2026-03-03 cs.CV 57%

VoxelDiffusionCut: Non-destructive Internal-part Extraction via Iterative Cutting and Structure Estimation

VoxelDiffusionCut:通过迭代切割和结构估计实现非破坏性内部部件提取

Takumi Hachimine, Yuhwan Kwon, Cheng-Yu Kuo, Tomoya Yamanokuchi, Takamitsu Matsubara

机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,奈良科学技术研究所)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 VoxelDiffusionCut通过迭代切割和结构估计方法,利用扩散模型估计体素结构以实现非破坏性内部部件提取。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02131 2026-03-03 stat.AP cs.SI physics.soc-ph stat.OT 50%

Socio-Spatial Patterns of Suicide Mortality in the United States

美国自杀死亡的社会空间模式

Kushagra Tiwari, M. Amin Rahimian, Marie-Laure Charpignon, Philippe J. Giabbanelli, Praveen Kumar

专题命中 扩散模型 :diffusion(abstract)

AI总结 研究通过整合社会连通性指数与自杀死亡数据,揭示了社会网络对自杀风险及保护性政策传播的影响。

Comments Code and data: https://github.com/kut97/suicide-sci

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02015 2026-03-03 cs.LG 50%

CausalWrap: Model-Agnostic Causal Constraint Wrappers for Tabular Synthetic Data

CausalWrap: 用于表格合成数据的模型无关因果约束包装器

Amir Asiaee, Zhuohui J. Liang, Chao Yan

机构 * Department of Biostatistics, Vanderbilt University Medical Center(生物统计学系,范德比尔特大学医学中心) Department of Biomedical Informatics, Vanderbilt University Medical Center(生物医学信息学系,范德比尔特大学医学中心)

专题命中 扩散模型 :diffusion(abstract)

AI总结 CausalWrap通过注入部分因果知识提升表格合成数据的因果忠实度,减少治疗效应误差,提升因果推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01935 2026-03-03 cs.LG cs.AI 50%

Dream2Learn: Structured Generative Dreaming for Continual Learning

Dream2Learn: 结构化生成式做梦用于持续学习

Salvatore Calcagno, Matteo Pennisi, Federica Proietto Salanitri, Amelia Sorrenti, Simone Palazzo, Concetto Spampinato, Giovanni Bellitto

机构 * PeRCeiVe Lab, University of Catania, Italy(卡塔尼亚大学PeRCeiVe实验室)

专题命中 扩散模型 :diffusion(abstract)

AI总结 Dream2Learn通过生成梦中类别来提升持续学习的适应性,利用内部合成数据自我训练,实现更好的知识迁移和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19638 2026-03-03 cond-mat.mtrl-sci 50%

DFT-informed Design of Radiation-Resistant Dilute Ternary Cu Alloys

基于DFT的辐射抗性稀释三元铜合金设计

Vaibhav Vasudevan, Thomas Schuler, Pascal Bellon, Robert Averback

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究利用DFT方法设计抗辐射的稀释三元铜合金,通过添加结合空位的溶质减少空位迁移,抑制溶质拖拽效应。

Journal ref Physical Review Materials, 10(3), 033602 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19739 2026-03-03 eess.SP cs.ET 50%

Molecular Communication for Gastroretentive Drug Delivery

分子通信用于胃滞留药物递送

Sebastian Lotter, Marco Seiter, Maryam Pirmoradi, Lukas Brand, Dagmar Fischer, Robert Schober

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出了一种基于物理的模型,用于研究细菌纳米纤维素作为药物递送系统中聚合物涂层对药物释放的影响,以实现可控的药物释放。

Comments 6 pages, 2 figures, This paper has been accepted as Transactions Letter at IEEE Transactions on Molecular, Biological, and Multi-Scale Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04071 2026-03-03 quant-ph cond-mat.stat-mech 50%

Destruction and recovery of the entanglement entropy of a many-body quantum system after a single measurement

单次测量后多体量子系统的纠缠熵破坏与恢复

Bo Fan, Can Yin, Antonio M. García-García

专题命中 扩散模型 :diffusion(abstract)

AI总结 研究单次测量对多体量子系统纠缠熵的影响,发现不同测量协议下分布呈现高斯或指数特征,边界站点主导分布特性。

Comments v4: update manuscript, version as published

Journal ref Phys. Rev. B 113, 054304 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02645 2026-03-03 physics.comp-ph cs.AI cs.LG cs.NA math.NA 50%

Astral: training physics-informed neural networks with error majorants

Astral: 通过误差上界训练物理信息神经网络

Vladimir Fanaskov, Tianchi Yu, Alexander Rudikov, Ivan Oseledets

机构 * INM(研究所)

专题命中 扩散模型 :diffusion(abstract)

AI总结 Astral通过误差上界训练物理信息神经网络,实现更准确的误差估计和更快的收敛速度。

Comments Accepted to ICLR 2026 workshop AI&PDE, reviewed at https://openreview.net/forum?id=TcFpJK2FcN

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01727 2026-03-03 physics.flu-dyn physics.comp-ph 50%

Numerical method for strongly variable-density flows at low Mach number: flame-sheet regularisation and a mass-flux immersed boundary method

低马赫数强变密度流动的数值方法:火焰片正则化与质量通量浸入边界法

Matheus P. Severino, Fernando F. Fachini, Elmer M. Gennaro, Daniel Rodríguez, Leandro F. Souza

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出一种针对低马赫数强变密度流动的数值方法,结合火焰片正则化和质量通量浸入边界法,用于模拟燃烧系统中的复杂流动问题。

Comments 49 pages, 14 figures, 8 tables. This work is part of the first author's Ph.D. thesis at the University of São Paulo (USP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01077 2026-03-03 math.DS cs.NA cs.SY eess.SY math.NA stat.ML 50%

Kernel Methods for Stochastic Dynamical Systems with Application to Koopman Eigenfunctions: Feynman-Kac Representations and RKHS Approximation

用于随机动力系统核方法的Koopman特征函数:Feynman-Kac表示与RKHS近似

Boumediene Hamzi, Houman Owhadi, Umesh Vaidya

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出用于随机动力系统的核方法,通过Feynman-Kac表示和RKHS近似,扩展了Koopman特征函数的构造,并分析了扩散对数值稳定性的提升作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09755 2026-03-03 cond-mat.mtrl-sci 50%

Giant thermopower changes related to the resistivity maximum and colossal magnetoresistance in EuCd2P2

与电阻率最大值相关的巨大热电功率变化及巨磁阻效应在EuCd2P2中

Judith Grafenhorst, Sarah Krebber, Kristin Kliemt, Cornelius Krellner, Elena Hassinger, Ulrike Stockert

专题命中 扩散模型 :diffusion(abstract)

AI总结 EuCd2P2中观察到的巨热电功率效应源于电阻率最大值和巨磁阻效应,通过电子性质梯度实现巨大热电功率值。

Comments 5 Figures, 21 pages; caption of table 1 corrected, slightly shortened introduction + discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01307 2026-03-03 math.PR math-ph math.MP 50%

Inversions of stochastic processes from ergodic measures of Nonlinear SDEs

非线性随机微分方程的厄米特测度下的随机过程反问题

Hongyu Liu, Zhihui Liu

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出了一种基于非线性随机微分方程厄米特测度的反问题研究方法,通过分析静止福克-平克方程的解的唯一性,探讨了漂移和扩散项的恢复问题,并提供了唯一恢复失败的反例。

Journal ref J. Inverse Ill-Posed Probl., 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19473 2026-03-03 cs.LG cs.AI 50%

WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning

WavefrontDiffusion: 动态解码调度以提升推理性能

Haojin Yang, Rui Hu, Zequn Sun, Rui Zhou, Yujun Cai, Yiwei Wang

机构 * School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学软件新技术国家重点实验室) The University of Queensland(昆士兰大学) University of California, Merced(加州大学梅尔德分校)

专题命中 扩散模型 :diffusion(abstract)

AI总结 WavefrontDiffusion通过动态解码调度提升推理和代码生成的性能与语义连贯性。

Comments 19 pages. 3 figures

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06740 2026-03-03 math.NA cs.NA 50%

Elliptic Interface Problem approximated by CutFEM: II. A posteriori error analysis based on equilibrated fluxes

用CutFEM方法近似椭圆界面问题:II. 基于平衡通量的后验误差分析

Daniela Capatina, Aimene Gouasmi

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文提出基于平衡通量的后验误差分析方法,用于解决具有不连续扩散系数的椭圆界面问题,通过CutFEM方法实现高精度求解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00888 2026-03-03 cs.LG stat.ML 50%

Probabilistic Learning and Generation in Deep Sequence Models

深度序列模型中的概率学习与生成

Wenlong Chen

专题命中 扩散模型 :diffusion(abstract)

AI总结 本研究通过深度序列模型的架构设计,改进概率推断和结构,提升模型的不确定性和生成能力。

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏