arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-03-03 至 2026-03-03 共收录 9 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 9 篇

2603.00483 2026-03-03 cs.CV cs.AI 85%

RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment

RAISE:基于需求的进化精炼用于无训练文本到图像对齐

Liyao Jiang, Ruichen Chen, Chao Gao, Di Niu

机构 * Department of ECE, University of Alberta, Canada(阿尔伯塔大学电子工程系) Huawei Technologies, Canada(华为技术有限公司)

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV

AI总结 RAISE是一种无训练、需求驱动的进化框架,通过动态调整生成过程实现高效且通用的文本到图像对齐。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04687 2026-03-03 cs.CL cs.CV cs.CY cs.HC 83%

Investigating Disability Representations in Text-to-Image Models

探究文本到图像模型中的残疾表示

Yang Tian, Yu Fan, Liudmila Zavolokina, Sarah Ebling

机构 * Department of Computational Linguistics University of Zurich(计算语言学系 苏黎世大学) Center for Law & Economics ETH Zurich(法律与经济学中心 伯尔尼联邦理工学院) Department of Information Systems University of Lausanne(信息系统系 瑞士洛桑大学)

专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 本文研究了文本到图像模型中残疾表示的问题,通过分析不同提示下的图像输出,揭示了残疾表示的不平衡,并提出通过缓解策略提升包容性。

Comments 21 pages, 9 figures. References included

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00523 2026-03-03 cs.CV 83%

SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation

SenseFlow: 为基于流的文本到图像蒸馏扩展分布匹配

Xingtong Ge, Xin Zhang, Tongda Xu, Yi Zhang, Xinjie Zhang, Yan Wang, Jun Zhang

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) SenseTime Research(商汤科技研究院) Vivix AI(维维克斯人工智能) Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究所)

专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 SenseFlow通过引入隐式分布对齐和段内指导,解决大规模基于流的文本到图像模型蒸馏中的收敛问题,提升蒸馏效果。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01579 2026-03-03 cs.CV cs.AI 79%

SkeleGuide: Explicit Skeleton Reasoning for Context-Aware Human-in-Place Image Synthesis

SkeleGuide: 基于显式骨骼推理的上下文感知人体图像合成

Chuqiao Wu, Jin Song, Yiyun Fei

机构 * Alibaba Group(阿里巴巴集团)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

AI总结 SkeleGuide通过显式骨骼推理提升上下文感知的人体图像合成质量,提供高保真且结构合理的生成结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03516 2026-03-03 cs.CV 79%

Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

绘画比思考更容易:文本到图像模型能否铺垫,却无法主导?

Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang, Rui Chen, Jiarong Ou, Xin Tao, Pengfei Wan, Xiaojuan Qi, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) The University of Hong Kong(香港大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出T2I-CoReBench基准测试,用于评估文本到图像模型的组合与推理能力,揭示现有模型在高组合场景和推理任务中的局限性。

Comments Accepted to ICLR 2026. Project Page: https://t2i-corebench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21278 2026-03-03 cs.CV cs.AI cs.LG 70%

Does FLUX Already Know How to Perform Physically Plausible Image Composition?

FLUX 是否已经能够进行物理上合理的图像合成?

Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang, Chen Zhao, Adams Wai-Kin Kong

机构 * Nanyang Technological University(南洋理工大学) Nanjing University(南京大学)

专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 FLUX能否通过SHINE框架实现物理合理的图像合成,通过引入无训练框架和降质抑制指导,提升高保真度和背景完整性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02026 2026-03-03 cs.CV cs.CL cs.LG 57%

Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT

学习如何注视:面向3D CT的疾病感知视觉-语言预训练

Simon Ging, Philipp Arnold, Sebastian Walter, Hani Alnahas, Hannah Bast, Elmar Kotter, Jiancheng Yang, Behzad Bozorgtabar, Thomas Brox

机构 * Computer Vision Group, University of Freiburg, Germany Adaptive \& Agentic AI (A3) Lab, Aarhus University, Denmark Department of Radiology, Medical Center -- University of Freiburg, Germany Chair of Algorithms Data Structures, University of Freiburg, Germany ELLIS Institute Finland School of Electrical Engineering, Aalto University, Finland

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 本文提出了一种面向3D CT的疾病感知视觉-语言预训练模型,通过对比预训练和基于提示的疾病监督,实现了文本到图像检索、疾病分类及内扫描片段定位的统一模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00882 2026-03-03 eess.IV cs.CV eess.SP 57%

Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priors

求解带标签MRI的非线性盲逆问题:结合物理和深度生成先验

Zhangxing Bian, Shuwen Wei, Samuel W. Remedios, Junyu Chen, Aaron Carass, Blake E. Dewey, Jerry L. Prince

机构 * Johns Hopkins University(约翰霍普金斯大学) Johns Hopkins School of Medicine(约翰霍普金斯医学院)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 本文提出了一种结合物理和深度生成先验的非线性盲逆框架,用于带标签MRI,实现了解剖恢复、高分辨率 cine 图像合成和运动估计的统一处理。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06566 2026-03-03 cs.CV 57%

Dynamic Uncertainty Learning with Noisy Correspondence for Text-Based Person Search

基于噪声对应关系的动态不确定性学习用于基于文本的人脸搜索

Zequn Xie, Haoming Ji, Chengxuan Li, Lingwei Meng

机构 * Zhejiang University(浙江大学) Beijing University of Posts Telecommunications(北京邮电大学) Beijing Forestry University(北京林业大学) Northwest Normal University(西北师范大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 本文提出DURA框架,通过动态不确定性学习和关系对齐方法,提升基于文本的人脸搜索在噪声环境下的鲁棒性和检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏