arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86539 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

1907.13418 2019-08-01 eess.IV cs.CV cs.LG stat.ML 70%

Uncertainty Quantification in Deep Learning for Safer Neuroimage Enhancement

Ryutaro Tanno, Daniel Worrall, Enrico Kaden, Aurobrata Ghosh, Francesco Grussu, Alberto Bizzi, Stamatios N. Sotiropoulos, Antonio Criminisi, Daniel C. Alexander

专题命中 文生图 :diffusion(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.04378 2019-07-11 cs.CV cs.CL cs.LG eess.AS eess.IV 70%

M3D-GAN: Multi-Modal Multi-Domain Translation with Universal Attention

Shuang Ma, Daniel McDuff, Yale Song

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01478 2019-07-03 cs.CV cs.LG 70%

Obj-GloVe: Scene-Based Contextual Object Embedding

Canwen Xu, Zhenzhong Chen, Chenliang Li

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV

Comments 14 pages; not the final version

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01187 2019-07-03 cs.CV 70%

Generative Guiding Block: Synthesizing Realistic Looking Variants Capable of Even Large Change Demands

Minho Park, Hak Gu Kim, Yong Man Ro

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV

Comments This work is accepted in ICIP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.00286 2019-05-02 cs.CV 70%

Learn to synthesize and synthesize to learn

Behzad Bozorgtabar, Mohammad Saeed Rad, Hazım Kemal Ekenel, Jean-Philippe Thiran

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV

Comments Accepted to Computer Vision and Image Understanding (CVIU)

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05729 2019-04-12 cs.CV 70%

FTGAN: A Fully-trained Generative Adversarial Networks for Text to Face Generation

Xiang Chen, Lingbo Qing, Xiaohai He, Xiaodong Luo, Yining Xu

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.08170 2018-11-30 cs.CV 70%

Turbo Learning for Captionbot and Drawingbot

Qiuyuan Huang, Pengchuan Zhang, Dapeng Wu, Lei Zhang

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments in proceedings of NeurIPS 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04694 2018-04-16 cs.CV 70%

A Variational U-Net for Conditional Appearance and Shape Generation

Patrick Esser, Ekaterina Sutter, Björn Ommer

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV

Comments CVPR 2018 (Spotlight). Project Page at https://compvis.github.io/vunet/

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04338 2018-04-13 cs.CV 70%

MelanoGANs: High Resolution Skin Lesion Synthesis with GANs

Christoph Baur, Shadi Albarqouni, Nassir Navab

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.6497 2015-01-22 cs.LG cs.CV stat.ML 70%

Bayesian CP Factorization of Incomplete Tensors with Automatic Rank Determination

Qibin Zhao, Liqing Zhang, Andrzej Cichocki

专题命中 文生图 :inpainting(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17203 2026-08-19 stat.ML cs.LG math.ST stat.TH 新提交 67%

Expressivity In Multimodal Contrastive Learning

多模态对比学习中的表达能力

Andrew Stuart, Florian Wolf

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 该研究针对多模态对比学习的表达能力展开分析,明确CLIP架构的表达能力随模态数量变化,提出Hadamard-CLIP模型,实现任意数量模态联合分布的通用近似且保留CLIP的检索优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00807 2026-07-30 cs.AI 版本更新 67%

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

BioPro: 关于差分意识的性别公平性在视觉-语言模型中

Yujie Lin, Jiayao Ma, Qingguo Hu, Wenbo Li, Genji Li, Derek Wong, Jinsong Su

机构 * School of Informatics, Xiamen University(厦门大学信息学院) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 BioPro通过差分意识的性别公平性方法,在视觉-语言模型中实现选择性去偏,减少中性情境中的性别偏见同时保持显性情境中的性别忠实性。

Comments ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25479 2026-07-29 cs.CR cs.AI cs.LG 新提交 67%

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

通过表示引导在视觉语言模型供应链中植入架构后门

Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cinà, Luca Oneto, Iacopo Masi, Fabio Roli

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 研究视觉语言模型供应链安全问题,提出通过表示引导植入架构后门的攻击方法,该方法不影响训练数据等,通过触发控制模型表示转向攻击者目标,评估表明其损害模型多项性能,还提出审计防御方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08307 2026-06-30 stat.ML cs.LG 67%

Concentration bounds on response-based vector embeddings of black-box generative models

基于响应的生成模型向量嵌入的集中界

Aranyak Acharyya, Joshua Agterberg, Youngser Park, Carey E. Priebe

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

AI总结 本文研究了基于响应的生成模型向量嵌入的集中界,通过数据核视角空间嵌入方法,在适当正则条件下推导出样本向量嵌入的高概率集中界,并确定所需样本响应数量以实现目标精度的群体嵌入近似。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12688 2026-06-16 cs.LG cs.AI cs.DC 新提交 67%

M*: A Modular, Extensible, Serving System for Multimodal Models

M*: 一个模块化、可扩展的多模态模型服务系统

Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

AI总结 提出M*系统,通过将模型表示为数据流图并引入Walk Graph抽象,支持多模态复合模型的高效服务,在多个任务上降低延迟并提升吞吐量。

Comments The codebase is available at https://github.com/mstar-project/mstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12445 2026-06-16 cs.AI cs.LG stat.ML 版本更新 67%

Computational Safety for Generative AI: A Hypothesis Testing Perspective

生成式AI的计算安全性:假设检验视角

Pin-Yu Chen

机构 * IBM Research(IBM研究院)

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

AI总结 本文从假设检验角度形式化生成式AI的计算安全性,提出基于信号处理的方法检测恶意输入和AI生成内容。

Comments Extended version of the paper presented at the ICML 2026 Workshop on Hypothesis Testing

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06295 2026-05-08 cs.LG cs.AI stat.ML 67%

Attributions All the Way Down? The Metagame of Interpretability

逐层归因?可解释性中的元游戏

Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli

机构 * University of Warsaw(华沙大学) Centre for Credible AI, Warsaw University of Technology(可信AI中心,华沙技术大学) Bielefeld University(比勒菲尔德大学) LMU Munich, MCML(慕尼黑大学LMU,MCML)

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

AI总结 本文提出元游戏框架,通过将归因方法视为合作博弈计算Shapley值,研究模型解释的第二阶交互效应,揭示归因的层级分解及其在多领域可解释性应用中的价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06061 2026-04-30 cs.LG 67%

PromptEvolver: Prompt Inversion through Evolutionary Optimization in Natural-Language Space

PromptEvolver:通过自然语言空间中的进化优化实现提示倒置

Asaf Buchnick, Aviv Shamsian, Aviv Navon, Ethan Fetaya

机构 * Bar-Ilan University(巴伊兰大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文提出PromptEvolver,通过进化优化生成自然语言提示,实现高保真图像重建,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00319 2026-04-02 cs.AI cs.MA 67%

Collaborative AI Agents and Critics for Fault Detection and Cause Analysis in Network Telemetry

协同AI代理与批评者用于网络遥测中的故障检测与原因分析

Syed Eqbal Alam, Zhan Shu

机构 * Department of Electrical and Computer Engineering, University of Alberta(阿尔伯塔大学电气与计算机工程系) SheQAI Research(SheQAI研究)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文提出一种多代理联邦系统,通过AI代理与批评者协作完成网络遥测中的故障检测、严重性和原因分析等多模态任务,利用多时间尺度随机逼近技术保证收敛性,通信开销低且隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23140 2026-03-26 cs.LG 67%

DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative Models

DAK-UCB: 为LLMs和生成模型的多样性感知提示路由

Donya Jafari, Farzan Farnia

机构 * Sharif University of Technology(谢赫拉扎德技术大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文提出DAK-UCB方法,结合保真度与多样性指标,用于在线选择生成模型,以提升生成结果的多样性同时保持保真度。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12101 2026-03-17 cs.RO cs.AI cs.LG 67%

Decoupled Action Expert: Confining Task Knowledge to the Conditioning Pathway

解耦动作专家:将任务知识限制在条件路径中

Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU

专题命中 文生图 :diffusion(abstract);image synthesis(abstract)

AI总结 本文提出解耦训练方法,将任务知识限制在条件路径中,使动作主干网络无任务依赖,通过Diffusion Policy验证,证明动作专家编码较少任务特定知识。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13116 2026-03-16 cs.HC 67%

Memory Printer: Exploring Everyday Reminiscing by Combining Slow Design with Generative AI-based Image Creation

记忆打印机:通过结合缓慢设计与基于生成AI的图像创作探索日常回忆

Zhou Fang, Janet Yi-Ching Huang

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文提出Memory Printer,结合丝网印刷隐喻与文本到图像生成,通过分层重建、物理木刮刀和内置打印,探索慢互动如何重塑人机关系,发现记忆唤起与控制感提升等机遇及算法偏见等挑战。

Comments Accepted to CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17028 2026-02-25 cs.LG cs.AI 67%

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

通过Cauchy-Schwarz散度进行分布式视觉语言对齐

Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu, Jiayi Shen, Jan-Jakob Sonke, Efstratios Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) The Netherlands Cancer Institute(荷兰癌症研究所) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) The Arctic University of Norway(挪威北极大学) Singapore Management University(新加坡管理大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文提出CS-Aligner框架,通过整合Cauchy-Schwarz散度与互信息,实现更紧密的视觉语言分布对齐,提升跨模态生成与检索性能。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18216 2026-02-23 cs.LG 67%

Generative Model via Quantile Assignment

通过分位数分配生成模型

Georgi Hrusanov, Oliver Y. Chén, Julien S. Bodelet

专题命中 文生图 :diffusion(abstract);image synthesis(abstract)

AI总结 NeuroSQL通过隐式学习低维潜在表示,无需辅助网络,实现高效稳定的合成数据生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00849 2026-02-03 cs.LG cs.AI cs.NA math.NA 67%

RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation

RMFlow:通过噪声注入步骤细化均流以实现多模态生成

Yuhao Huang, Shih-Hsin Wang, Andrea L. Bertozzi, Bao Wang

机构 * Department of Mathematics and Scientific Computing and Imaging (SCI) Institute University of Utah(数学与科学计算及成像学院(SCI)院,犹他大学) Department of Mathematics, UCLA(数学系,加州大学洛杉矶分校)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 RMFlow通过引入噪声注入步骤,改进均流模型,实现高效多模态生成,仅需单次功能评估即可达到接近最先进的性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02725 2026-02-02 cs.AI 67%

Advances in Artificial Intelligence: A Review for the Creative Industries

人工智能进展:面向创意产业的综述

Nantheera Anantrasirichai, Fan Zhang, David Bull

机构 * Visual Information Laboratory, University of Bristol, Bristol, UK(布里斯托大学视觉信息实验室)

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

AI总结 本文综述了自2022年以来人工智能在创意产业中的进展,探讨了生成式AI、大语言模型和扩散模型等技术对创意生产流程的影响,并分析了人类与AI协作的新趋势及面临的挑战。

Comments This is an updated review of our previous paper (see https://doi.org/10.1007/s10462-021-10039-7), and has been accepted by Artificial Intelligence Review journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10096 2026-01-22 cs.LG 67%

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

多语言到多模态(M2M):通过单语文本解锁新语言

Piyush Singh Pasi

机构 * Amazon(亚马逊)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 M2M通过单语文本学习多语言多模态对齐,实现多语言文本到图像检索的零样本迁移

Comments EACL 2026 Findings accepted. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18867 2026-01-07 cs.AI 67%

Topological Perspectives on Optimal Multimodal Embedding Spaces

拓扑视角下的最优多模态嵌入空间

Abdul Aziz A. B, A. B Abdul Rahim

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文通过拓扑数据分析比较CLIP和CLOOB的嵌入空间,揭示其模态差距驱动因素和维度坍缩的影响,为多模态模型优化提供新视角。

Comments This manuscript contains substantive technical inaccuracies and an incomplete treatment of the stated topic. Subsequent developments and a reassessment of the problem indicate that the scope and framing of the work do not adequately reflect the current state of research, and the analysis is therefore incomplete and outdated

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02219 2026-01-06 eess.SP 67%

Beam-Brainstorm: A Generative Site-Specific Beamforming Approach

Beam-Brainstorm: 一种生成式特定地点波束成形方法

Zihao Zhou, Zhaolin Wang, Yuanwei Liu

专题命中 文生图 :diffusion(abstract);image synthesis(abstract)

AI总结 本文提出了一种生成式特定地点波束成形方法,通过联合结构建模和自定义扩散模型,实现高效且高质量的用户特定波束生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15442 2025-12-18 cs.LG 67%

Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting

通过链式推理和任务指令提示降低版权侵权风险

Neeraj Sarna, Yuanyuan Li, Michael von Gablenz

机构 * Munich RE(慕尼黑RE)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文通过链式推理和任务指令提示结合负提示和提示重写,降低生成图像的版权侵权风险,并评估不同模型复杂度下的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏