Uncertainty Quantification in Deep Learning for Safer Neuroimage Enhancement
专题命中 文生图 :diffusion(abstract);image synthesis(abstract);分类 cs.CV
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :diffusion(abstract);image synthesis(abstract);分类 cs.CV
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV
Comments 14 pages; not the final version
专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV
Comments This work is accepted in ICIP 2019
专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV
Comments Accepted to Computer Vision and Image Understanding (CVIU)
专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
Comments in proceedings of NeurIPS 2018
专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV
Comments CVPR 2018 (Spotlight). Project Page at https://compvis.github.io/vunet/
专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV
专题命中 文生图 :inpainting(abstract);image synthesis(abstract);分类 cs.CV
多模态对比学习中的表达能力
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 该研究针对多模态对比学习的表达能力展开分析,明确CLIP架构的表达能力随模态数量变化,提出Hadamard-CLIP模型,实现任意数量模态联合分布的通用近似且保留CLIP的检索优势。
BioPro: 关于差分意识的性别公平性在视觉-语言模型中
机构 * School of Informatics, Xiamen University(厦门大学信息学院) ; Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 BioPro通过差分意识的性别公平性方法,在视觉-语言模型中实现选择性去偏,减少中性情境中的性别偏见同时保持显性情境中的性别忠实性。
Comments ACM Multimedia 2026
通过表示引导在视觉语言模型供应链中植入架构后门
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 研究视觉语言模型供应链安全问题,提出通过表示引导植入架构后门的攻击方法,该方法不影响训练数据等,通过触发控制模型表示转向攻击者目标,评估表明其损害模型多项性能,还提出审计防御方法。
基于响应的生成模型向量嵌入的集中界
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 本文研究了基于响应的生成模型向量嵌入的集中界,通过数据核视角空间嵌入方法,在适当正则条件下推导出样本向量嵌入的高概率集中界,并确定所需样本响应数量以实现目标精度的群体嵌入近似。
M*: 一个模块化、可扩展的多模态模型服务系统
机构 * Stanford University(斯坦福大学) ; University of Washington(华盛顿大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 提出M*系统,通过将模型表示为数据流图并引入Walk Graph抽象,支持多模态复合模型的高效服务,在多个任务上降低延迟并提升吞吐量。
Comments The codebase is available at https://github.com/mstar-project/mstar
生成式AI的计算安全性:假设检验视角
机构 * IBM Research(IBM研究院)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 本文从假设检验角度形式化生成式AI的计算安全性,提出基于信号处理的方法检测恶意输入和AI生成内容。
Comments Extended version of the paper presented at the ICML 2026 Workshop on Hypothesis Testing
逐层归因?可解释性中的元游戏
机构 * University of Warsaw(华沙大学) ; Centre for Credible AI, Warsaw University of Technology(可信AI中心,华沙技术大学) ; Bielefeld University(比勒菲尔德大学) ; LMU Munich, MCML(慕尼黑大学LMU,MCML)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 本文提出元游戏框架,通过将归因方法视为合作博弈计算Shapley值,研究模型解释的第二阶交互效应,揭示归因的层级分解及其在多领域可解释性应用中的价值。
PromptEvolver:通过自然语言空间中的进化优化实现提示倒置
机构 * Bar-Ilan University(巴伊兰大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文提出PromptEvolver,通过进化优化生成自然语言提示,实现高保真图像重建,优于现有方法。
协同AI代理与批评者用于网络遥测中的故障检测与原因分析
机构 * Department of Electrical and Computer Engineering, University of Alberta(阿尔伯塔大学电气与计算机工程系) ; SheQAI Research(SheQAI研究)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文提出一种多代理联邦系统,通过AI代理与批评者协作完成网络遥测中的故障检测、严重性和原因分析等多模态任务,利用多时间尺度随机逼近技术保证收敛性,通信开销低且隐私保护。
DAK-UCB: 为LLMs和生成模型的多样性感知提示路由
机构 * Sharif University of Technology(谢赫拉扎德技术大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文提出DAK-UCB方法,结合保真度与多样性指标,用于在线选择生成模型,以提升生成结果的多样性同时保持保真度。
Comments Accepted at ICLR 2026
解耦动作专家:将任务知识限制在条件路径中
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 本文提出解耦训练方法,将任务知识限制在条件路径中,使动作主干网络无任务依赖,通过Diffusion Policy验证,证明动作专家编码较少任务特定知识。
记忆打印机:通过结合缓慢设计与基于生成AI的图像创作探索日常回忆
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文提出Memory Printer,结合丝网印刷隐喻与文本到图像生成,通过分层重建、物理木刮刀和内置打印,探索慢互动如何重塑人机关系,发现记忆唤起与控制感提升等机遇及算法偏见等挑战。
Comments Accepted to CHI 2026
通过Cauchy-Schwarz散度进行分布式视觉语言对齐
机构 * University of Amsterdam(阿姆斯特丹大学) ; The Netherlands Cancer Institute(荷兰癌症研究所) ; Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) ; The Arctic University of Norway(挪威北极大学) ; Singapore Management University(新加坡管理大学)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文提出CS-Aligner框架,通过整合Cauchy-Schwarz散度与互信息,实现更紧密的视觉语言分布对齐,提升跨模态生成与检索性能。
Comments Accepted by ICLR2026
通过分位数分配生成模型
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 NeuroSQL通过隐式学习低维潜在表示,无需辅助网络,实现高效稳定的合成数据生成。
RMFlow:通过噪声注入步骤细化均流以实现多模态生成
机构 * Department of Mathematics and Scientific Computing and Imaging (SCI) Institute University of Utah(数学与科学计算及成像学院(SCI)院,犹他大学) ; Department of Mathematics, UCLA(数学系,加州大学洛杉矶分校)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 RMFlow通过引入噪声注入步骤,改进均流模型,实现高效多模态生成,仅需单次功能评估即可达到接近最先进的性能。
Comments Accepted to ICLR 2026
人工智能进展:面向创意产业的综述
机构 * Visual Information Laboratory, University of Bristol, Bristol, UK(布里斯托大学视觉信息实验室)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract)
AI总结 本文综述了自2022年以来人工智能在创意产业中的进展,探讨了生成式AI、大语言模型和扩散模型等技术对创意生产流程的影响,并分析了人类与AI协作的新趋势及面临的挑战。
Comments This is an updated review of our previous paper (see https://doi.org/10.1007/s10462-021-10039-7), and has been accepted by Artificial Intelligence Review journal
多语言到多模态(M2M):通过单语文本解锁新语言
机构 * Amazon(亚马逊)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 M2M通过单语文本学习多语言多模态对齐,实现多语言文本到图像检索的零样本迁移
Comments EACL 2026 Findings accepted. Camera-ready version
拓扑视角下的最优多模态嵌入空间
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文通过拓扑数据分析比较CLIP和CLOOB的嵌入空间,揭示其模态差距驱动因素和维度坍缩的影响,为多模态模型优化提供新视角。
Comments This manuscript contains substantive technical inaccuracies and an incomplete treatment of the stated topic. Subsequent developments and a reassessment of the problem indicate that the scope and framing of the work do not adequately reflect the current state of research, and the analysis is therefore incomplete and outdated
Beam-Brainstorm: 一种生成式特定地点波束成形方法
专题命中 文生图 :diffusion(abstract);image synthesis(abstract)
AI总结 本文提出了一种生成式特定地点波束成形方法,通过联合结构建模和自定义扩散模型,实现高效且高质量的用户特定波束生成。
通过链式推理和任务指令提示降低版权侵权风险
机构 * Munich RE(慕尼黑RE)
专题命中 文生图 :image generation(abstract);text-to-image(abstract)
AI总结 本文通过链式推理和任务指令提示结合负提示和提示重写,降低生成图像的版权侵权风险,并评估不同模型复杂度下的效果。