Copyright-Aware Incentive Scheme for Generative Art Models Using Hierarchical Reinforcement Learning
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments 9 pages, 9 figures
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments 9 pages, 9 figures
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR、cs.MM
Comments Tech report
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments ICML 2024
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
Journal ref Foundations and Trends in Privacy and Security 6 (2023) 1-52
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments 11 pages, 9 figures
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
Comments Technical report. Project page at https://minidalle3.github.io/
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
专题命中 文生图 :text-to-image(abstract);diffusion(abstract);image synthesis(abstract)
Comments Accepted for "Cultures in AI/AI in Culture" NeurIPS 2022 Workshop
专题命中 文生图 :image generation(abstract);text-to-image(abstract);inpainting(abstract)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract)
前沿大语言模型能否媲美原生多模态嵌入?在难负例文本到图像检索上的对比
机构 * Westcliff University(韦克利夫大学)
专题命中 文生图 :text-to-image(title);分类 cs.CV
AI总结 本研究在Flickr30k数据集上对比Gemini Embedding 2等原生多模态嵌入与GPT-4.1、Claude Sonnet 4.6等前沿LLM的难负例文本到图像检索性能,发现二者表现相当,且多模态嵌入更适配低延迟应用
从基线到随访:利用因果层次变分自编码器在UK Biobank中生成脊柱DXA图像
机构 * School of Electronics and Computer Science(电子与计算机科学学院) ; University of Southampton(萨塞克斯大学) ; MRC Lifecourse Epidemiology Centre(英国医学研究理事会生命周期流行病学中心) ; University of Southampton, Southampton General Hospital(萨塞克斯大学索马塞特医院) ; Computer Science University of Southampton(计算机科学萨塞克斯大学)
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 本文提出了一种基于元数据的因果层次变分自编码器,用于在UK Biobank中生成一致的脊柱DXA图像,通过基线到随访的设置评估因果一致性,展示了年龄干预下关键椎体形态学变量的高一致性,支持了在解剖上合理的DXA图像合成。
Comments 7 pages, 4 figures, 3 tables. Accepted at the 48th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2026)
Lens:重新思考基础文本到图像模型的训练效率
机构 * Microsoft Lens Team(微软Lens团队)
专题命中 文生图 :text-to-image(title);分类 cs.CV
AI总结 本文提出Lens,一个具有38亿参数的文本到图像模型,在多种基准测试中表现与超过60亿参数的最新模型相当甚至更优,同时训练计算需求显著降低。通过最大化训练批次的数据信息密度和改进收敛速度的架构选择,实现了高效的训练和优化。
Comments Project Page: https://github.com/microsoft/Lens
PureCC: 文本到图像概念定制的纯学习
机构 * Tsinghua University(清华大学) ; School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) ; Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(深圳大学智能信息处理广东省重点实验室) ; Kling Team, Kuaishou Technology(快手科技Kling团队) ; University of Exeter(埃克塞特大学) ; S-Lab, Nanyang Technological University(南洋理工大学S实验室)
专题命中 文生图 :text-to-image(title);分类 cs.CV
AI总结 本文提出PureCC,一种用于文本到图像概念定制的纯学习方法,通过分离学习目标来平衡概念定制的保真度与模型保留。
Comments Accepted to CVPR 2026
利用3D多对比自注意GAN进行脑部MRI图像合成
机构 * School of Computing and Mathematical Sciences, University of Leicester(利兹大学计算与数学科学学院)
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 本文提出3D-MC-SAGAN框架,通过单张T2w图像生成高保真多对比MRI图像,保留肿瘤特征,结合多种损失函数提升全局真实感和肿瘤结构保真度。
Comments Note: This work has been submitted to the IEEE for possible publication
文本到自动机图:比较TikZ代码生成与直接图像合成
机构 * Marshall University(马歇尔大学) ; West Virginia State University(西弗吉尼亚州立大学)
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 本研究比较了通过视觉-语言模型生成TikZ代码与直接图像合成在生成自动机图描述中的效果,发现人工修正能显著提高生成准确性。
Comments Accepted to ASEE North Central Section 2026
VisionDirector: 生成图像合成中的视觉-语言引导闭环细化
机构 * The Hong Kong University of Science and Technology(香港理工大学) ; The Chinese University of Hong Kong(香港中文大学) ; Huawei Research(华为研究)
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 VisionDirector通过视觉-语言引导闭环细化方法,提升生成图像合成中多目标任务的完成度和质量。
由双阶段生成对抗网络驱动的肺结节图像合成
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 本文提出双阶段生成对抗网络TSGAN,通过解耦肺结节形态结构与纹理特征,提升合成数据多样性及检测模型性能。
从易到难++:通过空间-频率课程促进差分隐私图像合成
机构 * University of Virginia(弗吉尼亚大学) ; Microsoft Research(微软研究院)
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 FETA-Pro通过引入频率特征作为训练捷径,提升差分隐私图像合成的保真度和实用性。
Comments Accepted at Usenix Security 2026; code available at https://github.com/2019ChenGong/Feta-Pro
通过生成式人工智能图像合成促进AI皮肤病变分类器的公平性评估
机构 * DFKI GmbH(德累斯顿理工大学人工智能研究所)
专题命中 文生图 :image synthesis(title);分类 cs.CV
AI总结 本文通过生成式人工智能图像合成方法,评估AI皮肤病变分类器的公平性,验证合成图像在公平性测试中的有效性。
零shot分层植物分割:通过基础分割模型和文本到图像注意力
机构 * The University of Osaka(大阪大学) ; Phytometrics ; Nagoya University(名古屋大学)
专题命中 文生图 :text-to-image(title);分类 cs.CV
AI总结 ZeroPlantSeg通过基础分割模型和文本到图像注意力实现零shot分层植物分割,有效提取植物个体,优于现有方法。
Comments WACV 2026 Accepted
略多一些像这个:利用视觉语言模型进行文本到图像检索的相关反馈
机构 * Maastricht University(马斯特里赫特大学)
专题命中 文生图 :text-to-image(title);分类 cs.CV
AI总结 本文提出基于相关反馈机制提升视觉语言模型文本到图像检索性能的方法,通过四种反馈策略在Flickr30k和COCO数据集上验证,提升了检索效果并增强了多轮检索的鲁棒性。
Comments Accepted to WACV'26
专题命中 文生图 :text-to-image(title);分类 cs.MM
机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) ; Department of Electrical Engineering and Computer Science, Case Western Reserve University(电气工程与计算机科学系,凯斯西储大学)
专题命中 文生图 :image synthesis(title);分类 cs.CV
机构 * School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院) ; National Engineering Research Center of Advanced Magnetic Resonance Technologies for Diagnosis and Therapy (NERC-AMRT)(先进磁共振诊断与治疗技术国家工程研究中心) ; Med-X Research Institute, Shanghai Jiao Tong University(医-X研究院) ; Kashi Prefecture Second People’s Hospital(喀什县第二人民医院) ; Renji Hospital, Shanghai, China(仁济医院) ; Shanghai Sixth People’s Hospital, Shanghai, China(上海第六人民医院)
专题命中 文生图 :image synthesis(title);分类 cs.CV
Journal ref IEEE Journal of Biomedical and Health Informatics, 2025
机构 * STORM Lab UK, School of Electronic and Electrical Engineering, University of Leeds(STORM实验室(英国)、电子与电气工程学院、利兹大学)
专题命中 文生图 :image synthesis(title);分类 cs.CV
Comments 13 pages, 9 figures
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; Shanghai Innovation Institute(上海创新研究院) ; Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院) ; Department of Computer Science, The University of Hong Kong(香港大学计算机科学系)
专题命中 文生图 :image synthesis(title);分类 cs.CV
Comments iccv 2025, camera-ready version
机构 * Hanyang University(翰阳大学) ; Hanyang University College of Medicine(翰阳大学医学院) ; Chonnam National University Medical School(全南国立大学医学院) ; Keimyung University School of Medicine(庆尚大学医学院) ; Yonsei Wonju College of Medicine(延世Wonju医学院) ; Yonsei University College of Medicine(延世大学医学院)
专题命中 文生图 :image synthesis(title);分类 cs.CV
Comments Single column, 28 pages, 7 figures