arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86585 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2604.02406 2026-05-19 cs.CY 50%

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics

基于社区指导的评估文化文物AI生成图像

Nari Johnson, Deepthi Sudharsan, Hamna, Samantha Dalal, Theo Holroyd, Anja Thieme, Hoda Heidari, Daniela Massiceti, Jennifer Wortman Vaughan, Cecily Morrison

专题命中 文生图 :text-to-image(abstract)

AI总结 本文通过三个社区案例研究,探讨如何将社区经验融入AI生成图像的文化适当性评估体系,提出基于多模态LLM的自动化测量方法及挑战。

Comments Published at ACM FAccT 2026. 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14296 2026-05-18 cs.CY 50%

On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

生成基础模型的可信度:指南、评估与视角

Yue Huang, Chujie Gao, Siyuan Wu, Haoran Wang, Xiangqi Wang, Yujun Zhou, Yanbo Wang, Jiayi Ye, Jiawen Shi, Qihui Zhang, Yuan Li, Han Bao, Zhaoyi Liu, Tianrui Guan, Dongping Chen, Ruoxi Chen, Kehan Guo, Andy Zou, Bryan Hooi Kuen-Yew, Caiming Xiong, Elias Stengel-Eskin, Hongyang Zhang, Hongzhi Yin, Huan Zhang, Huaxiu Yao, Jaehong Yoon, Jieyu Zhang, Kai Shu, Kaijie Zhu, Ranjay Krishna, Swabha Swayamdipta, Taiwei Shi, Weijia Shi, Xiang Li, Yiwei Li, Yuexing Hao, Zhihao Jia, Zhize Li, Xiuying Chen, Zhengzhong Tu, Xiyang Hu, Tianyi Zhou, Jieyu Zhao, Lichao Sun, Furong Huang, Or Cohen Sasson, Prasanna Sattigeri, Anka Reuel, Max Lamparth, Yue Zhao, Nouha Dziri, Yu Su, Huan Sun, Heng Ji, Chaowei Xiao, Mohit Bansal, Nitesh V. Chawla, Jian Pei, Jianfeng Gao, Michael Backes, Philip S. Yu, Neil Zhenqiang Gong, Pin-Yu Chen, Bo Li, Dawn Song, Xiangliang Zhang

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出生成基础模型的可信度框架,通过多学科协作制定指导原则,引入动态评估平台TrustGen,揭示可信度进展与挑战,并探讨未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13113 2026-05-14 cs.CY cs.AI 50%

Context Matters: Auditing Gender Bias in T2I Generation through Risk-Tiered Use-Case Profiles

语境至关重要:通过风险分层用例配置审计T2I生成中的性别偏见

Jose Luna, Yankun Wu, Xiaofei Xie, Noa Garcia

机构 * Singapore Management University(新加坡国立大学) The University of Osaka(大阪大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出一个风险对齐的审计框架,用于评估T2I模型中的性别偏见,通过风险分层用例配置、指标目录和危害类型,系统化地审计性别偏见。

Comments FAccT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08280 2026-05-12 cs.LG cs.AI 50%

Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

超越虚假的权衡:适应性EWC用于隐蔽且可泛化的T2I后门

Lu Bowen, Xinyu Tang, Yin Yin Low, Shu-Min Leong

机构 * Monash University(墨尔本大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出Cosine-Aware Adaptive EWC,通过动态调整EWC正则化提升T2I后门攻击的ASR与模型保真度平衡,增强对域外数据集的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00699 2026-05-08 cs.CR 50%

STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack

STARE:分步时间对齐与红队引擎用于多模态毒性攻击

Xutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng, Hongying Zan, Cong Wang

专题命中 文生图 :image generation(abstract)

AI总结 STARE通过分步时间对齐和红队引擎,提升多模态毒性攻击的成功率,揭示优化诱导的相位对齐现象,为安全机制提供理论基础。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05414 2026-04-27 cs.CL cs.AI stat.ML 50%

Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions

大语言模型是糟糕的骰子玩家:LLM在生成统计分布的随机数时表现不佳

Minda Zhao, Yilun Du, Mengyu Wang

机构 * Harvard University(哈佛大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 研究发现大语言模型在生成随机数时存在显著缺陷,其采样能力随分布复杂度和采样范围增加而下降,导致下游任务出现系统性偏差。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26007 2026-04-17 cs.SD cs.AI cs.LG 50%

MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms

MARS:通过多通道自回归在频谱图上生成声音

Eleonora Ristori, Luca Bindini, Paolo Frasconi

机构 * AI Lab, DINFO Università di Firenze Florence, Italy(火nze大学人工智能实验室)

专题命中 文生图 :image synthesis(abstract)

AI总结 MARS通过多通道自回归在频谱图上生成声音,利用通道复用策略提升频谱图生成的效率和质量,实验表明其在多个评估指标上表现优异。

Comments Accepted at IJCNN 2026 (to appear in IEEE/IJCNN proceedings). This arXiv submission corresponds to the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09706 2026-04-14 quant-ph 50%

Hybrid Quantum-Classical Generative Adversarial Networks with Transfer Learning

混合量子-经典生成对抗网络与迁移学习

Asma Al-Othni, Saif Al-Kuwari, Mohammad Mahdi Nasiri Fatmehsari, Kamila Zaman, Ebrahim Ardeshir-Larijani

专题命中 文生图 :image synthesis(abstract)

AI总结 本文研究了结合迁移学习的混合量子-经典GAN架构,探讨了在生成器、判别器或两者中加入变分量子电路对性能的影响,发现全混合模型在图像质量和量化指标上表现更优。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08637 2026-04-10 cs.CY cs.AI cs.CR 50%

How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets

数据所有者如何说不?Web抓取视觉-语言AI训练数据集中的数据同意机制案例研究

Chung Peng Lee, Rachel Hong, Harry H. Jiang, Aster Plotnik, William Agnew, Jamie Morgenstern

机构 * Princeton University(普林斯顿大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学) University of Toronto(多伦多大学) Amazon AWS AI/ML(亚马逊AWS AI/ML)

专题命中 文生图 :text-to-image(abstract)

AI总结 研究探讨了数据所有者在AI训练数据集中的同意表达方式,分析了DataComp数据集中样本层面和网络层面的信息,揭示了当前数据收集流程对数据同意的不尊重问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12018 2026-03-24 cs.HC cs.AI cs.ET 50%

An Intent of Collaboration: On Agencies between Designers and Emerging (Intelligent) Technologies

协作的意图:设计者与新兴(智能)技术之间的机构

Pei-Ying Lin, Julie Heij, Iris Borst, Britt Joosten, Kristina Andersen, Wijnand IJsselsteijn

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 研究探讨设计者与智能技术协作中的权力动态,提出通过反思创作过程、理解技术特性及调整人机关系来恢复创作自主性。

Comments Accepted by IASDR Conference 2025, Taipei, Taiwan 16 pages excluding references, 8 figures

Journal ref Proceedings of IASDR 2025: Design Next

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15724 2026-03-18 cs.LG cs.AI 50%

Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models

Meta-TTRL:一种用于统一多模态模型中自改进测试时间强化学习的元认知框架

Lit Sin Tan, Junzhe Chen, Xiaolong Fu, Lichen Ma, Junshi Huang, Jianzhong Shi, Yan Li, Lijie Wen

机构 * Tsinghua University(清华大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 Meta-TTRL通过元认知信号实现测试时间参数优化,提升统一多模态模型的自改进能力,在文本到图像生成任务中取得显著成果。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07335 2026-03-10 cs.AI 50%

VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models

VisualScratchpad: 视觉语言模型推理时的视觉概念分析

Hyesu Lim, Jinho Choi, Taekyung Kim, Byeongho Heo, Jaegul Choo, Dongyoon Han

机构 * KAIST AI(韩国科学技术院人工智能研究所) NAVER AI Lab(NAVER人工智能实验室)

专题命中 文生图 :text-to-image(abstract)

AI总结 VisualScratchpad通过交互式界面分析视觉语言模型推理过程中的视觉概念,揭示三种未被充分探索的失败模式,帮助理解模型内部机制并改进调试。

Comments Accetped at ICLR 2026 Turstworthy AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04759 2026-02-19 eess.IV eess.SP 50%

Dilated POCS: Minimax Convex Optimization

扩张POCS:极小极大凸优化

Albert R. Yu, Robert J. Marks, Keith E. Schubert, Charles Baylis, Austin Egbert, Adam Goad, Sam Haug

专题命中 文生图 :image synthesis(abstract)

AI总结 本文提出扩张POCS方法,用于生成极小极大解,改进图像重建中的信号处理

Comments 8 pages, 12 figures

Journal ref IEEE Access, Volume 11, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15555 2026-02-16 eess.IV cs.AI 50%

Whole-Body Image-to-Image Translation for a Virtual Scanner in a Healthcare Digital Twin

全身体图像到图像翻译用于医疗数字孪生中的虚拟扫描仪

Valerio Guarrasi, Francesco Di Feola, Rebecca Restivo, Lorenzo Tronchin, Paolo Soda

机构 * Department of Computing Science, Umeå University(计算科学系,乌梅大学) Department of Diagnostics and Intervention, Biomedical Engineering and Radiation Physics, Umeå University(诊断与介入系,生物医学工程与放射物理,乌梅大学)

专题命中 文生图 :image synthesis(abstract)

AI总结 本文提出了一种基于区域分割的GAN方法,用于全身体CT到PET的图像翻译,通过分割身体为四个区域并使用区域特定的GANs提升翻译精度,从而在医疗数字孪生中实现更准确的虚拟PET扫描。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10023 2026-02-11 cs.CL 50%

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

MEVER:基于图的证据检索的多模态和可解释性声明验证

Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang, Dongwon Lee

机构 * University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) University of California San Diego(加州大学圣地亚哥分校) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 MEVER通过多模态图检索和解释生成,实现了准确且可解释的声明验证,同时创建了AI领域的科学数据集AIChartClaim。

Comments Accepted to EACL-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26861 2026-01-27 cs.IR cs.CL 50%

Evaluating Perspectival Biases in Cross-Modal Retrieval

评估跨模态检索中的视角偏差

Teerapol Saengsukhiran, Peerawat Chomphooyod, Narabodee Rodjananant, Chompakorn Chaksangchaichot, Patawee Prakrankamanant, Witthawin Sripheanpol, Pak Lovichit, Sarana Nutanong, Ekapol Chuangsuwanich

机构 * Department of Computer Engineering, Chulalongkorn University(楚莱隆坤大学计算机工程系) School of Information Science and Technology, VISTEC(信息科学与技术学院,VISTEC)

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出3XCM基准,揭示跨模态检索中语言和文化偏见的影响,强调需通过解耦策略实现公平检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22931 2025-12-10 cs.CY 50%

Visual Orientalism in the AI Era: From West-East Binaries to English-Language Centrism

人工智能时代视觉沙文主义:从东西二元对立到英语中心主义

Zhilong Zhao, Yindi Liu

专题命中 文生图 :text-to-image(abstract)

AI总结 本文探讨了人工智能时代视觉沙文主义的演变,揭示AI通过视觉表示强化东西方二元对立,转向英语中心主义,强调需重新审视算法治理与训练数据中的地缘政治结构。

Comments 43 pages, 8 figures, supplementary materials included

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15061 2025-12-05 cs.AI 50%

Empowering Clients -- Transformation of Design Processes Due to Generative AI

赋能客户——生成式人工智能对设计过程的转变

Johannes Schneider, Kilic Sinem, Daniel Stockhammer

机构 * University of Liechtenstein(利希滕斯坦大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 研究探讨生成式AI如何改变建筑设计过程,发现AI使客户能更快参与设计,但可能阻碍创新,同时引发建筑师角色转变和文化同质化风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21547 2025-11-27 cs.HC 50%

Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies

双倍观察:如何通过并列的T2I比较改变审计策略

Matheus Kunzler Maldaner, Wesley Hanwen Deng, Jason I. Hong, Kenneth Holstein, Motahhare Eslami

专题命中 文生图 :text-to-image(abstract)

AI总结 MIRAGE工具通过并列展示多个文本到图像模型,帮助用户发现生成模型的偏见,改变审计策略。

Comments 8 pages, 6 figures. Presented at ACM Collective Intelligence (CI), 2025. Available at https://ci.acm.org/2025/wp-content/uploads/101-Maldaner.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03263 2025-11-26 cs.LG cs.AI 50%

Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models

记忆自我再生:在未学习模型中揭示隐藏知识

Agnieszka Polowczyk, Alicja Polowczyk, Joanna Waczyńska, Piotr Borycki, Przemysław Spurek

机构 * Silesian University of Technology(谢勒斯大学技术学院) Jagiellonian University(雅盖隆大学) IDEAS Research Institute(IDEAS研究所)

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出记忆自我再生任务和MemoRa策略,旨在解决模型遗忘难题,通过增强知识检索的鲁棒性提升反学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17358 2025-11-24 cs.CL 50%

Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding

不要学习,而是依托:自然语言推理与视觉依托的案例

Daniil Ignatev, Ayman Santeer, Albert Gatt, Denis Paperno

机构 * Utrecht University(乌特勒支大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出了一种基于视觉依托的零样本自然语言推理方法,通过生成视觉表示并比较与假设的相似度,实现高精度推理,展示了对文本偏见的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16038 2025-11-21 cs.HC 50%

Panel-by-Panel Souls: A Performative Workflow for Expressive Faces in AI-Assisted Manga Creation

面板式灵魂:一种用于AI辅助漫画创作中表现力面部的表演性工作流程

Qing Zhang, Jing Huang, Yifei Huang, Jun Rekimoto

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出了一种交互式工作流程,用于AI辅助漫画创作中表现力面部的生成,通过双混合流程实现艺术家意图与AI执行的高效衔接。

Comments NeurIPS 2025 Creative AI Track, The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10268 2025-11-14 cs.AI 50%

Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention

Zhe Xu, Zhicai Wang, Junkang Wu, Jinda Lu, Xiang Wang

专题命中 文生图 :text-to-image(abstract)

Comments accepted for publication in the Association for the Advancement of Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12170 2025-11-03 cs.NE 50%

Language Model Crossover: Variation through Few-Shot Prompting

Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K. Hoover, Joel Lehman

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02631 2025-10-31 cs.HC cs.AI 50%

Reflection on Data Storytelling Tools in the Generative AI Era from the Human-AI Collaboration Perspective

Haotian Li, Yun Wang, Huamin Qu

机构 * Microsoft Research Asia(微软亚洲研究院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 文生图 :text-to-image(abstract)

Comments This paper is a sequel to the CHI 24 paper "Where Are We So Far? Understanding Data Storytelling Tools from the Perspective of Human-AI Collaboration (https://doi.org/10.1145/3613904.3642726), aiming to refresh our understanding with the latest advancements. It is accepted at IEEE VIS 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15551 2025-10-27 cs.LG 50%

PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors

Sepehr Dehdashtian, Mashrur M. Morshed, Jacob H. Seidman, Gaurav Bharaj, Vishnu Naresh Boddeti

机构 * Michigan State University(密歇根州立大学) Reality Defender

专题命中 文生图 :text-to-image(abstract)

Comments Accepted as NeurIPS 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18450 2025-09-30 cs.CL 50%

BRIT: Bidirectional Retrieval over Unified Image-Text Graph

Ainulla Khan, Yamada Moyuru, Srinidhi Akella

机构 * Fujitsu Research India(富士通印度研究)

专题命中 文生图 :text-to-image(abstract)

Comments Accepted in EMNLP-2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21874 2025-09-29 cs.LG 50%

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

Yifei Peng, Yaoli Liu, Enbo Xia, Yu Jin, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15432 2025-09-22 cs.IR 50%

SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models

Thong Nguyen, Yibin Lei, Jia-Huei Ju, Andrew Yates

专题命中 文生图 :text-to-image(abstract)

Comments Accepted

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15154 2025-09-09 cs.AI cs.LG 50%

Online Prompt Pricing based on Combinatorial Multi-Armed Bandit and Hierarchical Stackelberg Game

Meiling Li, Hongrun Ren, Haixu Xiong, Zhenxing Qian, Xinpeng Zhang

专题命中 文生图 :text-to-image(abstract)

Comments The paper has been withdrawn by the authors because the current experimental results are not sufficiently reliable. Further optimization and refinement of the methodology are required before the work can be disseminated

详情

展开后加载摘要…

URL PDF HTML 收藏