arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-21 至 2026-04-21 共收录 12 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 12 篇

2604.17768 2026-04-21 cs.AI 85%

When Vision-Language Models Judge Without Seeing: Exposing Informativeness Bias

当视觉-语言模型评判而不看:揭示信息性偏差

Xiaohan Zou, Roshan Sridhar, Mohammadtaher Safarzadeh, Dan Roth

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Oracle AI

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.AI

AI总结 研究揭示视觉-语言模型在评判时存在的信息性偏差问题,提出BIRCH方法通过修正图像与答案的一致性提升评判可靠性,实验显示偏差降低17%,性能提升9.8%。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17375 2026-04-21 cs.CV cs.AI 81%

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models

当文本劫持视觉:基准测试与减轻文本叠加引起的视觉语言模型幻觉

Cui Yakun, Xingqun Qi, TianTian Geng, Yuyao Zhang, Sirui Han, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Birmingham(伯明翰大学)

专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出VisualTextTrap基准测试,通过大规模人工验证样本和定制评估指标,减轻文本叠加引起的幻觉问题,并提出VTHM-MoE框架以提升视觉-文本解耦能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17318 2026-04-21 cs.CV 79%

When Background Matters: Breaking Medical Vision Language Models by Transferable Attack

背景的重要性:通过可转移攻击打破医疗视觉语言模型

Akash Ghosh, Subhadip Baidya, Sriparna Saha, Xiuying Chen

机构 * Indian Institute of Technology Patna(印度理工学院帕纳分校) Indian Institute of Technology Kanpur(印度理工学院坎普尔分校) MBZUAI(穆桑大学人工智能研究所)

专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);分类 cs.CV

AI总结 本文提出MedFocusLeak攻击方法,通过在非诊断背景区域注入协调扰动并利用注意力分散机制,使模型产生看似合理但错误的诊断,揭示了现代临床VLMs推理能力的弱点。

Comments ACL Main 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18444 2026-04-21 cs.LG cs.AI cs.CV 67%

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification

ProtoCLIP:基于原型对齐的潜在细化用于鲁棒零样本胸部X光分类

Florian Kittler, Sheethal Bhat, Andreas Maier

机构 * Friedrich-Alexander University Erlangen-Nuremberg(埃朗根-纽伦堡弗里德里希-亚历山大大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 ProtoCLIP通过针对性数据筛选和锚点对齐优化,提升零样本胸部X光分类的鲁棒性,在VinDr-CXR数据集上提升AUC2-10个百分点,尤其在肺部气胸检测中达到0.94的SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16504 2026-04-21 cs.CV cs.LG 62%

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms

从手写到结构化数据:基准测试AI手写文档数字化

Nicholas Pather, Joshua Fouché, Sitwala Mundia, Karl-Günter Technau, Thokozile Malaba, Alex Welte, Ushma Mehta, Bruce A. Bassett

机构 * CSAM, University of the Witwatersrand(沃特沙尔德大学计算机科学与数学系) Grai Labs(Grai实验室) Faculty of Health Sciences, University of the Witwatersrand(沃特沙尔德大学健康科学学院) Empilweni Services and Research Unit, Department of Paediatrics and Child Health, University of the Witwatersrand(沃特沙尔德大学儿科与儿童健康部门Empilweni服务与研究单位) Division of Epidemiology and Biostatistics, School of Public Health, Faculty of Health Sciences, University of Cape Town(开普敦大学公共卫生学院流行病学与生物统计学系) Discipline of Public Health, School of Medicine, University of KwaZulu-Natal(夸祖鲁-纳塔尔大学医学院公共卫生系) WITS MIND Institute and CSAM, University of the Witwatersrand, University of Cape Town & Grai Labs(沃特沙尔德大学WITS MIND研究所和CSAM、开普敦大学及Grai实验室)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV、cs.LG

AI总结 本文评估了17种前沿多模态大语言模型在处理复杂手写表格任务中的表现,发现最新版Google和OpenAI模型在准确率和F1分数上达到85%和90%,并展示了提示优化对性能的显著提升。

Comments 19 Pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17873 2026-04-21 cs.CV 57%

Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

时空趋炎附势:基于否定的欺骗性在视频大语言模型中的表现

Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University Shanghai Key Laboratory of Multimodal Embodied AI(可信具身人工智能研究所,复旦大学上海多模态具身人工智能重点实验室)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.CV

AI总结 本文研究视频大语言模型在面对否定性欺骗时的脆弱性,揭示其在时空推理中的错误修正机制,并提出GasVideo-1000基准测试集以评估模型的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15690 2026-04-21 cs.AI stat.AP 57%

From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models

从被动度量到主动信号:不确定性量化在大语言模型中的演变角色

Jiaxin Zhang, Wendi Cui, Zhuohang Li, Lifu Huang, Bradley Malin, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce AI研究院) Intuit(Intuit公司) Vanderbilt University(范德比大学) University of California, Davis(加州大学戴维斯分校) Vanderbilt University Medical Center(范德比大学医学中心)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

AI总结 本文探讨大语言模型中不确定性量化从被动诊断指标到主动控制信号的演变,分析其在高级推理、自主代理和强化学习中的应用及贡献。

Comments This paper has been accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16884 2026-04-21 cs.CV 57%

Bias-constrained multimodal intelligence for equitable and reliable clinical AI

具有偏见约束的多模态智能用于公平且可靠的临床AI

Cheng Li, Weijian Huang, Jiarun Liu, Hao Yang, Qi Yang, Song Wu, Ye Li, Hairong Zheng, Shanshan Wang

机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(Paul C. Lauterbur生物医学成像研究中心) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Department of Radiology, Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院放射科) Department of Urology, South China Hospital, Medical School, Shenzhen University(深圳大学南方医院泌尿科) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 幻觉与鲁棒性 :visual question answering(abstract);分类 cs.CV

AI总结 本文提出BiasCareVL框架,通过在模型设计中直接引入偏见控制,解决医疗影像与文本整合中公平性和可靠性问题,展示了其在多任务和基准测试中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16606 2026-04-21 cs.CR cs.LG 57%

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

SafeLM:面向可信联邦大语言模型的统一隐私意识优化

Noor Islam S. Mohammad, Uluğ Bayazıt

机构 * Istanbul Technical University(伊斯坦布尔技术大学)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.LG

AI总结 SafeLM通过整合隐私、安全、虚假信息和对抗鲁棒性四大支柱,提升联邦大语言模型的可信度,实现98%有害内容检测准确率,降低通信开销并提升隐私保护效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16541 2026-04-21 cs.CV 57%

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

BOOKAGENT:通过多智能体认知校准 orchestrate 安全意识的视觉叙述

Bo Gao, Chang Liu, Yuyang Miao, Siyuan Ma, Ser-Nam Lim

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Science and Technology of China(中国科学技术大学) Imperial College London(伦敦帝国理工学院) Nanyang Technological University(南洋理工大学) University of Central Florida(佛罗里达中央大学)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.CV

AI总结 本文提出BOOKAGENT,一种安全意识的多智能体协作框架,用于高质量的视觉叙述生成,通过联合规划、脚本、插图和全局修复不一致,提升叙述连贯性和视觉一致性。

Comments 18 pages, Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04043 2026-04-21 cs.CL 50%

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

当助手变成危险:一个多模态大语言模型在日常生活中的安全分析基准

Xinyue Lou, Jinan Xu, Jingyi Yin, Xiaolong Wang, Zhaolu Kang, Youwei Liao, Yixuan Wang, Xiangyu Shi, Fengran Mo, Su Yao, Kaiyu Huang

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能在交通运输中的关键实验室(北京交通大学),教育部) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学) Tsinghua University(清华大学) Peking University(北京大学) University of Montreal(蒙特利尔大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

AI总结 本文提出SaLAD基准,通过2013个真实世界图像-文本样本评估多模态大语言模型在日常生活中的安全影响,揭示模型在识别危险行为方面的局限性。

Comments Accepted by ACL 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16374 2026-04-21 cs.CY 50%

Automating Sexual Injustice: Epistemic Injustice in Fembot Design and Feminist Directions for Equitable HRI

自动化性不公:仿生人设计中的知识不公与女性主义的公平人机交互方向

Surabhi Bhardwaj

专题命中 幻觉与鲁棒性 :grounding(abstract)

AI总结 本文分析仿生人设计中的知识不公问题,提出基于实证、知识多样性与主动同意的女性主义设计方向,旨在推动以知识正义为核心的公平人机交互。

Comments 5 pages, 1 figure. Peer-reviewed workshop paper presented at the Equitable Robotics for Wellbeing (EqRoW) Workshop at the ACM/IEEE International Conference on Human-Robot Interaction (HRI 2026), Edinburgh, UK, March 16, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏