arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-14 至 2026-04-14 共收录 12 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 12 篇

2602.21428 2026-04-14 cs.CV cs.LG 84%

PSF-Med: Measuring and Explaining Paraphrase Sensitivity in Medical Vision Language Models

PSF-Med: 医学视觉语言模型中反语敏感性的测量与解释

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven(纽黑文大学SAIL实验室)

专题命中 幻觉与鲁棒性 :vision language model(title,abstract);grounding(abstract);分类 cs.CV、cs.LG

AI总结 本文提出PSF-Med基准,通过26850个胸部X光问题及92856个意义保持的改写句,评估九种VLMs在临床问题改写下的稳定性与视觉依赖性,发现模型在反语下的翻转率差异显著,且通过特征分析揭示了模型决策机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11589 2026-04-14 cs.CV 83%

MLLM-as-a-Judge Exhibits Model Preference Bias

基于大语言模型的自动评估存在模型偏好偏差

Shuitsu Koyama, Yuiga Wada, Daichi Yashima, Komei Sugiura

机构 * Keio University(庆应义塾大学)

专题命中 幻觉与鲁棒性 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文研究了基于多模态大语言模型(MLLM)的自动评估方法中存在模型特定偏好偏差的问题,提出Philautia-Eval工具量化这种偏差,并通过实验发现模型间存在相互偏好偏差,最后引入Pomms模型组合方法有效缓解了偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03402 2026-04-14 cs.AI cs.LG 81%

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

风险感知注入:为安全校准视觉-语言模型而不牺牲实用性

Mengxuan Wang, Yuxin Chen, Gang Xu, Tao He, Hongjie Jiang, Ming Li

机构 * Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(华南理工大学吴贤铭智能工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳)) Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) University of Electronic Science and Technology of China(电子科技大学)

专题命中 幻觉与鲁棒性 :vision-language model(title);vision language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出RAI框架,通过增强不安全信号恢复视觉-语言模型的安全识别能力,同时保持语义完整性,实验显示有效降低攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11576 2026-04-14 cs.CV 79%

Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models

微调如你预训练:提升视觉语言模型零样本对抗鲁棒性

Songlong Xing, Weijie Wang, Zhengyu Zhao, Jindong Gu, Philip Torr, Nicu Sebe

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Xi’an Jiaotong University(西安交通大学) University of Oxford(牛津大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出AdvFLYP方法,通过遵循CLIP预训练过程的训练配方,利用网络上收集的图像-文本对生成对抗样本,并通过对比损失匹配文本,提升视觉语言模型的零样本对抗鲁棒性。

Comments Accepted to CVPR Findings Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05510 2026-04-14 cs.CV 79%

Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality

在增强现实中评估视觉-语言模型在矛盾虚拟内容攻击下的基准测试

Yanming Xiu, Zhengyuan Jiang, Neil Zhenqiang Gong, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出ContrAR基准,评估视觉-语言模型在增强现实中的鲁棒性,通过312个真实AR视频验证,测试11种VLMs,发现现有模型在检测对抗性内容操纵方面仍有改进空间。

Comments CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10299 2026-04-14 cs.CV cs.CL 79%

Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking

见善不为:通过对抗性注意力劫持使大视觉-语言模型失明以确保安全

Jingru Li, Wei Ren, Tianqing Zhu

机构 * China University of Geosciences, Wuhan(中国地质大学(武汉)) City University of Macau(澳门城市大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出Attention-Guided Visual Jailbreaking方法,通过操控注意力模式规避安全对齐机制,减少梯度冲突并提高攻击成功率,揭示了安全失明的失败模式。

Comments Accepted to ACL 2026. Code: https://github.com/Landsayy/AttentionJailbreak

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10071 2026-04-14 cs.CV 79%

Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation

聚光与阴影:基于注意力的双锚点反思解码用于MLLM幻觉缓解

Yebo Wu, Han Jin, Zhijiang Guo, Li Li

机构 * State Key Laboratory of IOTSC, University of Macau(澳门大学物联网国家重点实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 幻觉与鲁棒性 :MLLM(title);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出双锚点反思解码框架DaID,通过动态校准每个token生成来缓解MLLM幻觉,利用视觉注意力分布指导双锚点选择,提升推理能力。

Comments Accepted for Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09749 2026-04-14 cs.CV 79%

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment

看清真相:公平关注提升 groundedness 并减少 hallucination 在视觉语言对齐中

Mohammad Anas Azeez, Ankan Deria, Zohaib Hasan Siddiqui, Adinath Madhavrao Dukre, Rafiq Ali, Sara Atito, Yutong Xie, Imran Razzak

机构 * Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE(穆罕默德·本·扎耶德人工智能大学,阿布扎比,阿联酋) King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia(法赫德国王石油矿产大学,达兰,沙特阿拉伯) Macquarie University, Sydney, Australia(麦考瑞大学,悉尼,澳大利亚) University of Surrey, Guildford, United Kingdom(萨里大学,吉尔福德,英国) University of New South Wales, Sydney, Australia(新南威尔士大学,悉尼,澳大利亚)

专题命中 幻觉与鲁棒性 :grounding(title);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出DOP-OBC方法,通过公平分配注意力减少视觉语言对齐中的幻觉问题,提升生成的准确性与一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11144 2026-04-14 cs.CV cs.CL cs.MM 57%

Hierarchical Textual Knowledge for Enhanced Image Clustering

层次化文本知识用于增强图像聚类

Yijie Zhong, Yunfan Gao, Weipeng Jiang, Haofen Wang

机构 * Tongji University(同济大学) Huawei Technologies Ltd.(华为技术有限公司)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 本文提出KEC方法,通过大语言模型构建层次化概念-属性知识,提升图像聚类的准确性与鲁棒性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03765 2026-04-14 cs.CV 57%

ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs

ITIScore: 一种图像到文本到图像评分框架,用于评估多模态大语言模型的图像描述能力

Zitong Xu, Huiyu Duan, Shengyao Qin, Guangyu Yang, Guangji Ma, Xiongkuo Min, Ke Gu, Guangtao Zhai, Patrick Le Callet

机构 * Shanghai Jiao Tong University(上海交通大学) University of Electronic and Science Technology of China(电子科技大学) Beijing University of Technology(北京工业大学) Institut Universitaire de France (IUF), University of Nantes(法国大学研究院(IUF),南特大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出ICBenchmark图像描述基准,包含12类内容和2K图像生成的40K短长描述,提出ITIScore自动评估指标,通过图像到文本到图像框架衡量描述质量,实验显示其与人类判断一致且具有强泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11258 2026-04-14 cs.CL 50%

Dialectic-Med: Mitigating Diagnostic Hallucinations via Counterfactual Adversarial Multi-Agent Debate

辩证-调解:通过反事实对抗多智能体辩论缓解诊断幻觉

Zhixiang Lu, Jionglong Su

机构 * Xi’an Jiaotong-Liverpool University(西交利物浦大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

AI总结 本文提出Dialectic-Med框架,通过多智能体对抗辩论机制缓解医疗多模态大语言模型的诊断幻觉问题,提升推理可信度和解释忠实度。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05896 2026-04-14 cs.RO cs.HC 50%

Dialogue based Interactive Explanations for Safety Decisions in Human Robot Collaboration

基于对话的交互式解释用于人机协作中的安全决策

Yifan Xu, Xiao Zhan, Akilu Yunusa Kaltungo, Ming Shan Ng, Tsukasa Ishizawa, Kota Fujimoto, Clara Cheung

机构 * Department of Civil Engineering and Management, Faculty of Science and Engineering, The University of Manchester(曼彻斯特大学科学与工程学院土木工程与管理系) VRAIN, Universitat Politècnica de València(瓦伦西亚理工大学VRAIN研究所) Department of Engineering, University of Cambridge(剑桥大学工程系) Department of Mechanical and Aerospace Engineering, Faculty of Science and Engineering, The University of Manchester(曼彻斯特大学科学与工程学院机械与航空航天工程系) Center for the Possible Futures, Kyoto Institute of Technology(京都工艺纤维大学未来可能性中心) Institute of Industrial Science, The University of Tokyo(东京大学生产技术研究所) Graduate School of Frontier Sciences, The University of Tokyo(东京大学新领域创成科学研究科)

专题命中 幻觉与鲁棒性 :grounding(abstract)

AI总结 本文提出一种基于对话的交互式解释框架,用于人机协作中的安全决策,通过约束基础的安全评估实现解释与决策的紧密耦合,支持因果、对比和反事实查询,提升安全干预的透明度和协作效率。

Comments This paper has been accepted by the 2nd InterAI workshop, HRI conference 26'

详情

展开后加载摘要…

URL PDF HTML 收藏