arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 2242 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 2242 篇

2507.09279 2025-10-14 cs.CV cs.AI cs.CL 86%

Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models

Anita Kriz, Elizabeth Laura Janes, Xing Shen, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克AI研究所)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);visual question answering(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025 Workshop CVAMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00827 2025-09-26 cs.CV cs.AI 86%

IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves

Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Huawei Technologies Ltd.(华为技术有限公司) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00373 2025-09-03 cs.CV cs.AI 86%

Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models

Sihao Wu, Gaojie Jin, Wei Huang, Jianhong Wang, Xiaowei Huang

机构 * University of Liverpool(利物浦大学) University of Exeter(埃克塞特大学) Purple Mountain Laboratories(紫金山实验室) University of Bristol(布里斯托大学)

专题命中 幻觉与鲁棒性 :vision language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08982 2025-07-15 eess.IV cs.CV cs.LG 86%

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models

Hanene F. Z. Brachemi Meftah, Wassim Hamidouche, Sid Ahmed Fezza, Olivier Déforges

机构 * Univ. Rennes, INSA Rennes, CNRS, IETR - UMR 6164(里昂大学、里昂国家理工学院、国家科学研究中心、IETR - UMR 6164)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05163 2025-07-08 cs.CV cs.LG 86%

Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models

Aishwarya Venkataramanan, Paul Bodesheim, Joachim Denzler

机构 * Computer Vision Group, Friedrich Schiller University Jena(计算机视觉组,费迪里奇·施勒尔大学耶纳)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV、cs.LG

Comments UAI 2025, 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03788 2025-05-08 cs.CL cs.AI cs.CV 86%

Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding

Trilok Padhi, Ramneet Kaur, Adam D. Cobb, Manoj Acharya, Anirban Roy, Colin Samplawski, Brian Matejek, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha

机构 * Georgia State University(佐治亚州立大学) Computer Science Lab, SRI(SRI计算机科学实验室) Army Cyber Institute, United States Military Academy(美国陆军网络学院)

专题命中 幻觉与鲁棒性 :grounding(title,abstract);LLaVA(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14003 2024-03-22 cs.CV cs.CL cs.LG 86%

Multi-Modal Hallucination Control by Visual Information Grounding

Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto

专题命中 幻觉与鲁棒性 :grounding(title);vision-language model(abstract);VLM(abstract);LLaVA(abstract)

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07447 2026-05-11 cs.CV cs.AI cs.CL cs.LG 86%

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

稀疏自编码器作为视觉语言模型中对抗攻击检测的即插即用防火墙

Hao Wang, Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh, Daisuke Kawahara

机构 * Magellan Technology Research Institute (MTRI)(马杰伦技术研究 institute) Waseda University(早稻田大学)

专题命中 幻觉与鲁棒性 :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出基于稀疏自编码器的轻量级对抗攻击检测框架SAEgis,通过插入预训练VLM中的稀疏自编码模块,利用学习到的稀疏潜在特征检测对抗扰动输入,实验显示其在跨领域和跨攻击设置中表现优异,且无需额外对抗训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07525 2026-08-11 cs.CL cs.AI 新提交 85%

Unified Hallucination Fuzzing for Multimodal Large Language Models

面向多模态大语言模型的统一幻觉模糊测试

Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);grounding(abstract);分类 cs.AI

AI总结 针对多模态大语言模型的幻觉问题,提出含UniHall基准与SAMF自演化模糊测试的评估框架,发现SOTA模型在模糊测试下性能显著下降,存在有用性-幻觉权衡。

Comments 47 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21619 2026-07-27 cs.CL cs.AI 新提交 85%

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

对抗风格优化:通过基于GRPO的风格触发优化增强VLM越狱

Bingjun Luo, Jialin Guo, Yue Yao, Xinpeng Ding

机构 * Tsinghua University(清华大学) Harbin Engineering University(哈尔滨工程大学) Shandong University(山东大学) Xidian University(西安电子科技大学)

专题命中 幻觉与鲁棒性 :VLM(title,title_cn);multimodal large language model(abstract);分类 cs.AI

AI总结 研究MLLMs安全对齐易受越狱攻击问题,提出基于GRPO的对抗风格优化(ASO)方法,通过优化风格触发增强视觉越狱,实验证明该方法显著提高攻击成功率,凸显风格偏差对MLLMs红队测试的作用。

Comments Accepted by CVPR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31976 2026-07-01 cs.AI cs.MA 新提交 85%

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models

TreeAgent: 一种基于编译专家规则和视觉语言模型的通用多智能体框架,用于林业自动化偏差标注

Shiyi Chen, Nicholas Saban, Collin Hargreaves, Huiqi Wang

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出TreeAgent多智能体框架,结合专家决策树与视觉语言模型,通过解耦声明式决策实现零修改泛化,在树木偏差分类中超越监督学习基线并降低标注成本。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05017 2026-06-15 cs.CV cs.CL 版本更新 85%

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

通过细化文本嵌入缓解大型视觉语言模型中的幻觉

Aakriti Agrawal, Gouthaman KV, Rohith Aralikatti, Gauri Jagatap, Jiaxin Yuan, Sarvesh Baskar, Vijay Kamarshi, Andrea Fanelli, Furong Huang

机构 * University of Maryland(马里兰大学) Dolby Laboratories(杜比实验室) Capital One

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 针对大型视觉语言模型因过度依赖文本先验而忽视视觉线索导致的幻觉问题,提出一种简单有效的视觉特征融入方法,通过学习视觉信息化的文本嵌入来平衡注意力分布,显著降低幻觉并提升多模态推理能力。

Comments Accepted at The 64th Annual Meeting of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23853 2026-05-29 cs.AI cs.MA 85%

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

SCoOP: 多视觉-语言模型系统中用于不确定性量化的语义一致意见池化

Chung-En Johnny Yu, Brian Jalaian, Nathaniel D. Bastian

机构 * University of West Florida(西佛罗里达大学) United States Military Academy(美国军事学院)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出SCoOP框架,通过不确定性加权的线性意见池化聚合多个视觉-语言模型的输出,实现无训练的不确定性量化,有效检测幻觉并支持高不确定性样本的弃权。

Comments Accepted to ICLR 2026 Workshop on Agentic AI in the Wild: From Hallucinations to Reliable Autonomy

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28609 2026-05-28 cs.CV 85%

JECA^2: Judgment-Explanation Consistent Adversarial Attack against Forensic Vision-Language Models

JECA^2: 面向取证视觉语言模型的判断-解释一致对抗攻击

Jiachen Qian

机构 * City University of Hong Kong(香港城市大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 针对取证视觉语言模型,提出一种白盒对抗攻击方法JECA^2,通过Grad-CAM引导的视觉扰动和令牌邻近约束的文本嵌入优化,实现判断与解释的一致性,实验表明攻击成功率和一致性优于基线。

Comments 37 pages, 6 figures. Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26992 2026-05-27 cs.CV 85%

On the Robustness of Machine Unlearning for Vision-Language Models

机器遗忘在视觉-语言模型中的鲁棒性研究

Yujie Lin, Kaidi Jia, Jiayao Ma, Chengyi Yang, Jinsong Su

机构 * Xiamen University(厦门大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文首次系统调查了视觉-语言模型机器遗忘的鲁棒性,通过提出三种攻击范式揭示现有方法往往隐藏而非彻底移除目标知识。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16127 2026-05-18 cs.CV 85%

WeatherOcc3D: VLM-Assisted Adverse Weather Aware 3D Semantic Occupancy Prediction

WeatherOcc3D: 借助VLM的恶劣天气感知3D语义占用预测

A. Enes Doruk, Abdelaziz Hussein, Hasan F. Ates

机构 * Department of Artificial Intelligence(人工智能系) Data Engineering Ozyegin University Istanbul, Türkiye(数据工程奥祖根大学伊斯坦布尔,土耳其)

专题命中 幻觉与鲁棒性 :VLM(title,title_cn);分类 cs.CV

AI总结 本文提出一种借助预训练CLIP隐空间的框架,通过语言环境线索指导多传感器融合,解决恶劣天气下传感器可靠性问题,提升3D语义占用预测的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12258 2026-05-13 cs.LG 85%

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

指令透镜分数:您的指令为多模态大语言模型提供了一个强大的对象幻觉检测器

Runhe Lai, Xinhua Lu, Yanqi Wu, Jinlun Ye, Weijiang Yu, Ruixuan Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China(中山大学计算机科学与工程学院,广州,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) Key Laboratory of Machine Intelligence and Advanced Computing, MOE, Guangzhou, China(机器智能与高级计算关键实验室,教育部,广州,中国)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.LG

AI总结 本文提出InsLen,通过结合校准局部分数和上下文一致性分数,有效检测多模态大语言模型中的对象幻觉,无需额外训练或辅助模型。

Comments Accepted by ICML-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08841 2026-05-12 cs.CV 85%

Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models

基于 illusion 意识的视觉预处理与反 illusion 提示的经典 illusion 理解在视觉-语言模型中的应用

Junli Zha, Jiahui Wang, Xinkai Lu, Jinbo Wang

机构 * SF Technology Co., Ltd.(SF技术有限公司)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出无需训练的框架,通过图像预处理、反 illusion 提示工程和多票集成策略,解决视觉-语言模型对视觉错觉的感知与记忆冲突,实现90.48%和98.41%的准确率。

Comments Accepted at CVPR 2026 Workshop on 5th DataCV Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25884 2026-04-29 quant-ph cs.CV 85%

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

QCalEval:用于量子校准图理解的视觉-语言模型基准测试

Shuxiang Cao, Zijian Zhang, Abhishek Agarwal, Grace Bratrud, Niyaz R. Beysengulov, Daniel C. Cole, Alejandro Gómez Frieiro, Elena O. Glen, Hao Hsu, Gang Huang, Raymond Jow, Greshma Shaji, Tom Lubowe, Ligeng Zhu, Luis Mantilla Calderón, Nicola Pancotti, Joel Pendleton, Brandon Severin, Charles Etienne Staub, Sara Sussman, Antti Vepsäläinen, Neel Rajeshbhai Vora, Yilun Xu, Varinia Bernales, Daniel Bowring, Elica Kyoseva, Ivan Rungger, Giulia Semeghini, Sam Stanwyck, Timothy Costa, Alán Aspuru-Guzik, Krysta Svore

机构 * NVIDIA University of Toronto(多伦多大学) IQM Quantum Computers(IQM量子计算机) Lawrence Berkeley National Laboratory(伯克利国家实验室) Conductor Quantum(Conductor量子) National Physical Laboratory(国家物理实验室) Infleqtion Harvard University(哈佛大学) Fermi National Accelerator Laboratory(费米国家加速器实验室) Northwestern University(西北大学) EeroQ Corporation(EeroQ公司) Royal Holloway University of London(伦敦皇家霍洛威大学) Vector Institute for Artificial Intelligence(人工智能向量研究所)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出QCalEval,首个用于评估视觉-语言模型理解量子校准图能力的基准测试,包含243个样本和87种场景类型,测试零样本和上下文学习下的六种问题类型,展示了不同模型的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18275 2026-04-22 cs.CV 85%

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

面向自动驾驶的视觉对抗攻击研究

Tianyuan Zhang, Lu Wang, Xinwei Zhang, Yitong Zhang, Boyi Jia, Siyuan Liang, Shengshan Hu, Qiang Fu, Aishan Liu, Xianglong Liu

机构 * Beihang University(北航) National University of Singapore(国立新加坡大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出ADvLM框架,针对自动驾驶中视觉语言模型的特殊需求,解决文本指令变异性和视觉场景时间序列性问题,实现高效对抗攻击。

Comments Accepted by Machine Intelligence Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17768 2026-04-21 cs.AI 85%

When Vision-Language Models Judge Without Seeing: Exposing Informativeness Bias

当视觉-语言模型评判而不看:揭示信息性偏差

Xiaohan Zou, Roshan Sridhar, Mohammadtaher Safarzadeh, Dan Roth

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Oracle AI

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.AI

AI总结 研究揭示视觉-语言模型在评判时存在的信息性偏差问题,提出BIRCH方法通过修正图像与答案的一致性提升评判可靠性,实验显示偏差降低17%,性能提升9.8%。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07222 2026-04-17 cs.LG cs.CL 85%

Pay Less Attention to Function Words for Free Robustness of Vision-Language Models

减少对功能词的关注以实现视觉语言模型的自由鲁棒性

Qiwei Tian, Chenhao Lin, Zhengyu Zhao, Chao Shen

机构 * School of Cyber Science and Engineering, Xi’an Jiaotong University, Xi’an 710049, China(西安交通大学计算机科学与工程学院,西安 710049,中国)

专题命中 幻觉与鲁棒性 :vision-language model(title);VLM(abstract,abstract_cn);grounding(abstract);分类 cs.LG

AI总结 本文提出FDA方法,通过减少功能词的注意力影响,提升视觉语言模型在跨模态对抗攻击下的鲁棒性,实验显示在检索任务中ASR下降18%/13%/53%,性能损失极小。

Comments The paper has been accepted by ICLR26

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01010 2026-04-02 cs.CV cs.MM 85%

PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks

PDA:用于对抗图像攻击的文本增强防御框架

Jingning Xu, Haochen Luo, Chen Liu

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV

AI总结 PDA通过测试时文本增强提升视觉语言模型对多种对抗攻击的鲁棒性,无需修改底层模型,平衡了鲁棒性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20985 2026-03-24 cs.CV 85%

Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models

一致却危险:每样本安全分类揭示医疗视觉-语言模型中的虚假可靠性

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven(SAIL实验室,新罕布什尔大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV

AI总结 研究揭示医疗VLMs中一致性指标的缺陷,通过四象限分类发现危险样本高准确率且低熵,建议部署评估需结合文本基线以识别虚假可靠性。

Comments CVPR 2026 Workshop on Medical Reasoning with Vision Language Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16960 2026-03-24 cs.CR cs.AI 85%

Adversarial attacks against Modern Vision-Language Models

对抗现代视觉-语言模型的攻击

Alejandro Paredes La Torre

机构 * Duke University(杜克大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.AI

AI总结 研究开放源码视觉-语言模型在模拟真实预部署环境中的对抗鲁棒性,评估两种模型在三种梯度攻击下的表现,发现其存在显著的安全差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19983 2026-02-26 cs.RO cs.AI 85%

Contextual Safety Reasoning and Grounding for Open-World Robots

面向开放世界机器人的上下文安全推理与 grounding

Zachary Ravichandran, David Snyder, Alexander Robey, Hamed Hassani, Vijay Kumar, George J. Pappas

机构 * University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 幻觉与鲁棒性 :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.AI

AI总结 CORE框架通过视觉语言模型实现在线上下文推理与空间grounding,提供概率安全保证,在开放世界中有效执行上下文适应的安全行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26441 2026-01-27 cs.CV 85%

A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models

A-TPT:面向视觉语言模型测试时提示微调的角多样性校准特性

Shihab Aaqil Ahamed, Udaya S. K. P. Miriya Thanthrige, Ranga Rodrigo, Muhammad Haris Khan

机构 * Dept. of Electronic and Telecommunication Engineering, University of Moratuwa(摩图瓦大学电子与电信工程系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 A-TPT通过引入角度多样性提升视觉语言模型测试时提示微调的校准性能,有效减少校准误差并提升适应能力。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17394 2026-01-08 cs.CL cs.CV cs.CY 85%

Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?

视觉语言模型是否是跨文化理论思维推理者?

Zabir Al Nazi, GM Shahariar, Md. Abrar Hossain, Wei Peng

专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);VLM(abstract);MLLM(abstract)

AI总结 本文提出CulturalToM-VQA基准测试集,评估视觉语言模型在跨文化理论思维推理中的表现,发现前沿模型在准确性上显著提升,但存在假信念推理和区域差异等局限,揭示模型的社会可取性偏差问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12520 2025-10-28 cs.CV 85%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16805 2025-09-23 cs.CV 85%

Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models

Md. Atabuzzaman, Ali Asgarov, Chris Thomas

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);visual reasoning(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏