arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-08-18 至 2026-08-18 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 9 篇

2608.16690 2026-08-18 cs.CV 新提交 91%

AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty

AnchorScore:一种基于CLIP的多模态大语言模型(MLLM)标注难度诊断方法

Yan Ma, Lizhuo Zhang

机构 * School of Foreign Studies, Changsha University of Science and Technology(长沙理工大学外国语学院) School of Education, Hunan Agricultural University(湖南农业大学教育学院) School of Information and Intelligence, Hunan Agricultural University(湖南农业大学信息与智能科学学院)

专题命中 幻觉与鲁棒性 :MLLM(title,title_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出基于CLIP的AnchorScore,可低成本预先对类别按MLLM标注难度排名,在课堂行为、动作识别等数据上与MLLM准确率相关性高,可用于混合路由、提示消歧等场景。

Comments 37 pages, 7 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16081 2026-08-18 cs.CV 新提交 85%

SafeGesture: Evaluating Fine-Grained Hand Gesture Understanding in Vision-Language Models through Scenario-Conditioned Safety Interpretation

SafeGesture:通过场景条件安全解释评估视觉语言模型的细粒度手势理解能力

Taegang Kim, Saleh Afroogh, Junfeng Jiao

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Urban Information Lab, The University of Texas at Austin(德克萨斯大学奥斯汀分校城市信息实验室)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);LLaVA(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出SafeGesture基准,评估5款视觉语言模型的细粒度手势安全理解能力,发现模型存在感知与推理脱节,瓶颈为场景条件安全推理而非手势识别。

Comments 14 pages, 22 tables, 2 figures. Code and benchmark resources available at this https URL (https://github.com/The-Responsible-AI-Initiative/SafeGesture)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18094 2026-08-18 cs.CV cs.AI cs.DB 版本更新 76%

OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models

OODBench: 用于大型视觉-语言模型的分布外基准

Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen

机构 * University of Science and Technology of China Suzhou Institute for Advanced Research, USTC Key Laboratory of the Ministry of Education for Mathematical Foundations Unmanned System Research Institute, Northwestern Polytechnical University

专题命中 幻觉与鲁棒性 :vision-language model(title);分类 cs.CV、cs.AI

AI总结 OODBench提出了一种自动化方法,用于构建评估大型视觉-语言模型处理分布外数据能力的基准,并展示了当前模型在面对此类数据时的性能下降。

Comments 54 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14854 2026-08-18 cs.CV 新提交 70%

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

Zero-MELO:基于多模态大语言模型的测试时证据校准用于零样本微手势识别

Chengyan Wang, Hanliang Xie, Yueyi Yang, Haoyu Chen

机构 * University of Oulu(奥卢大学) Peking University(北京大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract,abstract_cn);分类 cs.CV

AI总结 该研究针对多模态大语言模型在微手势识别中局部证据不足、分数偏差的瓶颈,提出Zero-MELO框架,结合树搜索、测试时校准与多线索融合,在iMiGUE和MA-52数据集上显著优于Qwen2.5-VL基线。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04484 2026-08-18 cs.CV 版本更新 70%

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

TrustCLIP:通过对抗性重建学习隐私视觉特征

Nikos Athanasiou, Ilya A. Petrov, Angela Yao, Shugao Ma, Eric Sauser, Edoardo Remelli, Shreyas Hampali, Johannes Schönberger, Fadime Sener, Bugra Tekin

机构 * Meta Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) University of Tübingen(图宾根大学) National University of Singapore(新加坡国立大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 研究视觉和视觉语言模型中视觉特征隐私问题,提出TrustCLIP框架,以特征条件生成器为隐私对手,优化编码器特征与下游模块投影,降低生成式反演保真度并保持下游任务性能。

Comments this https URL (https://atnikos.github.io/trustclip/) Update affiliations

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01503 2026-08-18 cs.CV 版本更新 70%

Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

通过深度排序任务解构视觉语言模型中的图像线索理解与语言偏差

Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos

机构 * York University(约克大学) University of Guelph(圭尔夫大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV

AI总结 提出深度排序与异常检测任务,结合控制图像线索和语言表达,量化视觉语言模型的深度感知能力,发现模型对深度线索利用不足且存在语言偏差。

Comments 15 pages, 7 figures, accepted to ECCV 2026 (30 pages, 13 figures, supplementary materials included)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16663 2026-08-18 cs.DB cs.AI 新提交 57%

Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL

用于可靠企业级文本到SQL的有界语义规划与确定性编译

Yi Ai

专题命中 幻觉与鲁棒性 :grounding(abstract_cn);分类 cs.AI

AI总结 提出语义路径编译(SPC)系统,在ACME保险基准测试中,其文本到SQL任务正确率达97.4%,显著优于基线系统,且鲁棒性更强。

Comments 10 sections, 2 figures, 6 tables. Preprint. Code and research artifacts are described in the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15115 2026-08-18 cs.CV 新提交 57%

Perspective-Invariant Attack with Enhanced Transferability of Adversarial Examples

具有增强对抗样本迁移性的视角不变攻击

Kaisheng Liang, Yiming Cao, Bin Xiao

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

AI总结 针对对抗样本跨模型迁移性带来的安全威胁,提出视角不变攻击(PIA)及其扩展PIA-Mix,通过多自由度顶点采样策略提升对抗样本迁移性,实验显示其性能优于当前最优基于迁移的攻击方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11603 2026-08-18 cs.CL cs.AI 版本更新 57%

DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning

DR.GAP:采用解耦推理的性别感知提示缓解大语言模型中的偏见

Hongye Qiu, Yue Xu, Yi Wang, Meikang Qiu, Wenjie Wang

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

AI总结 该研究针对大语言模型的性别偏见问题,提出DR.GAP方法,通过生成无性别推理轨迹作为上下文示例,在不修改模型参数的情况下缓解偏见,且可扩展至视觉语言模型,经多任务多模型实验验证有效。

详情

展开后加载摘要…

URL PDF HTML 收藏