arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-03-12 至 2026-03-12 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 8 篇

2603.10360 2026-03-12 cs.CV 70%

One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination

一个token,两种命运:通过视觉token操控构建统一框架以对抗大语言模型幻觉

Zhan Fa, Yue Duan, Jian Zhang, Lei Qi, Yinghuan Shi

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室) School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院)

专题命中 幻觉与鲁棒性 :LLaVA(abstract);MLLM(abstract);分类 cs.CV

AI总结 通过视觉token操控构建统一框架,有效减少大语言模型幻觉,提升POPE准确性2%。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10436 2026-03-12 cs.RO cs.DC 67%

COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints

COHORT: 基于混合强化学习的多机器人系统实时协同大规模深度神经网络推理

Mohammad Saeid Anwar, Anuradha Ravi, Indrajeet Ghosh, Gaurav Shinde, Carl Busart, Nirmalya Roy

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract)

AI总结 COHORT通过混合强化学习策略,优化多机器人系统在实时约束下的大规模DNN推理,降低能耗并提升GPU利用率。

Comments Recently accepted at 27th IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks ( IEEE WoWMoM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10473 2026-03-12 cs.CL cs.AI 57%

Aligning Large Language Models with Searcher Preferences

将大型语言模型对齐于搜索者偏好

Wei Wu, Peilun Zhou, Liyi Chen, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Hui Xiong

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院) Xiaohongshu Inc.(小红书公司) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

AI总结 SearchLLM通过分层多维奖励系统提升开放式生成搜索的鲁棒性和用户需求对齐能力,实测有效消费率提升1.03%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10470 2026-03-12 cs.CV 57%

Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression

用反事实对抗幻觉:基于扩散引导的扰动用于LVLM幻觉抑制

Hamidreza Dastmalchi, Aijun An, Ali Cheraghian, Hamed Barzamini

机构 * York University(约克大学) Macquarie University(麦考瑞大学) Northern Illinois University(北伊利诺伊大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 CIPHER通过反事实图像扰动减少LVLM中的视觉幻觉,提升模型忠实性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18867 2026-03-12 cs.CV 57%

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

相似性作为证据:校准过度自信的视觉语言模型以实现可解释且标注高效的医学主动学习

Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi, Shaohua Hong, Shuo Li

机构 * School of Electronic Technology and Engineering, Xiamen University, Xiamen, China(厦门大学电子技术与工程学院,厦门,中国) School of Informatics, Xiamen University, Xiamen, China(厦门大学信息学院,厦门,中国) School of Engineering, Case Western Reserve University, Cleveland, USA(凯斯西储大学工程学院,克利夫兰,美国)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 SaE框架通过引入相似性证据头校准视觉语言模型,以提高医学主动学习的可解释性和标注效率。

Comments Accepted to CVPR 2026 (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04268 2026-03-12 cs.CV 57%

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

KVSmooth: 通过键值平滑缓解多模态大语言模型中的幻觉

Siyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun He

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

AI总结 KVSmooth通过键值平滑技术有效缓解多模态大语言模型中的幻觉问题,提升生成精度和召回率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01957 2026-03-12 cs.CV 57%

AFTER: Mitigating the Object Hallucination of LVLM via Adaptive Factual-Guided Activation Editing

AFTER: 通过自适应事实引导激活编辑缓解LVLM中的对象幻觉

Tianbo Wang, Yuqing Ma, Kewei Liao, Zhange Zhang, Simin Li, Jinyang Guo, Xianglong Liu

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 AFTER通过自适应事实引导激活编辑缓解LVLM中的对象幻觉,显著降低幻觉发生率。

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11099 2026-03-12 cs.CV 57%

A Survey on Interpretability in Visual Recognition

视觉识别中可解释性的综述

Qiyang Wan, Chengzhi Gao, Ruiping Wang, Xilin Chen

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

AI总结 本文综述了视觉识别中可解释性的发展,从意图、对象、呈现和方法学角度建立多维分类法,总结了关键评估指标,并探讨了多模态大语言模型的可解释性及实际应用。

Comments 20 pages, 8 figures, 7 tables. Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏