arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-07-03 至 2026-07-03 共收录 12 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 12 篇

2607.00483 2026-07-03 cs.RO 新提交 92%

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning

VLM-AR3L:用于强化学习中绝对和相对奖励的视觉-语言模型

Kuan-Chen Chen, Winston Chen, Wei-Fang Sun, Min-Chun Hu

机构 * National Tsing Hua University(国立清华大学) NVIDIA AI Technology Center (NVAITC)(英伟达AI技术中心)

专题命中 幻觉与鲁棒性 :VLM(title,title_cn);vision-language model(title,abstract)

AI总结 提出VLM-AR3L框架,利用视觉-语言模型从偏好标签中学习绝对和相对奖励,结合状态评估与比较监督,在多种基准任务中优于先前方法。

Comments Accepted at IJCAI 2026. Project website: https://vlm-ar3l.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01973 2026-07-03 cs.CV cs.AI cs.LG 新提交 89%

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

评估视觉语言模型在图像退化与偏差下进行医学图像质量评估的可靠性

Sofiane Ouaari, Kevin Vorwalder, Nico Pfeifer

机构 * University of Tuebingen(图宾根大学) Institute for Bioinformatics and Medical Informatics (IBMI), University of Tuebingen(图宾根大学生物信息学与医学信息学研究所)

专题命中 幻觉与鲁棒性 :VLM(title,abstract_cn);InternVL(abstract,abstract_cn);vision-language model(abstract);visual question answering(abstract)

AI总结 研究视觉语言模型在医学图像质量评估中,面对图像退化(如像素化)和文本属性偏差时的可靠性,发现像素化显著降低性能,而文本属性(如机构声望)引入偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18315 2026-07-03 cs.RO cs.AI cs.CV 版本更新 89%

DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving

DriveVLM-RL:受神经科学启发的强化学习与视觉语言模型用于安全可部署的自动驾驶

Zilin Huang, Zihao Sheng, Zhengyang Wan, Yansong Qu, Junwei You, Sicong Jiang, Sikai Chen

机构 * Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系) Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建设工程学院) Department of Civil Engineering, McGill University(麦吉尔大学土木工程系)

专题命中 幻觉与鲁棒性 :VLM(summary_cn,abstract);vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 提出DriveVLM-RL框架,通过双通路架构将VLM集成到RL中,实现安全可部署的自动驾驶,在CARLA中显著优于基线。

Comments 33 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02089 2026-07-03 cs.CV cs.AI cs.CL cs.LG cs.MM 新提交 88%

ESC: Emotional Self-Correction for Reliable Vision-Language Models

ESC:面向可靠视觉语言模型的情感自我纠正

Tien-Huy Nguyen, Minh-Nhat Nguyen, Nguyen Nhat Huy, Hung Viet Nguyen, Huy Nguyen Minh Nhat, Thanh-Huy Nguyen, Cuong Tuan Nguyen, Hoang M. Le, Dat Nguyen, Phat Kim Huynh, Min Xu, Ulas Bagci

机构 * 1 GenAI4E Lab 2 University of Information Technology, Ho Chi Minh City, Vietnam 3 Universit\"at Trier, Germany 4 Ho Chi Minh University of Technology, Ho Chi Minh City, Vietnam 5 PAMI Lab, Vietnamese German University, Vietnam 6 Vietnam National University, Ho Chi Minh City, Vietnam 7 Carnegie Mellon University, USA 8 Omoshiroi AI, USA 9 Harvard University, USA 10 Basis Research Institute 11 PASSIO Laboratory, North Carolina A\&T State University, USA 12 Mohamed bin Zayed University of Artificial Intelligence, UAE 13 Northwestern University, USA [4pt] Equal contribution. Corresponding author

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI、cs.LG

AI总结 提出训练无关的ESC框架,通过外部验证器检测错误初始响应并注入情感反馈,促使VLM自我纠正,提升安全、幻觉、感知和多模态推理等任务的可靠性。

Comments ECCV Main Track 2026 (113 pages, 15 tables, 65 figures). Project Page: https://genai4e.github.io/ESC/?

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01897 2026-07-03 cs.LG cs.AI 新提交 79%

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

先排序后行动:基于帧序进展的无奖励控制

Yuriy Maksyuta, George Bredis, Ruslan Rakhimov, Daniil Gavrilov

机构 * T-Tech

专题命中 幻觉与鲁棒性 :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 提出Rank-Then-Act框架,利用视觉语言模型从专家视频中学习进度排序器,通过斯皮尔曼等级相关奖励函数实现无环境奖励的策略学习,在离散和连续控制任务中达到或超越现有方法。

Comments 20 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01365 2026-07-03 cs.LG cs.AI cs.CV 新提交 75%

Multi-modal Rail Crossing Safety Analysis

多模态铁路道口安全分析

Paimon Goulart, Chansong Lim, Nícolas Roque dos Santos, Yue Dong, Sheldon Peterson, Jia Chen, Evangelos E. Papalexakis

机构 * University of California, Riverside(加州大学河滨分校) Riverside County Transportation Commission(河滨县交通委员会)

专题命中 幻觉与鲁棒性 :VLM(abstract,abstract_cn);分类 cs.CV、cs.AI、cs.LG

AI总结 提出多模态管道,利用视觉线索和事故报告评估铁路道口安全,基于FRA评分实现高风险/低风险分类(F1=0.757)和分数估计(RMSE=0.078)。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02494 2026-07-03 cs.CV cs.CL 新提交 70%

Towards Robustness against Typographic Attack with Training-free Concept Localization

基于训练无关的概念定位实现对抗印刷体攻击的鲁棒性

Bohan Liu, Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 幻觉与鲁棒性 :vision language model(abstract);visual question answering(abstract);分类 cs.CV

AI总结 提出一种无需训练的可解释性方法,通过分析注意力头对词汇和语义的编码差异,定位并干预ViT中的词汇偏置电路,从而在不额外训练的情况下显著提升对印刷体攻击的鲁棒性。

Comments 15 pages main text, provisionally accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02284 2026-07-03 cs.CV 新提交 70%

FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval

FlowCIR: 通过流匹配实现零样本组合图像检索的语义传输

Zhenqi He, Ziqi Jiang, Yuanpei Liu, Yanghao Wang, Teng Wang, Long Chen

机构 * The Hong Kong University of Science and Technology, Hong Kong SAR(香港科技大学) The University of Hong Kong, Hong Kong SAR(香港大学)

专题命中 幻觉与鲁棒性 :VLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出FlowCIR,利用条件流匹配将参考图像与指令组合成目标对齐查询嵌入,避免文本反转的信息损失,训练效率高,并引入多负向引导策略提升对否定指令的鲁棒性。

Comments Accept to ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01657 2026-07-03 cs.CV 新提交 57%

Domain Generalization via Text-Anchored Information Bottleneck

基于文本锚定信息瓶颈的域泛化

Eunyi Lyou, Yunjeong Choi, Junho Lee, Joonseok Lee

机构 * Seoul National University(首尔大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 提出通过语言嵌入空间作为信息瓶颈来抑制视觉表征中的虚假线索,从而提升域泛化性能,实验表明该方法达到最新最优水平。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01457 2026-07-03 cs.CL cs.AI 新提交 57%

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

接地优化:一种减少自动个人文档重写中LLM幻觉的分层工程框架

Shashank Indukuri, Adarsh Agrawal

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

AI总结 提出五层接地优化框架,通过时间验证、污染检测、结构不变性、提示接地和评估器,将简历重写中的幻觉率降至0.04-0.24。

Comments 13 pages, 1 figure. Equal contribution by both authors. Code and data: https://github.com/shashank-indukuri/grounded-optimization

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02329 2026-07-03 cs.AI cond-mat.mtrl-sci physics.comp-ph 新提交 57%

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

基于语料库的前沿计算物理容错LLM流水线:从语料到论文的自主研究

Haonan Huang

机构 * Princeton University(普林斯顿大学)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

AI总结 提出一个端到端LLM流水线,从11,083篇arXiv论文语料库出发,自主完成前沿计算物理研究,包括构思方向、复现文献、第一性原理计算和撰写论文,通过冗余设计实现容错。

Comments 39 pages, 5 figures. Accepted at the ICML 2026 AI for Science Workshop (https://openreview.net/forum?id=R5YXaPgUAx). Includes the pipeline-generated companion physics manuscript as an appendix. Data and scaffolding archive: https://doi.org/10.5281/zenodo.21126996

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07820 2026-07-03 cs.CL 50%

Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests

参考游戏作为模型不确定性与澄清请求对齐的测试平台

Manar Ali, Judith Sieker, Sina Zarrieß, Hendrik Buschmeier

机构 * Digital Linguistics Lab(数字语言实验室) Computational Linguistics Group(计算语言学小组)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

AI总结 本文通过参考游戏测试语言模型在不确定性识别与澄清请求表达上的能力,发现模型在简单任务中难以准确识别自身不确定性并转化为澄清行为。

Comments Accepted at GEM@ACL 2026, the 5th Generation, Evaluation & Metrics Workshop

Journal ref Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM 2026), pp. 990-998

详情

展开后加载摘要…

URL PDF HTML 收藏