arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Intel(英特尔)

2026-06-25 至 2026-06-25 共收录 2
2606.25760 2026-06-25 cs.LG cs.AI cs.CL cs.CV 新提交

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

计算机使用代理的不确定性量化:跨视觉语言模型和GUI基础数据集的基准测试

Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Intel Labs(英特尔实验室) Capital One AI Labs(Capital One AI实验室)

AI总结 提出跨机制基准Argus,评估27种后验不确定性量化方法在4个VLM代理和4个数据集上的表现,发现UQ排名在固定模型内跨数据集稳定,但跨模型类别和接口时下降,隐藏状态和密度方法最稳定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25954 2026-06-25 econ.TH cs.AI cs.LO math.CO math.LO 新提交

Measurable Majorities Are Not Finitely Axiomatizable

可测多数不是有限可公理化的

Lawrence S. Moss, Arthur Paul Pedersen

机构 * Dept. of Mathematics, Indiana University, Bloomington. Dept. of Computer Science \& the Intel Investigations Lab, the City College of New York the Graduate Center \& Remote Sensing Earth Systems Institute, the City University of New York.

AI总结 本文证明在有限社会决策框架中,严格多数推理的相干性准则无法被任何有界有限片段替代,从而否定了有限可公理化性。

详情

展开后加载摘要…

URL PDF HTML 收藏