arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Southern California(南加州大学)

2026-06-15 至 2026-06-15 共收录 3
2606.14240 2026-06-15 cs.AI 新提交

AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties

AFFORDANCE20Q:从物理属性评估可承担性推理

Yifan Jiang, Meige Yang, Zitong Li, Jay Pujara

机构 * Information Sciences Institute, University of Southern California(南加州大学信息科学研究所) University of Southern California(南加州大学)

AI总结 提出Affordance20Q基准,通过20个问题游戏评估模型从物理属性推理物体可承担性的能力,发现LLM与人类差距约20分,并开发KARI方法提升开源模型达15.2分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13934 2026-06-15 cs.AI 新提交

Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

对抗性概念搜索:从特征几何预测组合错误

Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee, David Alvarez-Melis, Ellie Pavlick, Naomi Saphra

机构 * Brown University(布朗大学) University of Southern California(南加州大学) Harvard University(哈佛大学) Boston University(波士顿大学)

AI总结 利用LLM的表征几何预测其组合失败模式,发现概念编码近正交时可靠组合,编码接近时因干扰导致失败,无需评估具体输入即可预测错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12430 2026-06-15 cs.CY cs.AI 新提交

Will AI Agents Free Us From Meaningless Work? A Human-Centered Analysis

AI代理能否让我们摆脱无意义的工作?一项以人为中心的分析

Davide Ghia, Jaspreet Ranjit, Tania Cerquitelli, Daniele Quercia

机构 * Politecnico di Torino(都灵理工大学) University of Southern California(南加州大学) Nokia Bell Labs(诺基亚贝尔实验室)

AI总结 基于Graeber的“狗屁工作”理论,通过任务级分析发现,工人感知的任务无意义程度强烈预测其对AI委托的意愿,且此类任务被认为需要较少人工监督。

Comments Improved overall writing; add details about task filtering and participants screening; add comments in the discussion about the subjective and context-specific nature of the scale introduced;

详情

展开后加载摘要…

URL PDF HTML 收藏