arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Southern California(南加州大学)

2026-05-25 至 2026-05-25 共收录 7
2605.23898 2026-05-25 cs.AI

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

SPACENUM: 重新审视视觉语言模型中的空间数值理解

Jianshu Zhang, Yijiang Li, Huifeixin Chen, Haoran Lu, Letian Xue, Bingyang Wang, Han Liu

机构 * Northwestern(西北大学) UCSD(加州大学圣地亚哥分校) USC(南加州大学) GaTech(佐治亚理工学院)

AI总结 提出SpaceNum框架,通过Num2Space和Space2Num双向任务评估VLM在动态过渡和静态布局中对空间数值的理解,发现模型依赖浅层空间线索,难以建立稳定的坐标感知表示。

Comments Project page: https://sterzhang.github.io/SpaceNum-Home

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18859 2026-05-25 cs.LG cs.AI

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

TwinRouterBench:面向现实智能体LLM路由的快速静态与实时动态评估

Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi

机构 * Gradient Soochow University(苏州大学) Independent Researcher(独立研究者) University of Southern California(南加州大学) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) Shanghai Jiao Tong University(上海交通大学) University of California, Berkeley(加州大学伯克利分校) University of the Chinese Academy of Sciences(中国科学院大学) University of California, Los Angeles(加州大学洛杉矶分校)

AI总结 提出TwinRouterBench基准,通过静态轨迹级前缀和动态智能体执行两轨,实现无需在线LLM评判的步骤级路由评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11490 2026-05-25 cs.LG stat.ML

Adaptive Calibration in Non-Stationary Environments

非平稳环境中的自适应校准

Junyan Liu, Haipeng Luo, Lillian J. Ratliff

机构 * University of Washington(华盛顿大学) University of Southern California(南加州大学)

AI总结 针对非平稳环境中的在线预测校准问题,提出基于epoch调度和非均匀预测空间划分的自适应算法,实现了校准误差在平稳与对抗场景之间的平滑插值,并达到最优或接近最优的遗憾界。

Comments Added results for piecewise-stationary environments and included a comparison with the concurrent work of Huang et al. (arXiv:2605.09273)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22880 2026-05-25 cs.CL cs.AI cs.CY

How Far Will They Go? Red-Teaming Online Influence with Large Language Models

它们会走多远?使用大型语言模型对在线影响力进行红队测试

Daniel C. Ruiz, Anna Serbina, Ashwin Rao, Emilio Ferrara, Luca Luceri

机构 * Information Sciences Institute University of Southern California(信息科学研究所 乌德穆尔特国立大学)

AI总结 本文提出一个红队测试框架,通过测量大型语言模型在争议话题上的政治观点表达范围(Overton Window),并量化简单自然语言越狱如何扩展该范围,发现开源LLM在政治表达上存在系统性不对称,且越狱效果因模型系列而异。

Comments 30 pages, 8 figures, submitted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01468 2026-05-25 cs.CL cs.LG

Remember what you did so you know what to do next

记住你做了什么,以便知道下一步该做什么

Manuel R. Ciosici, Alex Hedges, Yash Kankanampati, Justin Martin, Marjorie Freedman, Ralph Weischedel

机构 * Information Sciences Institute, University of Southern California(信息科学研究所,南加州大学)

AI总结 本研究使用中等规模的大语言模型GPT-J为模拟机器人在ScienceWorld文本游戏中的30类目标生成计划,通过填充更多历史步骤显著提升性能,并发现任务平均可能掩盖性能差异。

Comments Identical to EMNLP 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01552 2026-05-25 cs.CL cs.AI cs.LG

Perhaps PTLMs Should Go to School -- A Task to Assess Open Book and Closed Book QA

或许PTLMs应该去上学——一项评估开卷和闭卷问答的任务

Manuel R. Ciosici, Joe Cecil, Alex Hedges, Dong-Ho Lee, Marjorie Freedman, Ralph Weischedel

机构 * Information Sciences Institute, University of Southern California(信息科学研究所,南加州大学)

AI总结 提出一项基于大学教科书内容的新问答任务,通过闭卷和开卷测试评估预训练语言模型对教材的理解能力,发现模型在闭卷测试中表现接近随机,开卷测试略有提升。

Comments Identical to the EMNLP 2021 version

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05400 2026-05-25 cs.CL cs.AI cs.LG

Machine-Assisted Script Curation

机器辅助脚本编纂

Manuel R. Ciosici, Joseph Cummings, Mitchell DeHaven, Alex Hedges, Yash Kankanampati, Dong-Ho Lee, Ralph Weischedel, Marjorie Freedman

机构 * Information Sciences Institute, University of Southern California(信息科学研究所,南加州大学)

AI总结 提出MASC系统,通过自动化建议事件类型、链接维基数据和提醒遗漏子事件,实现人机协作脚本创作。

Comments Identical to the NAACL 2021 Demo version

详情

展开后加载摘要…

URL PDF HTML 收藏