CommentsAccepted for publication at the 42nd International Conference on Logic Programming (ICLP 2026). To appear in Theory and Practice of Logic Programming (TPLP)
CommentsSubmitted to Journal of Physics: Conference Series (Torque 2026). This is the Accepted Manuscript version of an article accepted for publication in Journal of Physics: Conference Series. IOP Publishing Ltd is not responsible for any errors or omissions in this version of the manuscript or any version derived from it. This Accepted Manuscript is published under a CC BY licence
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
福利、可改进性与方差:最优基准测试项聚合的主-代理方法
Andreas Haupt, Justin Hartenstein, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
机构
*
Department of Economics & Computer Science(经济与计算机科学系)
;
Institute for Computational and Mathematical Engineering(计算与数学工程研究所)
;
Department of Computer Science(计算机科学系)
;
Department of Aeronautics & Astronautics(航空与航天系)
CommentsAccepted to RLEval @ ACM CAIS 2026 (Workshop on Methods and RL Environments for Evaluating AI Agents) and selected for an invited talk based on reviewer ratings. 4-page short paper + appendix
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
DTBench:文档到表格提取的合成基准
Yuxiang Guo, Zhuoran Du, Nan Tang, Kezheng Tang, Congcong Ge, Yunjun Gao
机构
*
Zhejiang University(浙江大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science(香港科学与技术大学)
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
New York University(纽约大学)
;
Indiana University, Bloomington(印第安纳大学,布卢明顿)
;
Northeastern University(东北大学)
;
University College London(伦敦大学学院)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
语言模型智能体群体中的涌现语言:从令牌效率到监督规避
Stine Lyngsø Beltoft, William Brach, Federico Torrielli, Jacob Nielsen, Annemette Brok Pirchert, Filippo Tonini, Peter Schneider-Kamp, Lukas Galke Poech
机构
*
University of Southern Denmark(南丹麦大学)
;
Slovak University of Technology in Bratislava(布拉迪斯拉发技术大学)
;
University of Turin(都灵大学)
;
Ordbogen A/S(Ordbogen公司)
机构
*
X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系X-LANCE实验室)
;
MoE Key Lab of Artificial Intelligence(人工智能MOE重点实验室)
;
Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室)
;
AISpeech Ltd(AISpeech有限公司)
;
ETH Zürich(苏黎世联邦理工学院)
;
Nanjing University(南京大学)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
MosaicLeaks:深度研究代理的开放查询中的隐私风险
Alexander Gurung, Spandana Gella, Alexandre Drouin, Issam H. Laradji, Perouz Taslakian, Rafael Pardinas
机构
*
ServiceNow AI Research(ServiceNow AI研究院)
;
University of Edinburgh(爱丁堡大学)
;
Mila - Quebec AI Institute(魁北克AI研究所)
;
McGill University(麦吉尔大学)
;
University of British Columbia(不列颠哥伦比亚大学)
CommentsAccepted as a poster at the Foundation Models Meet Embodied Agents (FMEA) Workshop, CVPR 2026. 44 pages including appendix. Code: https://github.com/Avalon-S/PInVerify
The Surface You Test Is Not the Surface That Breaks
测试的表面并非断裂的表面
Shifat E Arman, Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin
机构
*
Department of Robotics and Mechatronics Engineering, University of Dhaka(达卡大学机器人与机电工程系)
;
Department of Computer Science and Engineering, University of Dhaka(达卡大学计算机科学与工程系)
专题命中
Agent评测
:agent(abstract);分类 cs.AI
AI总结
本文发现工具增强的LLM代理对提示注入的脆弱性依赖于攻击表面(工具输出 vs 工具描述),提出自适应攻击率并强调评估需报告每个表面的脆弱性。
Comments8 Figures, 8 Tables, Under Review at EMNLP
机构
*
Department of Electrical and Computer Engineering, Virginia Tech, Blacksburg, VA, USA(弗吉尼亚理工学院计算机工程系)
;
Center for Advanced AI, Accenture(Accenture高级人工智能中心)