Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
福利、可改进性与方差:最优基准测试项聚合的主-代理方法
Andreas Haupt, Justin Hartenstein, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
机构
*
Department of Economics & Computer Science(经济与计算机科学系)
;
Institute for Computational and Mathematical Engineering(计算与数学工程研究所)
;
Department of Computer Science(计算机科学系)
;
Department of Aeronautics & Astronautics(航空与航天系)
CommentsAccepted to RLEval @ ACM CAIS 2026 (Workshop on Methods and RL Environments for Evaluating AI Agents) and selected for an invited talk based on reviewer ratings. 4-page short paper + appendix
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
DTBench:文档到表格提取的合成基准
Yuxiang Guo, Zhuoran Du, Nan Tang, Kezheng Tang, Congcong Ge, Yunjun Gao
机构
*
Zhejiang University(浙江大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science(香港科学与技术大学)
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
New York University(纽约大学)
;
Indiana University, Bloomington(印第安纳大学,布卢明顿)
;
Northeastern University(东北大学)
;
University College London(伦敦大学学院)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
语言模型智能体群体中的涌现语言:从令牌效率到监督规避
Stine Lyngsø Beltoft, William Brach, Federico Torrielli, Jacob Nielsen, Annemette Brok Pirchert, Filippo Tonini, Peter Schneider-Kamp, Lukas Galke Poech
机构
*
University of Southern Denmark(南丹麦大学)
;
Slovak University of Technology in Bratislava(布拉迪斯拉发技术大学)
;
University of Turin(都灵大学)
;
Ordbogen A/S(Ordbogen公司)
机构
*
X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系X-LANCE实验室)
;
MoE Key Lab of Artificial Intelligence(人工智能MOE重点实验室)
;
Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室)
;
AISpeech Ltd(AISpeech有限公司)
;
ETH Zürich(苏黎世联邦理工学院)
;
Nanjing University(南京大学)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
MosaicLeaks:深度研究代理的开放查询中的隐私风险
Alexander Gurung, Spandana Gella, Alexandre Drouin, Issam H. Laradji, Perouz Taslakian, Rafael Pardinas
机构
*
ServiceNow AI Research(ServiceNow AI研究院)
;
University of Edinburgh(爱丁堡大学)
;
Mila - Quebec AI Institute(魁北克AI研究所)
;
McGill University(麦吉尔大学)
;
University of British Columbia(不列颠哥伦比亚大学)
CommentsAccepted as a poster at the Foundation Models Meet Embodied Agents (FMEA) Workshop, CVPR 2026. 44 pages including appendix. Code: https://github.com/Avalon-S/PInVerify
The Surface You Test Is Not the Surface That Breaks
测试的表面并非断裂的表面
Shifat E Arman, Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin
机构
*
Department of Robotics and Mechatronics Engineering, University of Dhaka(达卡大学机器人与机电工程系)
;
Department of Computer Science and Engineering, University of Dhaka(达卡大学计算机科学与工程系)
专题命中
Agent评测
:agent(abstract);分类 cs.AI
AI总结
本文发现工具增强的LLM代理对提示注入的脆弱性依赖于攻击表面(工具输出 vs 工具描述),提出自适应攻击率并强调评估需报告每个表面的脆弱性。
Comments8 Figures, 8 Tables, Under Review at EMNLP
机构
*
Department of Electrical and Computer Engineering, Virginia Tech, Blacksburg, VA, USA(弗吉尼亚理工学院计算机工程系)
;
Center for Advanced AI, Accenture(Accenture高级人工智能中心)
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
迈向人性化的社交媒体生态系统:面向安全、自主与福祉的AI增强人机交互设计模式
Mohd Ruhul Ameen, Akif Islam
机构
*
College of Engineering(工程学院)
;
Computer Sciences Marshall University Huntington, WV, USA(计算机科学马歇尔大学亨廷顿州威斯康星州)
;
Department of Computer Science(计算机科学系)
;
Engineering University of Rajshahi Rajshahi 6205, Bangladesh(工程 Rajshahi 大学 Rajshahi 6205 巴基斯坦)
Comments6 pages, 5 tables, 7 figures, and 2 algorithm tables. Accepted at International Conference on Signal Processing, Information, Communication and Systems (SPICSCON 2025)
Journal ref2025 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON)
LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification
LegSegNet:用于下肢CT组织分割与量化的公共深度学习系统
Yuwen Chen, Yaqian Chen, Roy Colglazier, Haoyu Dong, Hanxue Gu, Maciej A. Mazurowski, Kevin W. Southerland
机构
*
Department of Electrical and Computer Engineering, Duke University(杜克大学电气与计算机工程系)
;
Department of Biostatistics & Bioinformatics, Duke University(杜克大学生物统计与生物信息学系)
;
Department of Radiology, Duke University(杜克大学放射学系)
;
Department of Computer Science, Duke University(杜克大学计算机科学系)
;
Department of Surgery, Duke University(杜克大学外科系)
SAW-Bench: Learning Situated Awareness in the Real World
SAW-Bench:在现实世界中学习情境感知
Chuhan Li, Rilyn Han, Joy Hsu, Yongyuan Liang, Rajiv Dhawan, Jiajun Wu, Ming-Hsuan Yang, Xin Eric Wang
机构
*
University of California, Santa Barbara(加州大学圣芭芭拉分校)
;
Yale University(耶鲁大学)
;
Stanford University(斯坦福大学)
;
University of Maryland, College Park(马里兰大学学院市分校)
;
Amazon(亚马逊)
;
University of California, Merced(加州大学默塞德分校)