Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents
现在询问,以后使用:评估长期 LLM 代理中的主动性差距
Bin Wu, Guanyun Zou, Bingbing Wang, Huan Zhao, Chuan Shi
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Noumena AI
机构
*
University of New South Wales, NSW, Sydney, Australia(新南威尔士大学,新州,悉尼,澳大利亚)
;
Suzhou Institute for Advanced Research, University of Science(苏州先进研究院,科学大学)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
Cornell University(康奈尔大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.CL、cs.AI
AI总结
提出AtomWorld基准,通过十种基本原子结构操作评估LLM在材料科学中的空间推理能力,发现Claude Opus 4.6表现最佳但复杂空间关系操作成功率低,表明LLM更适合作为辅助工具而非完全自主的科研代理。
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
测量基于LLM的简历筛选中真实世界的提示注入攻击
Mohan Zhang, Yuqi Jia, Zhen Tan, Steven Jiang, Neil Zhenqiang Gong, Tianlong Chen, Dawn Song
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Duke University(杜克大学)
;
Arizona State University(亚利桑那州立大学)
;
hireEZ
;
University of California, Berkeley(加州大学伯克利分校)
Comments16 pages, 7 figures, 12 tables. Accepted to the ICML 2026 Workshop on Hypothesis Testing, Seoul, South Korea, 2026. Copyright 2026 by the author(s)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
基于LLM的自动评分中可学习的评估技能:通过迭代优化构建评分标准
Yun Wang, Xin Xia, Xuansheng Wu, Xiaoming Zhai, Ninghao Liu
机构
*
School of Computing, University of Georgia, Athens, GA, USA(佐治亚大学计算机学院)
;
AI4STEM Education Center, University of Georgia, Athens, GA, USA(佐治亚大学AI4STEM教育中心)
;
The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学)
机构
*
AI Lab, Princeton University Engineering and Applied Sciences(普林斯顿大学人工智能实验室、工程与应用科学学院)
;
RAND Corporation(RAND公司)
;
Engineering and Applied Sciences(工程与应用科学)
;
Science, Technology, and International Affairs(科学、技术与国际事务)
;
Georgetown University(乔治·华盛顿大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
机构
*
Instituto Superior Técnico & INESC-ID, Universidade de Lisboa, Portugal(里斯本大学技术高级学院及INESC-ID研究所,里斯本大学,葡萄牙)
;
Dept. of Computer Science and AI & DaSCI Institute, Universidad de Granada, Spain(计算机科学与人工智能系及DaSCI研究所,格拉纳达大学,西班牙)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI