Journal refProceedings of The fourth international workshop on the role of resources in the age of large language models RESOURCEFUL-2026 at LREC 2026, Palma de Mallorca, Spain, 2026
HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
HaluNet:从LLM问答内部信号学习幻觉风险
Chaodong Tong, Qi Zhang, Zhuojun Jiang, Lei Jiang, Yanbing Liu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
China Industrial Control Systems Cyber Emergency Response Team(中国工业控制系统网络应急响应团队)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
Comments16 pages, 12 tables, and 11 figures. This version includes a major revision of the manuscript and updates the author list with the consent of all involved authors
Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents
现在询问,以后使用:评估长期 LLM 代理中的主动性差距
Bin Wu, Guanyun Zou, Bingbing Wang, Huan Zhao, Chuan Shi
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Noumena AI
机构
*
University of New South Wales, NSW, Sydney, Australia(新南威尔士大学,新州,悉尼,澳大利亚)
;
Suzhou Institute for Advanced Research, University of Science(苏州先进研究院,科学大学)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
Cornell University(康奈尔大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.CL、cs.AI
AI总结
提出AtomWorld基准,通过十种基本原子结构操作评估LLM在材料科学中的空间推理能力,发现Claude Opus 4.6表现最佳但复杂空间关系操作成功率低,表明LLM更适合作为辅助工具而非完全自主的科研代理。
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
测量基于LLM的简历筛选中真实世界的提示注入攻击
Mohan Zhang, Yuqi Jia, Zhen Tan, Steven Jiang, Neil Zhenqiang Gong, Tianlong Chen, Dawn Song
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Duke University(杜克大学)
;
Arizona State University(亚利桑那州立大学)
;
hireEZ
;
University of California, Berkeley(加州大学伯克利分校)
Comments16 pages, 7 figures, 12 tables. Accepted to the ICML 2026 Workshop on Hypothesis Testing, Seoul, South Korea, 2026. Copyright 2026 by the author(s)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
基于LLM的自动评分中可学习的评估技能:通过迭代优化构建评分标准
Yun Wang, Xin Xia, Xuansheng Wu, Xiaoming Zhai, Ninghao Liu
机构
*
School of Computing, University of Georgia, Athens, GA, USA(佐治亚大学计算机学院)
;
AI4STEM Education Center, University of Georgia, Athens, GA, USA(佐治亚大学AI4STEM教育中心)
;
The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学)