AI Assistance Reduces Persistence and Hurts Independent Performance
人工智能辅助降低坚持性并损害独立表现
Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
University of Oxford(牛津大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of California, Los Angeles(加州大学洛杉矶分校)
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
从可验证奖励强化学习到自验证奖励强化学习:任务转换为开放式语言模型自我改进带来自验证奖励
Qinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao
机构
*
Duke University(杜克大学)
;
Adobe Inc.(奥多比公司)
;
Oregon State University(俄勒冈州立大学)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
National University of Singapore(新加坡国立大学)
;
Amazon(亚马逊)
机构
*
Grainger College of Engineering, Department of Civil and Environmental Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格拉inger工程学院土木与环境工程系)
CommentsUpdate to the latest ACLpublished version and add a link to the released code
Journal refProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, 2026, pp. 16764-16781
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
IMProofBench:在研究级数学证明生成上对人工智能进行基准测试
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, João Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontanari, Filippo Gaia, Baran Hashemi, Daniel Holmes, David Holmes, Aitor Iribar Lopez, Victor Jaeck, Martina Jørgensen, Steven Kelk, Martijn Kool, Stefan Kuhlmann, Adam Kurpisz, Johannes Lengler, Chiara Meroni, Ingmar Metzler, Martin Möller, Samuel Muñoz-Echániz, David Muñoz-Lahoz, Robert Nowak, Georg Oberdieck, Daniel Platt, Dylan Possamaï, Gabriel Ribeiro, Aluna Rizzoli, Daria Sakhanda, Raúl Sánchez Galán, Zheming Sun, Diaaeldin Taha, Josef Teichmann, Richard P. Thomas, Henk van der Pol, Michel van Garrel, Charles Vial, Ignacio Barros, Benjamin Doerr, Peter Grünwald, Henry Liu, David Martins, Aleksandar Mijatović, Sergej Monavari, Marc Roth, Patrick Schnider, Yannik Schuler, Pim Spelier, Yuuji Tanaka, Ronald van Luijk
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
Aarhus University(奥胡斯大学)
Commentsv2: benchmark expanded from 39 to 77 problems; evaluation extended to 14 models including GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6; new analyses (IRT-based score aggregation, inter-rater reliability, tool/token usage, non-agentic ablation); contributor author list updated
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
AfriqueLLM: 数据混合与模型架构如何影响非洲语言的持续预训练
Hao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani
机构
*
McGill University(麦吉尔大学)
;
Mila-Quebec AI Institute(魁北克AI研究所)
;
LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心)
;
Microsoft AI for Good Research Lab(微软AI for Good研究实验室)
Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery
Science Earth: 迈向面向AI原生科学发现的行星级操作系统
Zhe Zhao, Haibin Wen, Yingcheng Wu, Jiaming Ma, Yifan Wen, Jinglin Jian, Jiacheng Ge, Xiangru Tang, Bo An, Ming Yin, Sanfeng Wu, Mengdi Wang, Le Cong
机构
*
Department of Pathology, Department of Genetics, Stanford University School of Medicine(病理学系、遗传学系,斯坦福大学医学院)
;
Princeton AI Lab, Department of Electrical & Computer Engineering, Princeton University(普林斯顿人工智能实验室、电气与计算机工程系,普林斯顿大学)
;
Scripps Research, La Jolla, CA, USA(斯克里普斯研究机构,洛杉矶,加利福尼亚州,美国)
;
Division of Biostatistics, Department of Population Health, New York University Grossman School of Medicine(生物统计学部、人口健康系,纽约大学格罗斯曼医学院)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
;
Department of Computer Science, Yale University(计算机科学系,耶鲁大学)
;
Department of Physics, Princeton University(物理系,普林斯顿大学)
CommentsWithdrawn by the authors. (1) The author list and authorship roles had not been finalized and agreed upon by all listed authors prior to submission. (2) The specific contribution of the system in the K3 synchronization example (Section on Kuramoto/nonlinear physics) requires further validation before it can be reported. The authors are addressing both points and may resubmit a corrected version.
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Jinesis Lab, University of Toronto & Vector Institute(Jinesis实验室,多伦多大学及向量研究所)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Princeton University(普林斯顿大学)
;
Cornell University(康奈尔大学)
;
The University of Tokyo(东京大学)
;
RIKEN AIP(日本理化学研究所AIP)
;
Max Planck Institute for Intelligent Systems, Tübingen, Germany(德国图宾根最大计划智能系统研究所)
;
EuroSafeAI