arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Washington(华盛顿大学)

共收录 1148
2608.11368 2026-08-14 cs.LG 版本更新

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR

PAIR:用于RLVR中自适应rollout分配的成对感知包含重加权方法

Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of Washington(华盛顿大学)

AI总结 针对RLVR中自适应rollout分配的统计不匹配问题,提出PAIR方法,通过成对对比图校正梯度估计,在Qwen3模型上提升准确率同时减少token使用量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10860 2026-08-14 cs.RO cs.CV 版本更新

Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility

Flex-$π$:具备计算灵活性的多流世界-动作模型

Ge Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox

机构 * University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所)

AI总结 该研究提出Flex-$π$多流世界-动作模型,利用冻结的视频生成VAE实现多模态监督,通过混合专家Transformer与每流丢弃机制提升效率,在双臂操作任务上性能优于基线模型且运行更快。

Comments Project page: https://flex-pi.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09164 2026-08-14 cs.AI 版本更新

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

CIDER:用于隐私偏好对齐的上下文披露边界数据集

Bingcan Guo, Eryue Xu, Jijie Zhou, Zhiping Zhang, Tianshi Li

机构 * University of Washington(华盛顿大学) UIUC(伊利诺伊大学厄巴纳-香槟分校) Northeastern University(东北大学)

AI总结 本文推出包含169名用户标注的CIDER数据集,用于评估LLM隐私偏好对齐,发现上下文个性化可提升预测准确率,GPT-5.4和Claude Sonnet 4.6表现更优。

Comments Accepted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12253 2026-08-13 cs.CL cs.AI cs.LG 新提交

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

单个冻结模拟器不够:多智能体强化学习中的模拟器崩溃问题

Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi

机构 * Northeastern University(东北大学) New York University(纽约大学) UC Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

AI总结 针对人机交互多智能体强化学习中单个LLM模拟器导致的策略泛化缺陷,提出Verbalized Sampling和Co-Training两种方案,在多轮基准测试和真实用户研究中显著提升了性能,发布了开源框架SCOPE。

Comments 41 pages, 28 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12185 2026-08-13 cs.CV 新提交

GenFAR: A generalized representation of brain structure, derived from 49,246 multi-cohort MRIs via deep learning

GenFAR:基于49246例多队列MRI通过深度学习得到的脑结构通用表征

Vishnu M. Bashyam, Guray Erus, Junhao Wen, Pratik Chaudhari, Randa Melhem, Sindhuja Govindarajan Tirumalai, Gareth Harman, Yong Fan, Colin L. Masters, Paul Maruff, Sterling C. Johnson, Jurgen Fripp, Duygu Tosun, John C. Morris, Daniel S. Marcus, Pamela LaMontagne, Tammie Benzinger, Susan R. Heckbert, Mark Espeland, Marilyn S. Albert, Andrew J. Saykin, Paul M. Thompson, Timothy J. Hohman, Susan M. Resnick, R. Nick Bryan, Murat Bilgel, Yang An, David A. Wolk, Li Shen, Haochang Shou, Ilya M. Nasrallah, Christos Davatzikos

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Melbourne(墨尔本大学) University of Wisconsin School of Medicine and Public Health(威斯康星大学医学与公共卫生学院) CSIRO Health and Biosecurity(联邦科学与工业研究组织健康与生物安全部) CSIRO(联邦科学与工业研究组织) University of California, San Francisco(加利福尼亚大学旧金山分校) Washington University in St. Louis(圣路易斯华盛顿大学) University of Washington(华盛顿大学) Wake Forest School of Medicine(维克森林医学院) Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院) Indiana University(印第安纳大学) University of Southern California(南加利福尼亚大学)

AI总结 GenFAR是基于49246例多队列MRI的模块化深度学习框架,通过17类任务训练得到通用脑表征,可提升二级预测器的样本效率与准确性

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12236 2026-08-13 cs.RO cs.AI cs.LG 版本更新

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

TMRL: 差分时间步调节预训练实现高效策略微调的探索

Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta

机构 * University of Washington(华盛顿大学) Amazon FAR(亚马逊FAR)

AI总结 本文提出TMRL框架,通过结合行为克隆预训练与强化学习微调,解决预训练中动作分布狭窄的问题,提升机器人策略微调的样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12283 2026-08-13 cs.AI cs.MA 版本更新

Deep Fictitious Play-Based Potential Differential Games for Learning Human-Like Interaction at Unsignalized Intersections

基于深度虚拟博弈的势微分博弈:用于学习无信号交叉口的类人交互

Kehua Chen, Ryan Feng Lin, Shucheng Zhang, Yinhai Wang

机构 * Department of Civil and Environmental Engineering, University of Washington(华盛顿大学土木与环境工程系)

AI总结 本研究提出DFP-PDG框架,首次用深度虚拟博弈学习无信号交叉口类人交互驾驶策略,经INTERACTION数据集验证有效,可捕捉驾驶风格差异并保证收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10296 2026-08-12 cs.CL 新提交

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

基础的裂痕:看似微小的架构选择会影响长上下文扩展

Amanda Bertsch, Luca Soldaini, Matthew R. Gormley, Graham Neubig, Hannaneh Hajishirzi, Kyle Lo, Dirk Groeneveld

机构 * Ai2 Carnegie Mellon University(卡内基梅隆大学) University of Washington(华盛顿大学)

AI总结 该研究发现,Olmo、Llama、Qwen等稠密模型系列的四个微小架构决策会复合降低长上下文性能,结合三个及以上选择可使性能降47%,经17万余GPU小时训练发布OlmPool,其部分模型长上下文性能优于Llama 3。

Comments 29 pages; accepted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10251 2026-08-12 cs.CL cs.LG 新提交

Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

离轴,出于目的:Transformer在何处计算概念以及为何如此计算

Mark Oskin

机构 * University of Washington(华盛顿大学) School of Computer Science and Engineering(计算机科学与工程学院)

AI总结 该研究揭示Transformer分两阶段计算,概念阶段采用与读取轴近乎正交的子空间,强制该几何结构可提升稀疏旋转的收敛性,且不损害语言建模基准性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09273 2026-08-12 cs.LG 版本更新

Instance-Adaptive Online Multicalibration

实例自适应在线多校准

Zhiming Huang, Jamie Morgenstern, Aaron Roth, Claire Jie Zhang

机构 * Paul G. Allen School of Computer Science and Engineering, University of Washington(华盛顿大学保罗·G·阿伦计算机科学与工程学院) Department of Computer and Information Sciences, University of Pennsylvania(宾夕法尼亚大学计算机与信息科学系)

AI总结 本文提出了一种高效的实例自适应在线多校准算法,通过动态调整预测值的二进制网格来平衡最坏情况和易处理情况,实现了在不同实例下的最优误差控制。

Comments Affiliation update only; no changes to technical content

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21436 2026-08-12 stat.ML cs.GT cs.LG 版本更新

Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback

在带凸集的双线性鞍点问题中实现高效解耦学习动力学,具有$\tilde{O}\!\left(T^{-1/4}\right)$的最后迭代收敛性

Arnab Maiti, Claire Jie Zhang, Kevin Jamieson, Jamie Heather Morgenstern, Ioannis Panageas, Lillian J. Ratliff

机构 * University of Washington(华盛顿大学) University of California, Irvine(加州大学尔湾分校)

AI总结 本文提出了一种高效的解耦学习算法,在双线性鞍点问题中实现$\tilde{O}(T^{-1/4})$的最后迭代收敛性,通过结合实验设计和FTRL框架,利用定制正则化函数以高概率收敛到纳什均衡。

Comments 19 pages, accepted at AISTATS 2026. Affiliation update only; no changes to technical content

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05307 2026-08-12 cs.CL 版本更新

MentorCollab: Large-to-Small Inference-Time Mentorship for Concise Reasoning in Language Models

MentorCollab: 选择性大到小推理时指导以实现高效推理

Haojin Wang, Yike Wang, Shangbin Feng, Hannaneh Hajishirzi, Yulia Tsvetkov

机构 * UIUC(伊利诺伊大学香槟分校) University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能研究院)

AI总结 MentorCollab通过选择性推理时指导,使小型模型在多步骤推理任务中实现高效协作,提升性能的同时降低计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08566 2026-08-11 cs.CV cs.AI 新提交

On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation

基于不确定性校准的玻片级聚合的设备端多物种疟疾检测

Idaya Seidu, Ahmed Tahiru Issah, Charles B. Delahunt, Carine Mukamakuza

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校) University of Washington(华盛顿大学)

AI总结 针对资源有限地区疟疾检测的临床约束,开发基于YOLOv13n的设备端多物种疟疾检测系统,满足多物种鉴别等要求,在2739张图像上实现较高检测精度与聚合性能。

Comments Accepted at The Fifth Workshop on Applications of Medical AI (AMAI) 2026, a satellite event at MICCAI 2026. To appear in Springer Lecture Notes in Computer Science (LNCS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07886 2026-08-11 cs.CV cs.AI cs.CL 新提交

Vision-Language Grounding as Bidirectional Concept Correspondence

视觉-语言 Grounding 作为双向概念对应

Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所) FAIR at Meta(Meta FAIR实验室)

AI总结 本研究将视觉-语言 Grounding 建模为双向概念对应,提出 ConCor-1 模型统一相关任务,在长文本数据集和零样本 LVIS 上对应 F1 分别提升 48%、29%,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07567 2026-08-11 cs.CV cs.AI 新提交

Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark

基于fNIRS的自闭症分类中的时间泛化性:跨时间窗口迁移基准

Marios Petrov, Sahana Vinayak, Targol Bakhtiarvand, Moses Smith Guddah, Adham Atyabi, Frederick Shic, Kevin A. Pelphrey

机构 * University of Colorado Colorado Springs(科罗拉多大学科罗拉多斯普林斯分校) University of Washington School of Medicine(华盛顿大学医学院)

AI总结 该研究针对基于fNIRS的自闭症分类中的时间分布偏移问题,构建跨时间窗口迁移基准,发现域对抗等策略可在无目标受试者数据时实现较高准确率,为实际部署提供了路线图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17653 2026-08-11 cs.CL

Differences in Typological Alignment in Language Models' Treatment of Differential Argument Marking

语言模型处理差异论元标记中的类型学对齐差异

Iskar Deng, Nathalia Xu, Shane Steinert-Threlkeld

机构 * University of Washington(华盛顿大学)

AI总结 通过控制合成语料训练GPT-2模型,发现模型在标记方向(自然标记方向)上表现出类人偏好,但在论元角色偏好(宾语vs主语)上未复现人类语言的强烈宾语偏好。

Comments 16 pages, 8 figures, 7 tables. To appear at CoNLL 2026

Journal ref Proceedings of the 30th Conference on Computational Natural Language Learning (CoNLL 2026), pp. 268-283, San Diego, California, USA, Association for Computational Linguistics, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14382 2026-08-11 cs.CV cs.GR cs.MM 版本更新

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation

Delta Forcing:交互式自回归视频生成中的信任区域引导

Yuheng Wu, Xiangbo Gao, Tianhao Chen, Xinghao Chen, Qing Yin, Zhengzhong Tu, Dongman Lee

机构 * Texas A\&M University(德克萨斯A&M大学) University of Washington(华盛顿大学)

AI总结 本文提出Delta Forcing方法,通过约束不可靠的教师监督在适应性信任区域中,以提高自回归视频生成的一致性并保持对新事件的响应性。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14363 2026-08-11 cs.CL cs.AI cs.CV 版本更新

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models

语言的成本:质心擦除揭示并利用多模态语言模型中的模态竞争

Akshay Paruchuri, Ishan Chatterjee, Henry Fuchs, Ehsan Adeli, Piotr Didyk

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) UNC Chapel Hill(北卡罗来纳大学教堂山分校) USI Lugano(卢加诺大学)

AI总结 研究发现多模态语言模型在视觉任务中表现不佳,通过质心替换方法揭示语言表征优于视觉,通过文本质心对比解码提升准确率,不同训练方法影响结果。

Comments 37 pages, 8 figures, 28 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28590 2026-08-11 cs.AI 版本更新

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

MonitorBench: 一个用于大型语言模型链式思维可监控性的综合基准

Han Wang, Yifan Sun, Brian Ko, Mann Talati, Jiawen Gong, Zimeng Li, Naicheng Yu, Xucheng Yu, Wei Shen, Vedant Jolly, Huan Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Washington(华盛顿大学) University of California San Diego(加州大学圣地亚哥分校)

AI总结 本文提出MonitorBench,通过1514个测试实例和两种压力测试设置,评估LLM链式思维的可监控性,发现决策关键因素对中间推理过程的影响程度影响可监控性,更高级的LLM表现出更低的可监控性。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06933 2026-08-10 cs.CL cs.AI 新提交

Ask-E: An Environment for Calibrated Question Generation

Ask-E:一个用于校准型问题生成的环境

Sarah Pratt, Jae Sung Park, Scott Geng, Ali Farhadi

机构 * University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所)

AI总结 本研究提出Ask-E环境,以两个现有语言模型的能力范围定义目标技能水平,通过生成仅能被其中一个模型解决的问题来校准模型,该环境可用于基准测试与训练,且训练后模型在下游数学基准上表现提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06488 2026-08-10 cs.RO 新提交

A Disturbance in the Force: Force Actuation on the RAVEN II Surgical Robot with Parallel Motor-Cable Units

力的扰动:采用并联电机-线缆单元对RAVEN II手术机器人进行力驱动

Haonan Peng, Dun-Tin Chiang, Jordan Hendricks, Andrew Lewis, Jared Shing, Haokun Feng, Yun-Hsuan Su, Blake Hannaford

机构 * University of Washington(华盛顿大学) Mount Holyoke College(曼荷莲学院)

AI总结 针对手术机器人力反馈难题,开发含6个电机-线缆单元的并联系统,实现无干扰力驱动,力驱动误差小于1N,为获取训练数据提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02659 2026-08-10 cs.LG cs.AI cs.CE cs.NA math.NA 版本更新

In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization

隐式神经压缩器的原地训练:基于草图的正则化

Cooper Simpson, Stephen Becker, Alireza Doostan

机构 * Applied Mathematics, University of Washington, Seattle(华盛顿大学应用数学系,西雅图) Applied Mathematics, University of Colorado, Boulder(科罗拉多大学应用数学系,伯尔德) Ann and H.J. Smead Department of Aerospace Engineering Sciences, University of Colorado, Boulder(科罗拉多大学伯尔德分校安与H.J. 塞梅德航空航天工程科学系)

AI总结 本文提出基于草图正则化的隐式神经压缩器原地训练方法,通过有限内存缓冲区实现高压缩率下的重建性能,展示草图能近似匹配离线方法性能。

Comments 18 pages, 8 figures, 5 tables

Journal ref Journal of Computational Physics, Volume 566, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27435 2026-08-10 cs.CL cs.AI 版本更新

Improving Attributed Long-form Question Answering with Intent Awareness

通过意图意识提升带有属性的长形式问答

Xinran Zhao, Aakanksha Naik, Jay DeYoung, Joseph Chee Chang, Jena D. Hwang, Tongshuang Wu, Varsha Kishore

机构 * Allen Institute for AI(艾伦人工智能研究所) Carnegie Mellon University(卡内基梅隆大学) University of Washington(华盛顿大学)

AI总结 本文提出通过增强模型意图意识来提升长报告生成质量,通过结构化标签方案提取隐含意图,改进零样本生成能力并生成高质量合成数据,实验显示在科学报告生成任务中模型性能平均提升2.9和12.3个百分点。

Comments 39 pages, 7 figures

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05850 2026-08-07 cs.CL cs.AI 新提交

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

MameLoshnLM:意第绪语语言模型与评估基准

Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith

机构 * Bar-Ilan University(巴伊兰大学) University of Cambridge(剑桥大学) University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所)

AI总结 该研究推出首个专为意第绪语构建的开源8B参数语言模型MameLoshnLM,通过自研语料库与基准优化Llama 3.1 8B,其表现优于同规模开源基线,为意第绪语NLP及低资源语言模型开发提供了基础与模板。

Comments Accepted at the Conference on Language Modeling (COLM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05716 2026-08-07 cs.AI 新提交

BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming

BlockPython:一种支持从积木式编程向Python编程过渡的进程感知智能体平台

Jesse Yusuf Chan, Haoming Wang, Mingwei Xu, Xianlong Xu

机构 * East China Normal University(华东师范大学) Tsinghua University(清华大学) University of Washington(华盛顿大学)

AI总结 BlockPython是支持从积木式编程向Python过渡的平台,以双向转换为核心,通过四阶段引导学习者,利用进程证据诊断困难,为相关系统设计提供参考。

Comments AIED 2026 Interactive Event Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05122 2026-08-07 cs.CV 版本更新

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

IRIS:一种受视觉皮层启发的用于分析视觉Transformer中方向选择性的框架

Vaishnavi B Mohan, Vijayakrishna Naganoor, Yashas Annadani, Shashank Hegde

机构 * University of Washington(华盛顿大学) Microsoft(微软公司) TU Munich(慕尼黑工业大学) Gladstone Institute(格拉德斯通研究所) Nvidia(英伟达公司)

AI总结 该研究提出受视觉皮层启发的IRIS框架,通过RSS等神经科学指标分析ViT中方向选择性的形成,发现训练范式决定方向选择性、层的方向选择性变化规律,且指标可指导解冻层数以优化下游泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04511 2026-08-07 cs.SD 版本更新

A Dual Evaluation for Music Transcription

音乐转录的双重评估

Ping Wang, Guang Yang, Nazif Can Tamer, Victoria Ebert, Noah A. Smith

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(艾伦人工智能研究所)

AI总结 该研究针对音乐转录系统提出记谱相似度与回放相似度的双重评估框架,发现CLEWS指标相关性佳且成本低,不同评估维度偏好不同系统,新系统Rubato记谱相似度提升且回放具竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12519 2026-08-07 cs.CL cs.AI 版本更新

Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models

通过可验证的推理过程监督获得正确答案:语言模型的可验证过程监督

Kyuyoung Kim, Kevin Wang, Yunfei Xie, Peiyang Xu, Peiyao Sheng, Chen Wei, Zhangyang Wang, Jinwoo Shin, Pramod Viswanath, Sewoong Oh

机构 * KAIST AI(KAIST人工智能研究所) University of Texas, Austin(德克萨斯大学奥斯汀分校) Rice University(里士满大学) Princeton University(普林斯顿大学) University of Washington(华盛顿大学) Sentient Labs(Sentient实验室)

AI总结 本文提出VPS框架,通过联合优化预测准确性和推理质量,在可验证领域使语言模型同时获得准确和可靠的推理能力。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23565 2026-08-07 cs.LG cs.MA 版本更新

Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing

学习动态中的用户选择:过度专业化与同伴模型探测

Adhyyan Narang, Sarah Dean, Lillian J Ratliff, Maryam Fazel

机构 * Electrical and Computer Engineering, University of Washington(华盛顿大学电气与计算机工程学院) Computer Science, Cornell University(康奈尔大学计算机科学系)

AI总结 本文研究了用户选择对学习动态的影响,提出通过探测同伴模型预测来避免过度专业化陷阱,从而提升模型的全局性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16763 2026-08-07 cs.AI 版本更新

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

当AI基准测试达到平台期:基准饱和的系统性研究

Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman

机构 * University of California, Berkeley(加州大学伯克利分校) University of Toronto(多伦多大学) University of Washington(华盛顿大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Michigan(密歇根大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本研究定义并分析了60个语言模型基准的饱和现象,发现近半数基准出现饱和,且专家策划而非公开测试数据影响抗饱和能力,为延长基准寿命提供了设计建议。

Comments Published at ICML 2026 (Forty-Third International Conference on Machine Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏