arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

University of Washington(华盛顿大学)

2026-03-03 至 2026-03-03 共收录 12
2512.01351 2026-03-03 cs.AI

Benchmarking Overton Pluralism in LLMs

对大语言模型中Overton多元主义的基准测试

Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker

机构 * Massachusetts Institute of Technology(麻省理工学院) Brown University(布朗大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

AI总结 本文提出OVERTONBENCH框架,通过集合覆盖度量评估大语言模型中多元观点的代表性,揭示模型在多元主义对齐上的改进空间。

Comments Paper accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24119 2026-03-03 cs.AI cs.CL cs.LG

SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning

SPIRAL:通过多智能体多轮强化学习进行零和游戏的自我对战以促进推理

Bo Liu, Leon Guertler, Simon Yu, Zichen Liu, Penghui Qi, Daniel Balcells, Mickel Liu, Cheston Tan, Weiyan Shi, Min Lin, Wee Sun Lee, Natasha Jaques

机构 * National University of Singapore(新加坡国立大学) Northeastern University(东北大学) Sea AI Lab(Sea AI 实验室) Centre for Frontier AI Research (CFAR), A*STAR(前沿人工智能研究中心(CFAR),A*STAR) Plastic Labs University of Washington(华盛顿大学)

AI总结 SPIRAL通过多智能体多轮强化学习在零和游戏中促进推理,展示了模型在多个基准测试中的显著性能提升。

Comments Accepted at ICLR 2026. Code: https://github.com/spiral-rl/spiral

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01449 2026-03-03 eess.IV cs.CV

Revisiting Global Token Mixing in Task-Dependent MRI Restoration: Insights from Minimal Gated CNN Baselines

重新审视任务依赖性MRI修复中的全局token混合:来自最小门控CNN基线的见解

Xiangjian Hou, Chao Qin, Chang Ni, Xin Wang, Chun Yuan, Xiaodong Ma

机构 * Dept. of Radiology & Imaging Sciences, University of Utah, Salt Lake City, UT, USA(放射学与成像科学系,犹他大学,盐湖城,UT,USA) Dept. of Electrical & Computer Engineering, University of Utah, Salt Lake City, UT, USA(电气与计算机工程系,犹他大学,盐湖城,UT,USA) Dept. of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE(计算机视觉系,莫扎德人工智能大学,阿布扎赫,EUA) Dept. of Biomedical Engineering, University of Utah, Salt Lake City, UT, USA(生物医学工程系,犹他大学,盐湖城,UT,USA) Dept. of Electrical & Computer Engineering, University of Washington, Seattle, WA, USA(电气与计算机工程系,华盛顿大学,西雅图,WA,USA)

AI总结 本文探讨了MRI修复中全局token混合的效用,发现其在不同任务中表现各异,需根据具体退化结构和物理条件进行调整。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01343 2026-03-03 cs.CL cs.AI

PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology

PanCanBench: 用于评估大型语言模型在胰腺肿瘤学中的综合基准

Yimin Zhao, Sheela R. Damle, Simone E. Dekker, Scott Geng, Karly Williams Silva, Jesse J Hubbard, Manuel F Fernandez, Fatima Zelada-Arenas, Alejandra Alvarez, Brianne Flores, Alexis Rodriguez, Stephen Salerno, Carrie Wright, Zihao Wang, Pang Wei Koh, Jeffrey T. Leek

机构 * Department of Biostatistics, University of Washington(华盛顿大学生物统计学系) Clinical Research Division, Fred Hutch Cancer Center(Fred Hutch癌症中心临床研究部) Division of Hematology and Oncology, Department of Medicine, University of Washington(华盛顿大学医学系血液学与肿瘤学分会) Allen Institute for AI(Allen人工智能研究所) Department of Computer Science and Engineering, University of Washington(华盛顿大学计算机科学与工程系) Public Health Sciences, Biostatistics, Fred Hutchinson Cancer Center(Fred Hutchinson癌症中心公共卫生科学与生物统计学)

AI总结 PanCanBench通过评估22种LLM在胰腺肿瘤学问题上的表现,揭示了模型在事实准确性、临床完整性和网络搜索整合方面的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05724 2026-03-03 cs.AI

Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding

通过无损分层推测解码克服联合不可解性

Yuxuan Zhou, Fei Huang, Heng Li, Fengyi Wu, Tianyu Wang, Jianwei Zhang, Junyang Lin, Zhi-Qi Cheng

机构 * Qwen Team, Alibaba Inc.(阿里云团队,阿里巴巴公司) University of Washington(华盛顿大学)

AI总结 本研究提出分层推测解码方法,通过平衡概率质量克服联合不可解性,提升解码效率并保持分布保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02386 2026-03-03 cs.CR cs.AI cs.LG

On The Fragility of Benchmark Contamination Detection in Reasoning Models

在推理模型中基准污染检测的脆弱性

Han Wang, Haoyu Li, Brian Ko, Huan Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Washington(华盛顿大学)

AI总结 研究发现LRMs在基准污染检测中存在脆弱性,通过简单方法可轻易污染模型以提升排行榜表现,威胁评估公平性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01481 2026-03-03 cs.LG cs.CL

Intrinsic Entropy of Context Length Scaling in LLMs

大语言模型中上下文长度扩展的内在熵

Jingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao, Jeng-Neng Hwang, Lei Li

机构 * Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Carnegie Mellon University(卡内基梅隆大学) CPHOS University of Washington(华盛顿大学)

AI总结 本研究提出'内在熵'理论,通过实验验证长上下文对语言模型的影响,揭示训练数据集大小与最佳上下文长度的关系。

Comments 36 pages, 18 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01019 2026-03-03 cs.CR cs.LG

BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models

BadRSSD: 对于正则化自监督扩散模型的后门攻击

Jiayao Wang, Yiping Zhang, Mohammad Maruf Hasan, Xiaoying Lei, Jiale Zhang, Junwu Zhu, Qilin Wu, Dongfang Zhao

机构 * School of Information and Artificial Intelligence, Yangzhou University, China(扬州大学信息与人工智能学院) School of Computing and Artificial Intelligence, Chaohu University, China(池州学院计算机与人工智能学院) Tacoma School of Engineering and Technology, University of Washington, USA(华盛顿大学塔科马工程与技术学院)

AI总结 BadRSSD是一种针对自监督扩散模型表示层的后门攻击方法,通过在PCA空间中操控语义表示并应用跨空间约束,实现隐蔽的后门攻击并提升攻击效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00925 2026-03-03 cs.CL cs.CV cs.CY

The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors

DrawEduMath之后的后果:视觉语言模型在挣扎学生面前表现不佳且误判错误

Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo

机构 * University of Washington(华盛顿大学) Insource Services(Insource服务) Worcester Polytechnic Institute(沃斯特理工学院) Allen Institute for AI(人工智能联合研究所)

AI总结 本文研究了视觉语言模型在处理学生数学错误识别任务中的表现,发现其在挣扎学生面前表现不佳,需改进以支持教育应用。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00923 2026-03-03 cs.CL

Hybrid Neural-LLM Pipeline for Morphological Glossing in Endangered Language Documentation: A Case Study of Jungar Tuvan

混合神经-大语言模型管道在濒危语言文档中的形态学 glossing:朱加尔图瓦语案例研究

Siyu Liang, Talant Mawkanuli, Gina-Anne Levow

机构 * Department of Linguistics, University of Washington(语言学系,华盛顿大学) Department of Middle Eastern Languages and Cultures, University of Washington(中东语言与文化系,华盛顿大学)

AI总结 本文提出一种混合神经-大语言模型管道,用于朱加尔图瓦语的形态学 glossing,通过结合神经序列标注和 LLM 后修正,显著减少标注工作量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00822 2026-03-03 cs.SE cs.AI

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files

ContextCov: 从代理指令文件中推导并强制执行可执行约束

Reshabh K Sharma

机构 * University of Washington(华盛顿大学)

AI总结 ContextCov通过从代理指令文件中推导并强制执行可执行约束,解决代理在执行复杂软件任务时因上下文漂移导致的技术债务问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27492 2026-03-03 cs.CV

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

ThinkMorph:多模态交错链式推理中的涌现特性

Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学) absolute AI The Chinese University of Hong Kong(香港中文大学)

AI总结 ThinkMorph通过统一模型提升多模态推理性能,展现视觉操控与模式切换等新兴能力。

Comments project page: https://thinkmorph.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏