arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5813 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5813 篇

2605.09268 2026-05-12 cs.CL cs.AI 62%

Beyond Continuity: Challenges of Context Switching in Multi-Turn Dialogue with LLMs

超越连续性:在LLMs多轮对话中上下文切换的挑战

Aditya Sinha, Harald Steck, Vito Ostuni, Matteo Rinaldi

机构 * Netflix Inc.(Netflix公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了LLMs在多轮对话中处理上下文切换的挑战,通过构建合成基准测试,评估了十种LLMs在检测用户话题切换和筛选相关上下文方面的性能,发现只有部分具备推理能力的LLMs能准确检测话题切换。

Comments Accepted to the ICBINB Workshop @ ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06225 2026-05-12 cs.LG cs.AI 62%

Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

记忆 inception:用于引导 LLMs 的潜在空间 KV 缓存操作

Andy Zeyi Liu, Michael Zhang, Ilana Greenberg, Adam Alnasser, Lucas Baker, John Sous

机构 * Yale University(耶鲁大学) Princeton University(普林斯顿大学) Jump Trading

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出 memory inception 方法,通过在选定层插入文本衍生的键值对(KV)银行,在潜在注意力空间中引导 LLMs,实现更高效的控制与更少的存储消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13486 2026-05-12 cs.LG cs.AI cs.DC 62%

Preventing Rank Collapse in Federated Low-Rank Adaptation with Client Heterogeneity

防止联邦低秩适应中的秩坍缩与客户端异质性

Fei Wu, Jia Hu, Geyong Min, Shiqiang Wang

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文针对联邦低秩适应中因客户端异质性导致的秩坍缩问题,提出raFLoRA方法,通过分解本地更新并加权聚合提升模型性能和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12235 2026-05-12 cs.LG cs.AI 62%

RL Fine-Tuning Heals OOD Forgetting in SFT

强化学习修复SFT中的分布外遗忘

Hangzhan Jin, Sitao Luan, Tianwei Ni, Sicheng Lyu, Guillaume Rabusseau, Reihaneh Rabbany, Doina Precup, Mohammad Hamdaqa

机构 * Mila - Quebec AI Institute(魁北克人工智能研究所) Polytechnique Montréal(蒙特利尔理工学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) CIFAR AI Chair(CIFAR人工智能主席) Google DeepMind(谷歌DeepMind)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究通过分析SFT和RL对ID和OOD推理的影响,发现RL能修复后期SFT导致的OOD能力下降,且与奇异向量旋转相关。

Comments 31 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08301 2026-05-12 cs.LG cs.AI 62%

Priming: Hybrid State Space Models From Pre-trained Transformers

预训练:从预训练变换器中获得的混合状态空间模型

Aditya Chattopadhyay, Elvis Nunez, Prannay Kaul, Benjamin Bowman, Evan Becker, Luca Zancato, David Thomas, Wei Xia, Stefano Soatto

机构 * AWS Agentic AI(AWS 代理AI)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 预训练通过从预训练变换器初始化混合模型,利用短对齐和微调阶段,在少于源模型预训练token预算的情况下恢复下游性能,实现了大规模混合架构设计的可控比较。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08134 2026-05-12 cs.LG cs.AI 62%

DARE: Diffusion Language Model Activation Reuse for Efficient Inference

DARE:扩散语言模型激活重用以提高推理效率

Natalia Frumkin, Bokun Wang, Hung-Yueh Chiang, Chi-Chih Chang, Mohamed S. Abdelfattah, Diana Marculescu

机构 * Chandra Family Department of Electrical and Computer Engineering(查德拉家族电子与计算机工程系) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Cornell University(康奈尔大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 DARE通过利用扩散语言模型中token级冗余性,提高推理效率,减少计算量,同时保持生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07260 2026-05-11 cs.LG cs.CL 62%

When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

专家误路由何时发生?混合专家语言模型中的反事实路由分析

Youngsik Yoon, Siwei Wang, Wei Chen, Jungseul Ok

机构 * Department of Computer Science and Engineering, POSTECH, South Korea(韩国POSTECH计算机科学与工程系) Graduate School of Artificial Intelligence, POSTECH, South Korea(韩国POSTECH人工智能研究生院) Microsoft Research Asia, Beijing, China(中国北京微软亚洲研究院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 研究探讨了混合专家模型中路由选择的质量评估,发现标准路由在自信 token 上表现良好,但在脆弱 token 上存在误分配问题,通过更新路由层可提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06850 2026-05-11 cs.LG cs.AI 62%

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment

如何在RL后训练中压缩KV缓存?基于影子遮罩的内存高效对齐

Rui Zhu, Weiheng Bai, Qiushi Wu, Yang Ren, Haixu Tang, Yuchu Liu

机构 * Yale University(耶鲁大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出影子遮罩蒸馏方法,用于缓解RL后训练中KV缓存内存瓶颈问题,通过减少采样时的上下文密度来降低偏置,提升样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06654 2026-05-08 cs.LG cs.AI math.OC 62%

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

优化器-模型一致性:使用与预训练相同的优化器进行全微调可减少遗忘

Yuxing Liu, Jianyu Wang, Tong Zhang

机构 * UIUC(伊利诺伊大学香槟分校) Apple(苹果公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文发现使用与预训练相同的优化器进行全微调,在监督微调阶段能更少遗忘并保持性能,提出优化器-模型一致性概念,通过实验和理论分析揭示优化器对模型的影响及微调策略的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06196 2026-05-08 cs.AI cs.CL 62%

The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models

粒度轴:语言模型中社会角色的微到宏观潜在方向

Chonghan Qin, Xiachong Feng, Ziyun Song, Xiaocheng Feng, Jing Xiong, Lingpeng Kong

机构 * The University of Hong Kong(香港大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究揭示语言模型中社会角色粒度的潜在方向,通过定义粒度轴,发现其主导了角色表示空间的几何结构,并通过实验验证其在不同层级和模型间的稳定性与因果相关性。

Comments 28 pages, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05965 2026-05-08 cs.LG cs.AI 62%

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR

超越均匀信用分配:用于RLVR的选区资格痕迹

Chaoli Mou, Zhan Zhuang, Xinning Chen, Yu Zhang

机构 * Southern University of Science and Technology(南方科技大学) City University of Hong Kong(香港城市大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出选区资格痕迹(S-trace),通过稀疏资格痕迹机制减少方差,实现细粒度信用分配,实验显示其在Qwen3系列模型上表现优于GRPO,且具备更高的样本和令牌效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17866 2026-05-08 cs.CL cs.AI 62%

Latent Abstraction for Retrieval-Augmented Generation

检索增强生成的潜在抽象

Ha Lan N. T, Minh-Anh Nguyen, Dung D. Le

机构 * Center for AI Research, VinUniversity, Vietnam(越南Vin大学人工智能研究中心)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 LAnR通过在潜在空间中联合执行编码、检索和生成,提升检索增强生成的效率与效果,减少检索调用次数并增强模型整合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15270 2026-05-08 cs.CL cs.AI 62%

From Documents to Spans: Scalable Supervision for Evidence-Based ICD Coding with LLMs

从文档到跨度:基于LLM的证据导向ICD编码可扩展监督方法

Xu Zhang, Wenxin Ma, Chenxu Wu, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Kun Zhang, S. Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences and Medicine, USTC(生物医学工程学院,生命科学与医学系,中国科学技术大学) MIRACLE Center, Suzhou Institute for Advance Research, USTC(MIRACLE中心,苏州先进研究院,中国科学技术大学) Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology(江苏省多模态数字孪生技术重点实验室) State Key Laboratory of Precision and Intelligent Chemistry, USTC(精密与智能化学国家重点实验室)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Span-Centric Learning框架,通过局部跨度学习提升ICD编码能力,以更低成本实现宏F1提升8.2点,并提供明确证据支持预测代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05632 2026-05-08 cs.CR cs.CL cs.LG 62%

Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning

架构至关重要:在知识库中毒情况下比较RAG系统

Samuel Korn

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 研究探讨了在知识库中毒情况下不同RAG架构的鲁棒性,发现架构设计对攻击成功率影响显著,MADAM-RAG在矛盾检测上表现最佳但仍有不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03267 2026-05-05 cs.CL cs.AI 62%

OpenAI GPT-5 System Card

OpenAI GPT-5 系统卡片

Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Alexey Ivanov, Alexi Christakis, Alistair Gillespie, Allison Tam, Ally Bennett, Alvin Wan, Alyssa Huang, Amy McDonald Sandjideh, Amy Yang, Ananya Kumar, Andre Saraiva, Andrea Vallone, Andrei Gheorghe, Andres Garcia Garcia, Andrew Braunstein, Andrew Liu, Andrew Schmidt, Andrey Mereskin, Andrey Mishchenko, Andy Applebaum, Andy Rogerson, Ann Rajan, Annie Wei, Anoop Kotha, Anubha Srivastava, Anushree Agrawal, Arun Vijayvergiya, Ashley Tyra, Ashvin Nair, Avi Nayak, Ben Eggers, Bessie Ji, Beth Hoover, Bill Chen, Blair Chen, Boaz Barak, Borys Minaiev, Botao Hao, Bowen Baker, Brad Lightcap, Brandon McKinzie, Brandon Wang, Brendan Quinn, Brian Fioca, Brian Hsu, Brian Yang, Brian Yu, Brian Zhang, Brittany Brenner, Callie Riggins Zetino, Cameron Raymond, Camillo Lugaresi, Carolina Paz, Cary Hudson, Cedric Whitney, Chak Li, Charles Chen, Charlotte Cole, Chelsea Voss, Chen Ding, Chen Shen, Chengdu Huang, Chris Colby, Chris Hallacy, Chris Koch, Chris Lu, Christina Kaplan, Christina Kim, CJ Minott-Henriques, Cliff Frey, Cody Yu, Coley Czarnecki, Colin Reid, Colin Wei, Cory Decareaux, Cristina Scheau, Cyril Zhang, Cyrus Forbes, Da Tang, Dakota Goldberg, Dan Roberts, Dana Palmie, Daniel Kappler, Daniel Levine, Daniel Wright, Dave Leo, David Lin, David Robinson, Declan Grabb, Derek Chen, Derek Lim, Derek Salama, Dibya Bhattacharjee, Dimitris Tsipras, Dinghua Li, Dingli Yu, DJ Strouse, Drew Williams, Dylan Hunn, Ed Bayes, Edwin Arbus, Ekin Akyurek, Elaine Ya Le, Elana Widmann, Eli Yani, Elizabeth Proehl, Enis Sert, Enoch Cheung, Eri Schwartz, Eric Han, Eric Jiang, Eric Mitchell, Eric Sigler, Eric Wallace, Erik Ritter, Erin Kavanaugh, Evan Mays, Evgenii Nikishin, Fangyuan Li, Felipe Petroski Such, Filipe de Avila Belbute Peres, Filippo Raso, Florent Bekerman, Foivos Tsimpourlas, Fotis Chantzis, Francis Song, Francis Zhang, Gaby Raila, Garrett McGrath, Gary Briggs, Gary Yang, Giambattista Parascandolo, Gildas Chabot, Grace Kim, Grace Zhao, Gregory Valiant, Guillaume Leclerc, Hadi Salman, Hanson Wang, Hao Sheng, Haoming Jiang, Haoyu Wang, Haozhun Jin, Harshit Sikchi, Heather Schmidt, Henry Aspegren, Honglin Chen, Huida Qiu, Hunter Lightman, Ian Covert, Ian Kivlichan, Ian Silber, Ian Sohl, Ibrahim Hammoud, Ignasi Clavera, Ikai Lan, Ilge Akkaya, Ilya Kostrikov, Irina Kofman, Isak Etinger, Ishaan Singal, Jackie Hehir, Jacob Huh, Jacqueline Pan, Jake Wilczynski, Jakub Pachocki, James Lee, James Quinn, Jamie Kiros, Janvi Kalra, Jasmyn Samaroo, Jason Wang, Jason Wolfe, Jay Chen, Jay Wang, Jean Harb, Jeffrey Han, Jeffrey Wang, Jennifer Zhao, Jeremy Chen, Jerene Yang, Jerry Tworek, Jesse Chand, Jessica Landon, Jessica Liang, Ji Lin, Jiancheng Liu, Jianfeng Wang, Jie Tang, Jihan Yin, Joanne Jang, Joel Morris, Joey Flynn, Johannes Ferstad, Johannes Heidecke, John Fishbein, John Hallman, Jonah Grant, Jonathan Chien, Jonathan Gordon, Jongsoo Park, Jordan Liss, Jos Kraaijeveld, Joseph Guay, Joseph Mo, Josh Lawson, Josh McGrath, Joshua Vendrow, Joy Jiao, Julian Lee, Julie Steele, Julie Wang, Junhua Mao, Kai Chen, Kai Hayashi, Kai Xiao, Kamyar Salahi, Kan Wu, Karan Sekhri, Karan Sharma, Karan Singhal, Karen Li, Kenny Nguyen, Keren Gu-Lemberg, Kevin King, Kevin Liu, Kevin Stone, Kevin Yu, Kristen Ying, Kristian Georgiev, Kristie Lim, Kushal Tirumala, Kyle Miller, Lama Ahmad, Larry Lv, Laura Clare, Laurance Fauconnet, Lauren Itow, Lauren Yang, Laurentia Romaniuk, Leah Anise, Lee Byron, Leher Pathak, Leon Maksin, Leyan Lo, Leyton Ho, Li Jing, Liang Wu, Liang Xiong, Lien Mamitsuka, Lin Yang, Lindsay McCallum, Lindsey Held, Liz Bourgeois, Logan Engstrom, Lorenz Kuhn, Louis Feuvrier, Lu Zhang, Lucas Switzer, Lukas Kondraciuk, Lukasz Kaiser, Manas Joglekar, Mandeep Singh, Mandip Shah, Manuka Stratta, Marcus Williams, Mark Chen, Mark Sun, Marselus Cayton, Martin Li, Marvin Zhang, Marwan Aljubeh, Matt Nichols, Matthew Haines, Max Schwarzer, Mayank Gupta, Meghan Shah, Melody Y. Guan, Melody Huang, Meng Dong, Mengqing Wang, Mia Glaese, Micah Carroll, Michael Lampe, Michael Malek, Michael Sharman, Michael Zhang, Michele Wang, Michelle Pokrass, Mihai Florian, Mikhail Pavlov, Miles Wang, Ming Chen, Mingxuan Wang, Minnia Feng, Mo Bavarian, Molly Lin, Moose Abdool, Mostafa Rohaninejad, Nacho Soto, Natalie Staudacher, Natan LaFontaine, Nathan Marwell, Nelson Liu, Nick Preston, Nick Turley, Nicklas Ansman, Nicole Blades, Nikil Pancha, Nikita Mikhaylin, Niko Felix, Nikunj Handa, Nishant Rai, Nitish Keskar, Noam Brown, Ofir Nachum, Oleg Boiko, Oleg Murk, Olivia Watkins, Oona Gleeson, Pamela Mishkin, Patryk Lesiewicz, Paul Baltescu, Pavel Belov, Peter Zhokhov, Philip Pronin, Phillip Guo, Phoebe Thacker, Qi Liu, Qiming Yuan, Qinghua Liu, Rachel Dias, Rachel Puckett, Rahul Arora, Ravi Teja Mullapudi, Raz Gaon, Reah Miyara, Rennie Song, Rishabh Aggarwal, RJ Marsan, Robel Yemiru, Robert Xiong, Rohan Kshirsagar, Rohan Nuttall, Roman Tsiupa, Ronen Eldan, Rose Wang, Roshan James, Roy Ziv, Rui Shu, Ruslan Nigmatullin, Saachi Jain, Saam Talaie, Sam Altman, Sam Arnesen, Sam Toizer, Sam Toyer, Samuel Miserendino, Sandhini Agarwal, Sarah Yoo, Savannah Heon, Scott Ethersmith, Sean Grove, Sean Taylor, Sebastien Bubeck, Sever Banesiu, Shaokyi Amdo, Shengjia Zhao, Sherwin Wu, Shibani Santurkar, Shiyu Zhao, Shraman Ray Chaudhuri, Shreyas Krishnaswamy, Shuaiqi, Xia, Shuyang Cheng, Shyamal Anadkat, Simón Posada Fishman, Simon Tobin, Siyuan Fu, Somay Jain, Song Mei, Sonya Egoian, Spencer Kim, Spug Golden, SQ Mah, Steph Lin, Stephen Imm, Steve Sharpe, Steve Yadlowsky, Sulman Choudhry, Sungwon Eum, Suvansh Sanjeev, Tabarak Khan, Tal Stramer, Tao Wang, Tao Xin, Tarun Gogineni, Taya Christianson, Ted Sanders, Tejal Patwardhan, Thomas Degry, Thomas Shadwell, Tianfu Fu, Tianshi Gao, Timur Garipov, Tina Sriskandarajah, Toki Sherbakov, Tomek Korbak, Tomer Kaftan, Tomo Hiratsuka, Tongzhou Wang, Tony Song, Tony Zhao, Troy Peterson, Val Kharitonov, Victoria Chernova, Vineet Kosaraju, Vishal Kuo, Vitchyr Pong, Vivek Verma, Vlad Petrov, Wanning Jiang, Weixing Zhang, Wenda Zhou, Wenlei Xie, Wenting Zhan, Wes McCabe, Will DePue, Will Ellsworth, Wulfie Bain, Wyatt Thompson, Xiangning Chen, Xiangyu Qi, Xin Xiang, Xinwei Shi, Yann Dubois, Yaodong Yu, Yara Khakbaz, Yifan Wu, Yilei Qian, Yin Tat Lee, Yinbo Chen, Yizhen Zhang, Yizhong Xiong, Yonglong Tian, Young Cha, Yu Bai, Yu Yang, Yuan Yuan, Yuanzhi Li, Yufeng Zhang, Yuguang Yang, Yujia Jin, Yun Jiang, Yunyun Wang, Yushi Wang, Yutian Liu, Zach Stubenvoll, Zehao Dou, Zheng Wu, Zhigang Wang

机构 * OpenAI

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 GPT-5 是一个统一系统,具备快速回答问题的模型、深度推理模型和实时路由器,提升真实世界查询的实用性,减少幻觉并改进指令遵循。

Comments May 2026: Added monitorability evals and authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00140 2026-05-04 cs.LG cs.CL cs.CV 62%

Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization

技术报告:用于低比特大语言模型量化的小激活残差Hessian量化(ARHQ)

YiFeng Wang, Zhun Sun, Keisuke Sakaguchi

机构 * Graduate School of Information Sciences(信息科学研究生院)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 ARHQ通过构建输入侧残差Hessian来隔离误差敏感的权重方向,提升低比特激活权重量化的层间SNR并保持下游推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04212 2026-05-04 cs.CL cs.AI 62%

Language Models Struggle to Use Representations Learned In-Context

语言模型在上下文学习的表示使用上面临困难

Michael A. Lepori, Tal Linzen, Ann Yuan, Katja Filippova

机构 * Brown University(布朗大学) Google Research(谷歌研究) New York University(纽约大学) Google DeepMind(谷歌DeepMind)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了语言模型是否能利用上下文学习的表示完成简单下游任务,发现开放权重模型在处理新语义表示时存在困难,即使它们在潜在表示中编码了这些语义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28182 2026-05-01 cs.LG cs.CL 62%

Exploration Hacking: Can LLMs Learn to Resist RL Training?

探索黑客:大语言模型能否学会抵抗强化学习训练?

Eyon Jang, Damon Falck, Joschka Braun, Nathalie Kirch, Achu Menon, Perusha Moodley, Scott Emmons, Roland S. Zimmermann, David Lindner

机构 * MATS UC San Diego(UC圣迭戈大学) Google DeepMind(谷歌DeepMind) Anthropic

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 本文研究了大语言模型在强化学习训练中可能通过策略性调整探索行为导致失败的机制,通过微调创建了具有特定低效策略的模型,并评估了检测与缓解策略。

Comments 81 pages, 37 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27228 2026-05-01 cs.AI cs.CL cs.CY cs.MA 62%

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

当角色失效时:基于大语言模型的政治声明分析中的知识约束与倡导角色一致性

Juergen Dietrich

机构 * TRUST Project(TRUST项目)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了基于大语言模型的多智能体系统中倡导角色一致性的知识约束,通过TRUST管道测试发现两种失效模式,并揭示模型选择和事实核查提供商对角色一致性的影响。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26779 2026-04-30 cs.LG cs.CL 62%

Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding

通过系统集成的推测解码加速RL训练后的 rollout

Hayate Iso, Tiyasa Mitra, Sudipta Mondal, Rasoul Shafipour, Venmugil Elango, Terry Kong, Yuki Huang, Seonjin Na, Izzy Putterman, Benjamin Chislett, Maor Ashkenazi, Joseph Guman, Gerald Shen, Tugrul Konuk, Ashwath Aithal, Ritika Borkar, Ran Zilberstein, Bita Rouhani

机构 * NeMo-RL vLLM

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 本文研究了通过系统集成的推测解码技术提升RL训练后rollout效率的方法,展示了在大规模模型上实现1.8倍的吞吐量提升,并通过模拟预测异步RL可实现2.5倍的训练加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24005 2026-04-30 cs.LG cs.AI 62%

TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

TCOD:探索在线蒸馏在多轮自主代理中的时间课程

Jiaqi Wang, Wenhao Zhang, Weijie Shi, Yaliang Li, James Cheng

机构 * Tongyi Lab(通义实验室) Alibaba Group(阿里巴巴集团)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出TCOD框架,通过时间课程策略缓解在线蒸馏在多轮任务中的KL不稳定性,提升代理性能,实验显示性能提升达18个百分点,甚至超越教师模型。

Comments Update code, model weight

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13963 2026-04-30 cs.CL cs.LG 62%

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

通过压缩镜头:探讨量化对事实知识回忆的影响

Qianli Wang, Mingyang Wang, Nils Feldhus, Simon Ostermann, Yuan Cao, Hinrich Schütze, Sebastian Möller, Vera Schmitt

机构 * Quality and Usability Lab, Technische Universität Berlin(柏林技术大学质量与可用性实验室) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Saarland Informatics Campus(萨尔州信息学院) LMU Munich(慕尼黑大学) Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Centre for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心) BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究院) Technical University of Munich(慕尼黑技术大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 研究通过量化技术对大型语言模型的事实知识回忆能力影响,发现量化导致信息损失,但部分模型在低精度下表现更优,BitSandBytes保留最多原始精度的知识回忆能力。

Comments TrustNLP @ ACL 2026; camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00761 2026-04-30 cs.SE cs.AI cs.CL 62%

Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models

找出阿基里斯之踵:一种用于动态揭示大语言模型事实错误的迭代方法

Wenxuan Wang, Yuk-Kit Chan, Zixuan Ling, Juluan Shi, Youliang Yuan, Jen-tse Huang, Yifei Zhang, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu

机构 * Renmin University of China(中国人民大学) The Chinese University of Hong Kong(香港中文大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区) Johns Hopkins University(约翰霍普金斯大学) Nanyang Technological University(南洋理工大学) Xiaohongshu Inc.(小红书公司) Tencent Inc.(腾讯公司)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出HalluHunter框架,通过知识图谱和规则基NLP技术,系统揭示大语言模型的事实错误,测试表明可触发55%的问题错误。

Comments Accepted by Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24179 2026-04-29 cs.CL cs.AI 62%

MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection

MemeScouts@LT-EDI 2026:问对问题——针对弱监督的提示方法用于表情包仇恨言论检测

Ivo Bueno, Lea Hirlimann, Enkelejda Kasneci

机构 * Technical University of Munich(慕尼黑技术大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML))

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种提示弱监督方法,通过分解表情包理解为基于问题的标注函数,提升多语言多模态仇恨言论检测效果,尤其在中文和印地语中表现突出。

Comments Accepted at Sixth Workshop on Language Technology for Equality, Diversity and Inclusion at ACL2026 (LT-EDI@ACL26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25423 2026-04-29 cs.CL cs.AI 62%

Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives

大语言模型能捕捉具身认知和文化差异吗?来自跨语言证据的示范词研究

Yu Wang, Emmanuele Chersoni, Chu-Ren Huang

机构 * Department of Language Science and Technology(语言科学与技术系)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文通过示范词研究探讨大语言模型对具身认知和文化差异的捕捉能力,发现人类在视角转换和距离判断上存在文化差异,而LLMs则缺乏这种理解且默认英语中心化推理。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24715 2026-04-28 cs.CL cs.LG 62%

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling

长上下文意识的升级:混合大语言模型扩展的新前沿

Parsa Ashrafi Fashi, Utkarsh Saxena, Mehdi Rezagholizadeh, Aref Jafari, Akash Haridas, Mingyu Yang, Vansh Bhatia, Guihong Li, Vikram Appia, Emad Barsoum

机构 * AMD

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 本文提出HyLo方法,通过架构适应与高效Transformer块结合,提升长上下文能力,实现上下文长度扩展32倍,内存消耗降低90%,在混合架构中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24647 2026-04-28 cs.CL cs.AI 62%

DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference

DepthKV:面向长上下文LLM推理的分层KV缓存压缩

Zahra Dehghanighobadi, Asja Fischer

机构 * Ruhr University Bochum(博德姆鲁尔大学) UAR Research Center for Trustworthy Data Science and Security(UAR可信数据科学与安全研究中心)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出DepthKV,通过分层敏感度分配优化KV缓存预算,提升长上下文LLM推理效率,优于统一剪枝方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24623 2026-04-28 cs.AI cs.IR cs.LG 62%

XGRAG: A Graph-Native Framework for Explaining KG-based Retrieval-Augmented Generation

XGRAG:一种用于基于知识图谱检索增强生成的图原生解释框架

Zhuoling Li, Ha Linh Hong Tran Nguyen, Valeria Bladinieres, Maxim Romanovsky

机构 * Berlin Technology Centre(柏林技术中心) Deutsche Bank(德意志银行)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 XGRAG通过图基扰动策略生成因果解释,提升GraphRAG系统的可解释性,在多个问答数据集上实现解释质量提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22783 2026-04-28 cs.LG cs.AI 62%

Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation

参数效率不等于内存效率:重新思考设备端LLM适应的微调

Irene Tenison, Stella Ahn, Miriam Kim, Ebtisam Alshehri, Lalana Kagal

机构 * MIT CSAIL(MIT计算机科学与人工智能实验室) Harvard SEAS(哈佛大学科学、工程与应用数学系)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文挑战参数效率等同于内存效率的假设,提出LARS框架通过约束激活子空间降低内存消耗,实验证明在不同设备上均能提升内存效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22781 2026-04-28 cs.LG cs.AI 62%

BiTA: Bidirectional Gated Recurrent Unit-Transformer Aggregator in a Temporal Graph Network Framework for Alert Prediction in Computer Networks

BiTA:一种在时序图网络框架中用于计算机网络警报预测的双向门控循环单元-变换器聚合器

Zahra Makki Nayeri, Mohsen Rezvani

机构 * Faculty of Computer Engineering, Shahrood University of Technology(计算机工程学院,沙赫罗德理工大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出BiTA,通过联合编码双向序列依赖性和长距离上下文关系,改进时序图网络中的时间聚合,提升警报预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏