arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-24 至 2026-04-24 共收录 9 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 9 篇

2604.21223 2026-04-24 cs.CL cs.AI 92%

Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model

通过隐式奖励模型实现LLM生成文本的零样本检测

Runheng Liu, Heyan Huang, Xingchen Xiao, Zhijing Wu

机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)

专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出IRM方法,利用隐式奖励模型实现LLM生成文本的零样本检测,无需偏好收集或额外训练,在DetectRL基准上表现优于现有方法。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21209 2026-04-24 cs.AI cs.CL 88%

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management

对齐生成式人工智能与人类偏好:一种用于在线评论管理的新型大语言模型微调方法

Yanan Wang, Yong Ge

机构 * Department of Information Systems and Operations Management, The University of Texas at Arlington(信息系统与运营管理系,德克萨斯大学阿灵顿分校) Department of Management Information Systems, University of Arizona(管理信息学系,亚利桑那大学)

专题命中 后训练与偏好优化 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出一种新型微调方法,通过上下文增强和理论驱动的偏好微调,解决在线评论生成中的幻觉、偏好表示和保守优化问题,提升生成质量。

Comments Accepted to Information Systems Research (ISR). This is a preliminary version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00931 2026-04-24 cs.LG cs.AI 86%

Continuous-Utility Direct Preference Optimization

连续效用直接偏好优化

Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal, Zihao He, Muhammad Usman Rafique, Asad Aali, Muhammad Ali Jamshed, John M. Cioffi, Emily Fox

机构 * stanford(斯坦福大学) meta glasgow(格拉斯哥大学)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出CU-DPO框架,通过连续评分替代二元标签,提升模型在数学推理任务中的策略选择准确率和推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20904 2026-04-24 cs.LG cs.AI 82%

Reinforcing privacy reasoning in LLMs via normative simulacra from fiction

通过小说中的规范仿像强化大语言模型的隐私推理

Matt Franchi, Madiha Zahrah Choksi, Harold Triedman, Helen Nissenbaum

机构 * Cornell Tech(康奈尔科技)

专题命中 后训练与偏好优化 :LLM(abstract,abstract_cn);SFT(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本文通过从小说中提取规范仿像,结合监督学习和GRPO强化学习,提升大语言模型的隐私推理能力,验证了小说来源的规范仿像在现实领域中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05591 2026-04-24 cs.LG cs.CL 79%

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning

熵比裁剪作为稳定强化学习的软全局约束

Zhenpeng Su, Leiyu Pan, Minxuan Lv, Tiehua Mei, Zijia Lin, Yuntao Li, Wenping Hu, Ruiming Tang, Kun Gai, Guorui Zhou

机构 * Kuaishou Technology(快手科技) Tsinghua University(清华大学) Independent(独立)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.CL、cs.LG

AI总结 本文提出熵比裁剪机制,通过量化策略探索的相对变化,缓解强化学习中的分布偏移问题,提升训练稳定性。

Comments This paper has been accepted by ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03048 2026-04-24 cs.AI cs.CY cs.LG cs.MA 73%

The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment

规范陷阱:为何仅靠静态价值对齐无法实现稳健对齐

Austin Spizzirri

机构 * Belmont University(贝尔蒙特大学)

专题命中 后训练与偏好优化 :RLHF(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本文指出静态内容导向的人工智能价值对齐在能力扩展、分布偏移和自主性提升时无法实现稳健对齐,探讨了哲学难题及现有方法的结构性漏洞。

Comments 31 pages, no figures. Version 5. First posted as arXiv:2512.03048 in November 2025. First in a six-paper research program on AI alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20712 2026-04-24 cs.LG cs.CL 73%

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning

CE-GPPO:通过梯度保持剪裁策略优化协调熵在强化学习中

Zhenpeng Su, Leiyu Pan, Minxuan Lv, Yuntao Li, Wenping Hu, Fuzheng Zhang, Kun Gai, Guorui Zhou

机构 * Kuaishou Technology(快手科技) Independent(独立)

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出CE-GPPO算法,通过恢复剪裁令牌的梯度信号,协调探索与利用,缓解熵不稳定问题,并在数学推理任务中优于现有方法。

Comments This paper has been accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03248 2026-04-24 cs.CL 70%

STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning

STReasoner: 通过空间感知强化学习赋能LLMs进行时间序列的时空推理

Juntong Ni, Shiyu Wang, Qi He, Ming Jin, Wei Jin

机构 * Emory University(埃默里大学) Microsoft(微软) Griffith University(格里菲斯大学)

专题命中 后训练与偏好优化 :LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出STReasoner,通过整合时间序列、图结构和文本,提升LLMs的时空推理能力,采用S-GRPO算法增强空间逻辑,实验显示在低成本下取得显著准确率提升。

Comments ACL 2026 Main, we release our code publicly at https://github.com/LingFengGold/STReasoner

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21587 2026-04-24 cs.IT math.IT 50%

Generative Learning Enhanced Intelligent Resource Management for Cell-Free Delay Deterministic Communications

生成学习增强的智能资源管理用于无基站确定性通信

Shuangbo Xiong, Cheng Zhang, Wen Wang, Wenwu Yu, Yongming Huang

专题命中 后训练与偏好优化 :pretraining(abstract)

AI总结 本文研究了无基站MIMO系统中的资源分配问题,提出基于虚拟约束马尔可夫决策过程的离线预训练框架,通过证据感知条件高斯混合模型提升效率和安全性,实验表明其在能效和延迟约束方面表现优异。

Comments The paper has been submitted to IEEE Transactions on Wireless Communications

详情

展开后加载摘要…

URL PDF HTML 收藏