arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4535 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 4535 篇

2603.11901 2026-03-13 cs.LG 83%

FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning

FlexRec: 通过强化学习适应LLM推荐系统以满足灵活需求

Yijun Pan, Weikang Qiu, Qiyao Ma, Mingxuan Ju, Tong Zhao, Neil Shah, Rex Ying

专题命中 后训练与偏好优化 :LLM(title,abstract);post-training(abstract);分类 cs.LG

AI总结 FlexRec通过强化学习框架解决LLM推荐系统在需求特定排序和泛化设置中的挑战,提升NDCG@5和Recall@5性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06444 2026-03-09 cs.SD cs.AI 83%

Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input

面向流式文本输入的语调边界感知流式生成用于基于大语言模型的语音合成

Changsong Liu, Tianrui Wang, Ye Ni, Yizhou Peng, Eng Siong Chng

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) Tianjin University, China(天津大学) Southeast University, China(东南大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);post-training(abstract);分类 cs.AI

AI总结 本文提出一种面向语调边界的后训练策略,通过调整预训练模型以处理流式文本输入,从而提升语音合成的语调质量和长文本生成的稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05275 2026-03-06 cs.MM cs.CL cs.SD 83%

SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning

SarcasmMiner:一种双轨后训练框架,用于鲁棒的音频视觉讽刺推理

Zhu Li, Yongjian Chen, Huiyuan Lai, Xiyuan Gao, Shekhar Nayak, Matt Coler

机构 * Speech Technology Lab, University of Groningen, The Netherlands(格罗宁根大学语音技术实验室,荷兰) Center for Language and Cognition, University of Groningen, The Netherlands(格罗宁根大学语言与认知中心,荷兰)

专题命中 后训练与偏好优化 :post-training(title,abstract);foundation model(abstract);分类 cs.CL

AI总结 SarcasmMiner通过双轨蒸馏和组相对策略优化提升多模态讽刺检测性能,实现F1值显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21461 2026-02-26 cs.CL 83%

VecGlypher: Unified Vector Glyph Generation with Language Models

VecGlypher: 一种基于语言模型的统一向量图形单元生成方法

Xiaoke Huang, Bhavul Gauri, Kam Woh Ng, Tony Ng, Mengmeng Xu, Zhiheng Liu, Weiming Ren, Zhaochong An, Zijian Zhou, Haonan Qiu, Yuyin Zhou, Sen He, Ziheng Wang, Tao Xiang, Xiao Han

机构 * Meta AI

专题命中 后训练与偏好优化 :language model(title,abstract);post-training(abstract);分类 cs.CL

AI总结 VecGlypher通过多模态语言模型直接生成高质量向量图形,无需位图中间步骤,显著提升字体创建效率和可编辑性。

Comments Accepted to CVPR'26. Project page: https://xk-huang.github.io/VecGlypher/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14762 2026-02-20 cs.LG cs.CV 83%

Unlocking [CLS] Features for Continual Post-Training

解锁[CLS]特征以实现持续微调

Murat Onur Yildirim, Elif Ceren Gok Yildirim, Joaquin Vanschoren

专题命中 后训练与偏好优化 :post-training(title,abstract);foundation model(abstract);分类 cs.LG

AI总结 本文提出TOSCA方法,通过在[CLS]标记上部署稀疏LuCA模块,实现持续学习中稳定性与可塑性的平衡,减少参数量并提升性能。

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11596 2026-02-13 cs.AI 83%

MAPLE: Modality-Aware Post-training and Learning Ecosystem

MAPLE: 多模态感知的后训练与学习生态系统

Nikhil Verma, Minjung Kim, JooYoung Yoo, Kyung-Min Jin, Manasa Bharadwaj, Kevin Ferreira, Ko Keun Kim, Youngjoon Kim

机构 * LG Electronics Toronto AI Lab, Toronto, Canada(LG电子多伦多AI实验室) LG Electronics CTO AI Lab, Seoul, Republic of Korea(LG电子CTO AI实验室)

专题命中 后训练与偏好优化 :post-training(title,abstract);language model(abstract);分类 cs.AI

AI总结 MAPLE通过模态感知的后训练与学习生态系统,提升多模态强化学习的准确性、收敛速度和稳定性。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05910 2026-02-06 cs.LG 83%

Chunky Post-Training: Data Driven Failures of Generalization

块状后训练:数据驱动的泛化失败

Seoirse Murray, Allison Qi, Timothy Qian, John Schulman, Collin Burns, Sara Price

机构 * Thinking Machines Lab(思维机器实验室)

专题命中 后训练与偏好优化 :post-training(title,abstract);LLM(abstract);分类 cs.LG

AI总结 研究发现大语言模型在后训练过程中因数据块的不平衡或不明确导致泛化失败,提出SURF和TURF工具用于检测和追溯这些问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01581 2026-02-03 cs.LG 83%

Nearly Optimal Active Preference Learning and Its Application to LLM Alignment

近优主动偏好学习及其在大语言模型对齐中的应用

Yao Zhao, Kwang-Sung Jun

机构 * University of Arizona(亚利桑那大学) Department of Computer Science(计算机科学系)

专题命中 后训练与偏好优化 :LLM(title);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出两种主动学习算法,通过实例依赖标签复杂性保证和贪心方法提升大语言模型对齐的样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00603 2026-02-03 cs.LG 83%

Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains

直接偏好优化与评分信息:实用算法和可证明的收益

Luca Viano, Ruida Zhou, Yifan Sun, Mahdi Namazifar, Volkan Cevher, Shoham Sabach, Mohammad Ghavamzadeh

机构 * EPFL(苏黎世联邦理工学院) Amazon AGI(亚马逊人工智能实验室) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Cornell University(康奈尔大学) Qualcomm AI Research(高通人工智能研究)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);foundation model(abstract);分类 cs.LG

AI总结 本文提出利用评分间隙信息改进直接偏好优化算法,通过理论证明和实验验证,在准确评分间隙条件下实现更快的统计速率,并在多种LLM和基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23129 2026-02-02 cs.CL 83%

Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics

评估基于参考-free LLM的文档接地的效用

Yilun Hua, Giuseppe Castellucci, Peter Schulam, Heba Elfardy, Kevin Small

机构 * Department of Computer Science and Cornell Tech, Cornell University(计算机科学系和Cornell Tech,康奈尔大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);preference optimization(abstract);分类 cs.CL

AI总结 本文提出GroGU,一种无需注释的模型特定指标,用于评估RAG中文档接地的效用,通过优化查询重写器提升了检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16245 2026-02-02 cs.CL 83%

Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models

多样而非简短:一种长度控制的数据选择策略以提高语言模型的响应多样性

Vijeta Deshpande, Debasmita Ghose, John D. Patterson, Roger Beaty, Anna Rumshisky

机构 * University of Massachusetts Lowell(马萨诸塞大学洛维尔分校) Yale University(耶鲁大学) Pennsylvania State University(宾夕法尼亚州立大学) Amazon AGI(亚马逊人工智能研究院)

专题命中 后训练与偏好优化 :language model(title,abstract);preference optimization(abstract);分类 cs.CL

AI总结 Diverse-NS通过长度控制的数据选择策略提升语言模型响应多样性,适用于创造性生成任务,并在不同规模模型间展示出显著效果。

Comments Accepted to EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15596 2026-01-30 cs.LG 83%

Corrective Diffusion Language Models

纠正扩散语言模型

Shuibai Zhang, Fred Zhangzhi Peng, Yiheng Zhang, Jin Pan, Grigorios G. Chrysos

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Duke University(杜克大学)

专题命中 后训练与偏好优化 :language model(title,abstract);post-training(abstract);分类 cs.LG

AI总结 纠正扩散语言模型通过明确监督可见错误token,实现判别性置信度和针对性细化,显著提升代码修订和并行解码任务性能。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20966 2026-01-30 cs.RO cs.AI 83%

Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends

VLA模型后训练与人类运动学习的类比:进展、挑战与趋势

Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Shuang-Yi Wang, Sheng-Bin Duan, Fu-Chao Xie, Wen-Kai Wang, Si-Cheng Wang, Ling-Yun Li, Tian Tu, Zeng-Guang Hou

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) The Grainger College of Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格拉inger工程学院) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) Joint Laboratory of Intelligence Science and Technology, Institute of Systems Engineering, Macau University of Science and Technology(澳门科技大学系统工程学院智能科学与技术联合实验室)

专题命中 后训练与偏好优化 :post-training(title,abstract);language model(abstract);分类 cs.AI

AI总结 本文从人类运动学习角度综述了VLA模型后训练的进展、挑战与趋势,提出四类后训练方法并探讨了其在机器人操作中的应用与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19780 2026-01-22 cs.LG 83%

Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing

基于多维偏好混合的列表式直接偏好优化

Yuhui Sun, Xiyao Wang, Zixi Li, YiTian Ding, Tianyang Ling, Jialuo Chen, Tianyi Yu, Zhenlong Yuan, Jinman Zhao

机构 * University of Alberta(阿尔伯塔大学) University of Toronto(多伦多大学) Zhejiang University(浙江大学) School of Computer Science McGill University(麦吉尔大学计算机学院) Alibaba Group (Ant Group)(阿里巴巴集团(蚂蚁集团)) Chinese Academy of Sciences(中国科学院)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);RLHF(abstract);分类 cs.LG

AI总结 本文提出λ-DPO,通过多维偏好混合和自适应调度器提升模型对多维度偏好权衡的建模能力。

Comments 13 pages, 1 figures, appendix included

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21656 2026-01-06 cs.CV cs.CL 83%

Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs

细粒度偏好优化提升视觉语言模型的空间推理能力

Yifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao, Xiaofeng Zhang, Yixiao He, Wenming Ye, James Matthew Rehg, Ismini Lourentzou

专题命中 后训练与偏好优化 :preference optimization(title,abstract);language model(abstract);分类 cs.CL

AI总结 细粒度偏好优化方法SpatialReasoner-R1通过改进空间推理能力,在空间推理任务中取得显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06118 2025-12-23 cs.CV cs.AI 83%

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

ViGoR:通过细粒度奖励建模提升大视觉语言模型的视觉 grounding

Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, Li Erran Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) AWS AI(AWS人工智能)

专题命中 后训练与偏好优化 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 ViGoR通过细粒度奖励建模提升大视觉语言模型的视觉 grounding 能力,采用更经济的人类评估和自动化方法,有效提高视觉推理准确性。

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09485 2025-12-11 cs.CR cs.AI 83%

Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks

通过定制化群体相对策略优化推进基于大语言模型的安全自动化以实现零接触网络

Xinye Cao, Yihan Lin, Guoshun Nan, Qinchuan Zhou, Yuhang Luo, Yurui Gao, Zeliang Zhang, Haolang Lu, Qimei Cui, Yanzhao Hou, Xiaofeng Tao, Tony Q. S. Quek

机构 * National Engineering Research Center for Mobile Network Technologies, Beijing University of Posts and Telecommunications(移动网络技术国家工程研究中心,北京邮电大学) Beijing University of Posts and Telecommunications(北京邮电大学) Singapore University of Technology and Design(新加坡科技设计大学) Department of Electronic Engineering, Kyung Hee University(韩国庆熙大学电子工程系)

专题命中 后训练与偏好优化 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出SecLoop和SA-GRPO,通过定制化群体相对策略优化,解决6G零接触网络中的安全自动化挑战。

Comments Accepted by IEEE JSAC. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06920 2025-12-09 cs.LG 83%

Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models

基于嵌入的语义奖励模型(PGSRM):用于变换器语言模型强化学习的奖励函数

Alexandr Plashchinsky

机构 * VECTOR Labs(VECTOR实验室)

专题命中 后训练与偏好优化 :language model(title,abstract);RLHF(abstract);分类 cs.LG

AI总结 PGSRM通过嵌入相似度生成语义奖励,为变换器语言模型强化学习提供无需人工标注的替代方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12934 2025-12-02 cs.AI 83%

The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features

对齐的解剖:通过引导稀疏特征分解偏好优化

Jeremias Ferrao, Matthijs van der Lende, Ilija Lichkovski, Clement Neo

机构 * University of Groningen(格罗宁根大学) AI Safety Initiative Groningen(格罗宁根人工智能安全倡议) Apart Research(Apart研究)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);post-training(abstract);分类 cs.AI

AI总结 本文提出FSRL框架,通过引导稀疏特征来实现透明的偏好优化,揭示模型在风格化呈现上的偏倚,提供可解释的对齐控制方法。

Comments Spotlight at NeurIPS 2025 Mechanistic Interpretability Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23391 2025-12-01 cs.CL 83%

Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization

模糊性意识优化:朝着直接偏好优化的语义消歧

Jian Li, Shenglin Yin, Yujia Zhang, Alan Zhao, Xi Chen, Xiaohui Zhou, Pengfei Xu

机构 * AI Technology Center of OVB, Tencent, China(腾讯OVB人工智能技术中心,中国) School of Computer Science, Peking University, China(北京大学计算机学院,中国)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);RLHF(abstract);分类 cs.CL

AI总结 本文提出模糊性意识优化(AAO)方法,通过计算偏好对的语义相似性自动重新加权模糊内容,有效提升直接偏好优化的性能。

Comments Accepted at EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18212 2025-11-18 cs.CL 83%

Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries

Meiling Ning, Zhongbao Zhang, Junda Ye, Jiabao Guo, Qingyuan Guan

专题命中 后训练与偏好优化 :language model(title,abstract);RLHF(abstract);分类 cs.CL

Comments After further internal discussion, our author team has decided to withdraw this submission due to the need for several important refinements to the manuscript. All co-authors have been informed and agree with this decision

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10289 2025-11-14 eess.AS cs.CL 83%

Music Flamingo: Scaling Music Understanding in Audio Language Models

Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee, Zhifeng Kong, Joao Felipe Santos, Ramani Duraiswami, Dinesh Manocha, Wei Ping, Mohammad Shoeybi, Bryan Catanzaro

机构 * NVIDIA, CA, USA(NVIDIA公司) University of Maryland, College Park, USA(大学)

专题命中 后训练与偏好优化 :language model(title,abstract);post-training(abstract);分类 cs.CL

Comments Project Page: https://research.nvidia.com/labs/adlr/MF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08039 2025-11-05 cs.SD cs.CL cs.MM eess.AS 83%

Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning

Shu Wu, Chenxing Li, Wenfu Wang, Hao Zhang, Hualei Wang, Meng Yu, Dong Yu

专题命中 后训练与偏好优化 :language model(title,abstract);large language model(abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07127 2025-10-31 cs.RO cs.AI 83%

Human-assisted Robotic Policy Refinement via Action Preference Optimization

Wenke Xia, Yichu Yang, Hongtao Wu, Xiao Ma, Tao Kong, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing(中国人民大学北京校区人工智能学院) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索与推荐工程研究中心) Beijing Key Laboratory of Research on Large Models(北京大型模型研究重点实验室)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);foundation model(abstract);分类 cs.AI

Comments Accepted By NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14755 2025-10-30 cs.DC cs.AI 83%

Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models

Daoyuan Chen, Yilun Huang, Xuchen Pan, Nana Jiang, Haibin Wang, Yilei Zhang, Ce Ge, Yushuo Chen, Wenhao Zhang, Zhijian Ma, Jun Huang, Wei Lin, Yaliang Li, Bolin Ding, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

专题命中 后训练与偏好优化 :foundation model(title,abstract);post-training(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 (Spotlight). 43 pages, 16 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01381 2025-10-29 cs.CL 83%

AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation

Yilong Lai, Jialong Wu, Zhenglin Wang, Deyu Zhou

机构 * School of Computer Science and Engineering, Key Laboratory of Computer Network and Information Integration, Ministry of Education, Southeast University(计算机科学与工程学院、计算机网络与信息集成重点实验室、教育部、东南大学)

专题命中 后训练与偏好优化 :prompting(title,abstract);LLM(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04272 2025-10-14 cs.LG 83%

Understanding the Impact of Sampling Quality in Direct Preference Optimization

Kyung Rok Kim, Yumo Bai, Chonghuan Wang, Guanting Chen

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);RLHF(abstract);分类 cs.LG

Comments Submitted to AISTATS2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04089 2025-10-07 cs.AI 83%

SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows

Yitong Cui, Liu Liu, Baosheng Yu, Jiayan Qiu, Xikai Zhang, Likang Xiao, Yixing Liu, Quan Chen

机构 * Hangzhou International Innovation Institute and , Beihang University(杭州国际创新研究院和 北航) School of Artificial Intelligence, Beihang University(人工智能学院,北航) Nanyang Technological University(南洋理工大学) University of Leicester(莱斯特大学) China Mobile Communications Company Limited Research Institute(中国移动通信有限公司研究院)

专题命中 后训练与偏好优化 :preference optimization(title);large language model(abstract);language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07510 2025-10-07 cs.CV cs.LG 83%

Divergence Minimization Preference Optimization for Diffusion Model Alignment

Binxu Li, Minkai Xu, Jiaqi Han, Meihua Dang, Stefano Ermon

机构 * Stanford University(斯坦福大学)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20384 2025-09-26 cs.CR cs.AI cs.PL cs.SE 83%

R1-Fuzz: Specializing Language Models for Textual Fuzzing via Reinforcement Learning

Jiayi Lin, Liangcai Su, Junzhe Li, Chenxiong Qian

机构 * The University of Hong Kong(香港大学)

专题命中 后训练与偏好优化 :language model(title,abstract);post-training(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏