GameTalk: Training LLMs for Strategic Conversation
GameTalk: 训练 LLMs 进行战略性对话
机构 * University of Cambridge(剑桥大学)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 GameTalk 通过多轮互动训练 LLMs 实现战略性决策,优于传统方法,尤其在奖励塑造下表现突出。
Comments 32 pages, 8 figures
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
GameTalk: 训练 LLMs 进行战略性对话
机构 * University of Cambridge(剑桥大学)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 GameTalk 通过多轮互动训练 LLMs 实现战略性决策,优于传统方法,尤其在奖励塑造下表现突出。
Comments 32 pages, 8 figures
在数据稀缺条件下构建可执行领域特定大语言模型的通用框架:以半导体TCAD仿真为例
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
AI总结 本文提出一种通用框架,在数据稀缺条件下构建可执行的领域特定大语言模型,通过生成合成数据和代码优化流程,实现领域知识灌输和脚本可执行性提升。
Comments Submitted to Nature Computational Science
DB3团队的Meta KDD杯'25解决方案
机构 * Key Laboratory of High Confidence Software Technologies, CS, Peking University, China(高可信软件技术重点实验室,中国科学院,北京大学) ; Theory Lab, Central Research Institute, 2012 Labs, Huawei Technologies Co., Ltd, Shenzhen, China(理论实验室,中央研究院,2012实验室,华为技术有限公司,深圳)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 DB3团队通过多模态检索框架和拒绝训练方法,在KDD杯'25的CRAG-MM挑战中取得优异成绩,尤其在处理第一人称视角问题上表现突出。
SimLLM:用于基于SimPy的排队系统模拟的代码LLM微调
机构 * Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院) ; School of Information, Renmin University of China(中国人民大学信息学院) ; School of Management and Economics, University of Electronic Science and Technology of China(电子科技大学管理学院)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 SimLLM通过微调开源LLM提升SimPy排队系统模拟代码生成能力,提供替代闭源模型的可行方案。
Comments 33 pages, 10 figures
稳定偏好优化:一种针对灾难性偏好转移的双层方法
机构 * Tongji University(同济大学) ; AsiaInfo Technologies(亚洲信息科技)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出稳定偏好优化框架,通过约束偏好学习在安全对齐区域,解决灾难性偏好转移问题,提升偏好学习方法的稳定性和性能。
ReGal:基于PPO的法律AI在印度判决预测和摘要中的初步探索
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 ReGal通过PPO结合多任务指令微调与强化学习,探索法律AI在印度判决预测和摘要中的应用,揭示了强化学习在法律文本中的挑战与潜力。
Comments Accepted in AILaw @ AAAI 2026 conference
利用负信号:从教师数据中进行强化蒸馏以提升大语言模型推理能力
机构 * National University of Singapore(国立新加坡大学) ; INF AI
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出通过强化蒸馏利用正负推理轨迹提升LLM推理性能,实验表明在少量数据下可达到与大量数据训练模型相当的性能。
Comments 22 pages, 10 figures. Code available at https://github.com/Tim-Siu/reinforcement-distillation
Diffusion-SDPO:用于扩散模型的受保护直接偏好优化
机构 * School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) ; National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学) ; Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
AI总结 Diffusion-SDPO通过保护性更新规则提升扩散模型对人类偏好的对齐效果。
Comments The code is publicly available at https://github.com/AIDC-AI/Diffusion-SDPO
弥合差距:面向标准化考试题目的视觉语言模型数据驱动微调
机构 * organization= Department of Computer Engineering, Middle East Technical University (METU) , city= Ankara , country= Türkiye ; organization= METU-DTX Digital Transformation \& Innovation Centre, METU , city= Ankara , country= Türkiye
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 本研究通过高质量数据和优化语法提升视觉语言模型在标准化考试题目的多模态推理性能,达到接近SOTA的水平。
EAGER: 边缘对齐的LLM防御方法用于鲁棒、高效且准确的网络安全问答
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
AI总结 EAGER通过参数高效量化与领域偏好对齐,提升网络安全问答的鲁棒性、效率和准确性,降低对抗攻击成功率并优化响应延迟。
机构 * Department of Computer Science, University of New Hampshire(新罕布什尔大学计算机科学系) ; Materials Department, University of California, Santa Barbara(加州大学圣芭芭拉分校材料系) ; Department of Mechanical Engineering, University of California, Santa Barbara(加州大学圣芭芭拉分校机械工程系) ; Allen Institute for Artificial Intelligence(人工智能研究院)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥)) ; IBM(IBM公司)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Appears in AAAI 2026 in the Main Technical Track
机构 * Google DeepMind(谷歌DeepMind)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * BAISH | UBA | Apart Research(BAISH | UBA | Apart研究) ; University of São Paulo(圣保罗大学) ; Apart Research(Apart研究) ; Dovetail Research | Apart Research(Dovetail研究 | Apart研究)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Stanford University(斯坦福大学) ; University of Toronto(多伦多大学) ; University of Pennsylvania(宾夕法尼亚大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025
机构 * ETH Zürich(苏黎世联邦理工学院)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025
机构 * NVIDIA
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025 Datasets and Benchmarks Track Camera Ready, 46 pages, 2 figures
机构 * University of Maryland, College Park(马里兰大学学院公园分校)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments neurips cam-ready
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
机构 * Snap Research ; University of Toronto(多伦多大学) ; Vector Institute(向量研究所)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
Comments NeurIPS 2025 Spotlight. Project page: https://snap-research.github.io/DenseDPO/
机构 * Ramaiah Institute of Technology(拉马亚院技术学院) ; Ecofy Finance Private Limited(Ecofy金融私人有限公司)
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * University of Maryland, College Park(马里兰大学学院 park)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Google DeepMind(谷歌DeepMind)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to COLM 2025. Camera-ready version
机构 * Tsinghua University(清华大学)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
Comments ICCV 2025
机构 * University of California, Davis(加州大学戴维斯分校) ; WeChat, Tencent(微信、腾讯) ; Tsinghua University(清华大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
机构 * Boston University(波士顿大学) ; MIT(麻省理工学院) ; Monash University Indonesia(墨尔本大学印度尼西亚分校)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages, EMNLP PALS workshop 2025
机构 * Department of Statistics and Data Science, The Chinese University of Hong Kong(统计与数据科学系,香港中文大学) ; Department of Statistics, University of Wisconsin-Madison(统计系,威斯康星大学麦迪逊分校)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG