arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Transactions on Machine Learning Research · 期刊 · Machine Learning

2026-07-28 至 2026-07-28 共收录 3
2607.24057 2026-07-28 cs.LG 新提交

Constrained Reinforcement Learning Using Successor Representations

使用后继表示的约束强化学习

Michael Girstl, Alexander Mattick, Christopher Mutschler

机构 * Technical University of Darmstadt (TU Darmstadt)(达姆施塔特工业大学) Hessian Center for Artificial Intelligence (hessian.AI)(黑森州人工智能中心) Fraunhofer Institute for Integrated Circuits IIS, Fraunhofer IIS(弗劳恩霍夫集成电路研究所IIS) University of Technology Nuremberg (UTN)(纽伦堡工业大学)

AI总结 研究针对现实世界强化学习中策略难适应成本函数变化的问题,提出SafeDSR方法,通过引入可学习权重矩阵扩展深度后继表示到约束强化学习,能快速重训策略,在二维导航环境展示竞争力与灵活性。

Comments published in Transactions for Machine Learning Research 2026

Journal ref Michael Girstl, Alexander Mattick, & Christopher Mutschler (2026). Constrained Reinforcement Learning Using Successor Representations. Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00918 2026-07-28 cs.CL cs.AI cs.LG 版本更新

LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models

LIBMoE:一种用于大规模语言模型混合专家全面基准测试的库

Nam V. Nguyen, Thong T. Doan, Luong Tran, Van Nguyen, Quang Pham

AI总结 LibMoE为大规模语言模型混合专家提供统一框架,通过全面分析路由动态、初始化影响及训练模式差异,推动MoE研究标准化与创新。

Comments 40 pages

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05776 2026-07-28 stat.ML cs.LG

Dynamic Pricing in the Linear Valuation Model using Shape Constraints

在线性估值模型中使用形状约束进行动态定价

Daniele Bracale, Moulinath Banerjee, Yuekai Sun, Kevin Stoll, Salam Turki

机构 * University of Michigan(密歇根大学) Welltower Inc.(Welltower公司)

AI总结 本文提出了一种基于形状约束的动态定价方法,在线性估值模型中无需调节参数,通过α-霍尔德连续性假设推导后悔上界,并在实验中展示更低的后悔表现。

Journal ref Transactions on Machine Learning Research (TMLR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏