arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

International Conference on Machine Learning · 会议 · Machine Learning

2026-03-24 至 2026-03-24 共收录 2
2603.20217 2026-03-24 cs.CL cs.LG

Expected Reward Prediction, with Applications to Model Routing

预期奖励预测,及其在模型路由中的应用

Kenan Hasanaliyev, Silas Alberti, Jenny Hamer, Dheeraj Rajagopal, Kevin Robinson, Jasper Snoek, Victor Veitch, Alexander Nicholas D'Amour

机构 * Stanford University Inception Labs(斯坦福大学Inception实验室) Stanford University Cognition Labs(斯坦福大学Cognition实验室) Google DeepMind(谷歌DeepMind) University of Chicago(芝加哥大学)

AI总结 本文研究了基于响应级奖励模型预测模型对特定提示的适应性,提出了一种简单有效的预期奖励预测路由方法,实验证明其在模型路由中的优越性。

Comments ICML 2025 Workshop on Models of Human Feedback for AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14038 2026-03-24 cs.LG

Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning

滑动拼图健身房:一种可扩展的用于视觉强化学习状态表示的基准

Bryan L. M. de Oliveira, Luana G. B. Martins, Bruno Brandão, Murilo L. da Luz, Telma W. de L. Soares, Luckeciano C. Melo

机构 * Advanced Knowledge Center for Immersive Technologies -- AKCIT, Brazil(沉浸式技术高级知识中心 -- AKCIT,巴西) OATML, University of Oxford, United Kingdom(OATML,牛津大学,英国) Institute of Informatics, Federal University of Goiás, Goiânia, Brazil(信息学院,戈亚尼亚联邦大学,巴西)

AI总结 本文提出Sliding Puzzles Gym,通过调整网格大小和图像池来控制视觉表示复杂度,评估视觉表示学习能力,发现现有方法在处理视觉多样性时存在局限。

Comments Accepted at ICML 2025

Journal ref Proceedings of Machine Learning Research 267:12689-12717, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏