arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Oxford(牛津大学)

2026-04-22 至 2026-04-22 共收录 4
2604.19548 2026-04-22 cs.CL cs.AI cs.CY

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

通过辩证对齐平息代理中的行动者-观察者不对称性

Bobo Li, Rui Wu, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(国立新加坡大学) Sichuan University(四川大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) University of Oxford(牛津大学)

AI总结 本文提出ReTAS模型,通过辩证对齐方法减少代理中的行动者-观察者不对称性,提升故障解决能力。

Comments ACL 2026 Main Conference. Project page: https://unikcc.github.io/ReTAS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26238 2026-04-22 cs.LG

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

超越线性探测:语言模型的动态安全监控

James Oldfield, Philip Torr, Ioannis Patras, Adel Bibi, Fazl Barez

机构 * Queen Mary University of London(伦敦玛丽女王大学) University of Oxford(牛津大学) WhiteBox Martian

AI总结 本文提出Truncated Polynomial Classifiers (TPCs),通过动态激活监控提升安全性能,实现灵活的成本控制,在有害提示分类中表现优异且更具可解释性。

Comments ICLR 2026; Minor revisions and clarifications

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21184 2026-04-22 cs.CL cs.AI stat.ML

BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design

BED-LLM:利用LLM和贝叶斯实验设计进行智能信息收集

Deepro Choudhury, Sinead Williamson, Adam Goliński, Ning Miao, Freddie Bickford Smith, Michael Kirchhof, Yizhe Zhang, Tom Rainforth

机构 * University of Oxford(牛津大学) Apple(苹果公司) City University of Hong Kong(香港城市大学) University of Somewhere(某大学) Institute of Something(某研究所) Company Ltd(某有限公司)

AI总结 本文提出了一种基于贝叶斯实验设计框架的通用方法,通过迭代选择最大化信息增益的问题来提升LLM的信息收集能力,展示了BED-LLM在20 Questions游戏中优于传统提示方法的性能。

Comments Published at the International Conference on Learning Representations 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01687 2026-04-22 cs.CL

StochasTok: Improving Fine-Grained Subword Understanding in LLMs

StochasTok:提升大语言模型的细粒度子词理解

Anya Sims, Thom Foster, Klara Kaleb, Tuan-Duy H. Nguyen, Joseph Lee, Jakob N. Foerster, Yee Whye Teh, Cong Lu

机构 * FLAIR, University of Oxford(FLAIR,牛津大学) University of Oxford(牛津大学) National University of Singapore(新加坡国立大学) University of British Columbia(不列颠哥伦比亚大学)

AI总结 本文提出StochasTok,一种高效的随机分词方法,通过在训练中随机分割token,提升LLM在子词层面的任务表现,包括字符计数和数学任务。

详情

展开后加载摘要…

URL PDF HTML 收藏