arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Oxford(牛津大学)

共收录 186
2606.31085 2026-07-01 cs.AI 新提交

DDIAgents: Mechanism-Conditioned Context Flow for Drug-Drug Interaction Prediction

DDIAgents: 机制条件上下文流用于药物相互作用预测

Zhenqian Shen, Yu Liu, Xiaoyi Fu, Quanming Yao

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Institute of Biomedical Engineering, University of Oxford(牛津大学生物医学工程研究所) Division of Emerging Interdisciplinary Areas, Hong Kong University of Science and Technology(香港科技大学新兴跨学科领域学部)

AI总结 提出DDIAgents多智能体框架,通过动态知识编排预测药物相互作用,减少无关信息并生成可解释推理,性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17882 2026-07-01 cs.AI 新提交

Structural Preservation and the Logical Expressiveness of Graph Neural Networks

结构保持与图神经网络的逻辑表达能力

Przemysław Andrzej Wałęga, Bernardo Cuenca Grau

机构 * Queen Mary University of London(伦敦玛丽女王大学) University of Oxford(牛津大学)

AI总结 本文从语义角度研究图神经网络分类器在结构保持(嵌入、单同态、同态)下的逻辑表达能力,证明每种保持性质对应分级模态逻辑的一个片段,并给出相应GNN架构。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28204 2026-06-29 cs.GT cs.LG 新提交

Non-Linear Strategic Classification Made Practical

非线性的策略分类变得实用

Jack Geary, Boyan Gao, Henry Gouk

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院) Department of Engineering Science, University of Oxford(牛津大学工程科学系) School of Informatics University of Edinburgh(爱丁堡大学信息学院)

AI总结 针对策略分类中非线性分类器难以处理的问题,提出利用拉格朗日对偶近似最佳响应,结合隐函数定理计算损失梯度,实现非线性策略分类的实用训练算法。

Comments 15 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27460 2026-06-29 cs.CL 新提交

Developmental approach reveals the statistical learning of Neural Language Models: Transformers generalize from the most abstract statistical patterns

发展方法揭示神经语言模型的统计学习:Transformer从最抽象的统计模式中泛化

Wang Bojun, Holly Jenkins, Elizabeth Wonnacott

机构 * Department of Education, University of Oxford(牛津大学教育学院)

AI总结 本研究采用发展方法,通过训练生成式Transformer模型并分析其内部表征变化,发现神经语言模型先习得最抽象的全局统计知识,后习得局部统计依赖,并提出新的统计学习与语言认知框架。

Comments 10 pages, 7 figures, oral presentation at Interdisciplinary Advances in Statistical Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26236 2026-06-26 eess.IV cs.CV eess.SP 新提交

Rendering Novel Views of MRI Using 3D Gaussian Splatting

使用3D高斯泼溅渲染MRI的新视图

Robin Y. Park, Mark C. Eid, Rhydian Windsor, Amir Jamaludin, Ana I. L. Namburete, João F. Henriques, Andrew Zisserman

机构 * Visual Geometry Group, University of Oxford(视觉几何组,牛津大学) Oxford Machine Learning in NeuroImaging Lab, University of Oxford(牛津大学神经影像学机器学习实验室)

AI总结 通过3D高斯泼溅从稀疏各向异性MRI重建体积并采样与目标解剖对齐的平面,生成更优的放射学分级视图,相比原始扫描和体素插值方法提高了椎管狭窄分级准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27147 2026-06-26 cs.CV cs.AI 新提交

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

安全自回归图像生成与迭代自改进码本

Yunqi Xue, Zhijiang Li, Philip Torr, Jindong Gu

机构 * School of Information Management, Wuhan University(武汉大学信息管理学院) Torr Vision Group, University of Oxford(牛津大学Torr视觉组)

AI总结 针对自回归统一多模态模型生成图像的安全性问题,提出迭代自改进码本方法,利用模型自身判断不安全图像并修正码本映射,消除有害输出并保证生成质量。

Comments 10 pages including references, 8 figures, accepted for publication at the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27095 2026-06-26 cs.LG cs.AI 新提交

Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning

无数据储备特征用于高效长时域冷启动持续学习

Augustinas Jučas, Yangchen Pan

机构 * University of Oxford(牛津大学) Department of Computer Science(计算机科学系) Department of Engineering Sciences(工程科学系)

AI总结 提出CIRCLE方法,使用固定的双向二维储备特征和流式线性判别分析头,无需重放、预训练或骨干网络反向传播,在长任务序列中显著优于现有冷启动无样本类增量学习方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26836 2026-06-26 cs.AI 新提交

The Capability Frontier: Benchmarks Miss 82% of Model Performance

能力前沿:基准测试遗漏了82%的模型性能

Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Antía García, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay

机构 * Martian University of Oxford(牛津大学) ThoughtWorks

AI总结 提出能力前沿概念,通过帕累托前沿量化多模型多生成下的最佳性能,纠正单模型单次评估偏差,在16个基准上实现82%性能提升和85%成本降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26463 2026-06-26 cs.LG 新提交

Finding the Time to Think: Learning Planning Budgets in Real-Time RL

寻找思考时间:在实时强化学习中学习规划预算

Aneesh Muppidi, Firas Darwish, Dylan Cope, João F. Henriques, Jakob Nicolaus Foerster

机构 * British Open-ended Learning and Discovery Lab (BOLD)(英国开放性学习与探索实验室(BOLD)) University of Oxford(牛津大学) VGG(视觉信息小组)

AI总结 针对环境在决策期间持续运行的实时强化学习,提出一种轻量级门控策略,根据状态选择规划预算,在多个实时游戏中优于固定预算和启发式基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26350 2026-06-26 cs.AI cs.LG 新提交

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

OpenFinGym: 一个可验证的多任务Gym环境,用于评估量化智能体

Kaicheng Zhang, Wen Ge, Lei Jiang, Weixin Yang, Jordan Langham-Lopez, Jialin Yu, Lukasz Szpruch, Hao Ni

机构 * University of Edinburgh(爱丁堡大学) University College London(伦敦大学学院) Alan Turing Institute(艾伦·图灵研究所) University of Oxford(牛津大学)

AI总结 提出OpenFinGym统一环境,覆盖预测、市场生成、实时交易和欺诈检测等量化金融任务,支持自动化任务构建、容器化验证和延迟解析,以全面评估LLM智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25450 2026-06-26 cs.LG cs.CL 新提交

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms

泛化谱:一种评估学习算法的色谱方法

Jinghan Zhang, Zerui Cheng, Shiqi Chen, Ge Zhang, Wenhao Huang, Jiashuo Liu, Junxian He, Tianle Cai

机构 * ByteDance Seed(字节跳动Seed) Hong Kong University of Science and Technology(香港科技大学) Princeton University(普林斯顿大学) University of Oxford(牛津大学)

AI总结 提出泛化谱框架,通过构建从精确回忆到跨语言实现转移等不同转移距离的测试变体,揭示算法从单个样本泛化到其他样本的程度,并应用于竞争编程任务评估不同学习范式。

Comments Accepted at ICML 2026. 30 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14397 2026-06-26 cs.LG 新提交

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Running the Gauntlet: 重新评估智能体在陌生环境中的能力

Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna, Damian Rynczak, Shreyansh Padarha, Kumail Alhamoud, Zihao Fu, William Lugoloobi, Kai Rawal, Hanna Yershova, Xander Davies, Taras Rumezhak, Guohao Li, Fazl Barez, Baoyuan Wu, Arkadiusz Drohomirecki, Yarin Gal, Chris Russell, Christopher Summerfield, Adam Mahdi, Volodymyr Karpiv, Philip Torr, Adel Bibi

机构 * University of Oxford(牛津大学) SoftServe Massachusetts Institute of Technology(麻省理工学院) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) UK AI Security Institute(英国人工智能安全研究所) Ukrainian Catholic University(乌克兰天主教大学)

AI总结 提出GauntletBench基准,通过20个视觉密集型任务评估智能体在时间感知、图形理解和3D推理等未被充分探索的能力,发现最先进智能体成功率仅19.1%,远低于人类80%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25829 2026-06-25 cs.RO 新提交

Beyond a Shadow of a Doubt: Close Proximity Geometry Reconstruction Using FMCW Radar Shadow Effects

不容置疑:利用FMCW雷达阴影效应进行近距离几何重建

Felix de Trogoff du Boisguezennec, Benjamin Ramtoula, Daniele De Martini

机构 * Oxford Robotics Institute, University of Oxford, UK(牛津大学机器人研究所,英国) ETH Zurich, Switzerland(瑞士苏黎世联邦理工学院)

AI总结 针对恶劣环境下感知退化问题,提出利用车辆底盘遮挡形成的雷达阴影恢复近处细长垂直物体的三维倾斜角度,通过解析闭式映射实现,仿真和实验验证了可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25527 2026-06-25 cs.LG 新提交

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

超越一刀切:基于诊断的在线强化学习与离线先验知识

Guozheng Ma, Lu Li, Zilin Wang, Pierre-Luc Bacon, Dacheng Tao

机构 * Nanyang Technical University(南洋理工大学) University of Oxford(牛津大学)

AI总结 本文提出诊断驱动的张力管理框架,通过三种功能角色表征离线先验知识如何重塑在线优化,并展示跨领域证据,解决先验知识有效性随部署变化的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25197 2026-06-25 cs.LG stat.ML 新提交

Efficient Adaptive Data Acquisition via Pretrained Belief Representations

通过预训练信念表示的高效自适应数据获取

Daolang Huang, Zhuoyue Huang, Conor Hassan, Luigi Acerbi, Samuel Kaski, Tom Rainforth

机构 * ELLIS Institute Finland(芬兰ELLIS研究所) Department of Computer Science, Aalto University, Finland(芬兰阿尔托大学计算机科学系) Department of Computer Science, University of Helsinki, Finland(芬兰赫尔辛基大学计算机科学系) Department of Computer Science, University of Manchester, UK(英国曼彻斯特大学计算机科学系) Department of Statistics, University of Oxford, UK(英国牛津大学统计系)

AI总结 提出POLAR框架,利用预训练预测模型作为信念状态编码器,解耦表示学习与策略学习,在贝叶斯实验设计、贝叶斯优化和主动学习中实现高效自适应数据获取。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25179 2026-06-25 cs.RO 新提交

Learning Perceptive Platform Adaptive Locomotion Controllers for Quadrupedal Robots

学习四足机器人的感知平台自适应运动控制器

David Rytz, Kim Tien Ly, Ioannis Havoutis

机构 * Dynamic Robot Systems, Oxford Robotics Institute, University of Oxford(牛津大学牛津机器人研究所动态机器人系统)

AI总结 研究如何将感知融入形态感知强化学习架构,通过自适应地形课程训练跨形态通用控制器,发现仅评论家感知在鲁棒性和稳定性上优于完全感知策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23838 2026-06-25 cs.LG astro-ph.IM physics.comp-ph physics.data-an stat.ML 新提交

The Degeneracy Distillery

退化蒸馏器

T. Lucas Makinen, Deaglan J. Bartlett, Niall Jeffrey, Benjamin D. Wandelt

机构 * Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学应用数学与理论物理系) Imperial Centre for Inference and Cosmology (ICIC), Imperial College London(伦敦帝国理工学院帝国推理与宇宙学中心) Astrophysics, University of Oxford(牛津大学天体物理学系) CNRS & Sorbonne Université, Institut d’Astrophysique de Paris (IAP)(法国国家科学研究中心与索邦大学巴黎天体物理研究所) Department of Physics and Astronomy, University College London(伦敦大学学院物理与天文学系) Department of Physics & King’s Institute for Artificial Intelligence, King’s College London(伦敦国王学院物理系与国王人工智能研究所) Department of Physics and Astronomy, Johns Hopkins University(约翰霍普金斯大学物理与天文学系) Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系)

AI总结 提出退化蒸馏器方法,通过估计和展平Fisher信息矩阵,自动符号化检测并解决物理模型中的退化参数组合,降低神经后验估计所需的模拟预算。

Comments 30 pages, 10 figures. Supporting code found at https://github.com/tlmakinen/degeneracy_distillery

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24876 2026-06-24 cs.CV 新提交

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

FLAT: 前馈潜在三角形泼溅用于几何精确的场景生成

Orest Kupyn, Goutam Bhat, Philipp Henzler, Fabian Manhardt, Christian Rupprecht, Federico Tombari

机构 * Google Research(谷歌研究院) University of Oxford, Visual Geometry Group(牛津大学视觉几何组) TU Munich(慕尼黑工业大学)

AI总结 提出FLAT方法,首次从视频扩散潜在表示中直接解码三角形泼溅,通过射线中心旋转参数化和乘积窗口函数解决梯度流问题,在保持视觉质量的同时显著提升几何精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24759 2026-06-24 cs.CV cs.AI 新提交

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架

Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye

机构 * organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom organization= SpaceTimeLab, Department of Civil, Environmental Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom

AI总结 提出UniDrive框架,通过融合时序推理与高分辨率感知分支,联合生成风险描述和边界框定位,在DRAMA-Reasoning基准上超越现有方法,提升小目标定位和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24251 2026-06-24 cs.AI 新提交

Probing the Misaligned Thinking Process of Language Models

探究语言模型的错误对齐思维过程

Kaiwen Zhou, Constantin Venhoff, Jonathan Michala, Xin Eric Wang, William Saunders

机构 * University of Oxford(牛津大学)

AI总结 提出通过线性探针检测模型内部激活中的18种错误对齐指标,以可靠识别策略欺骗等行为,在分布外基准上达到0.935 AUROC。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24660 2026-06-24 q-bio.QM cs.LG cs.NA math.NA physics.bio-ph 新提交

Extended pseudo-spectral physics-informed neural networks for phase-field models

用于相场模型的扩展伪谱物理信息神经网络

Callum Marsh, Radek Erban, Andreas Munch

机构 * Mathematical Institute, University of Oxford(牛津大学数学研究所)

AI总结 提出扩展伪谱物理信息神经网络(ESPINN),从瞬态快照数据中同时恢复相场模型的体化学势和梯度系数,实现数据高效且物理一致的逆辨识。

Comments 20 pages, 10 figures, Data available: https://doi.org/10.5281/zenodo.20797058

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23662 2026-06-23 stat.ML cs.LG 新提交

Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives

Action-BED:具有单重难解目标的任务驱动贝叶斯实验设计

Tom Rossa, Angus Phillips, Tom Rainforth

机构 * Department of Statistics(统计系) University of Oxford(牛津大学)

AI总结 提出基于期望未来损失的任务驱动贝叶斯实验设计框架,将设计策略与下游行动策略联合优化,避免双重难解目标,仅需采样和评估损失函数。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23515 2026-06-23 stat.ML cs.LG 新提交

FairBED: A Bayesian Experimental Design Approach to Gathering Fairer Data

FairBED:一种用于收集更公平数据的贝叶斯实验设计方法

Marcel Hedman, Emily Alger, Brieuc Lehmann, Chris Holmes, Tom Rainforth

机构 * Department of Statistics(统计学系) CeBAM, Nuffield Department of Medicine(CeBAM医学系) University of Oxford(牛津大学) Department of Statistical Science(统计科学系) University College London(伦敦大学学院)

AI总结 提出FairBED方法,通过贝叶斯实验设计最大化目标信息增益并最小化敏感属性信息增益,以收集更公平的数据,改善公平性与准确性的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22239 2026-06-23 stat.ML cs.LG 新提交

Variance-Tilted Diffusion Models for Diverse Sampling

方差倾斜扩散模型用于多样化采样

Iskander Azangulov, Leo Zhang, Kianoosh Ashouritaklimi

机构 * Department of Statistics, University of Oxford, Oxford, UK(牛津大学统计学系)

AI总结 提出方差加权批分布,通过Doob h-变换导出交互粒子采样器,在扩散模型中实现多样化采样,具有透明概率目标。

Comments Accepted at SPIGM @ ICML workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23217 2026-06-23 cs.CL cs.AI 新提交

MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations

MuPPET:多轮对话中LLM助手上下文隐私的基准测试

Elena Sofia Ruzzetti, Cornelius Emde, Sangdoo Yun, Seong Joon Oh, Martin Gubri

机构 * Parameter Lab(参数实验室) University of Rome Tor Vergata(罗马第二大学) University of Oxford(牛津大学) NAVER AI Lab(NAVER AI实验室) KAIST AI(韩国科学技术院人工智能研究所)

AI总结 提出MuPPET基准,评估多用户对话中LLM代理的隐私泄露风险,发现模型在多用户场景下泄露更严重,现有防御措施效果有限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23213 2026-06-23 cs.LG 新提交

Deep learning-based detection of cessation of breathing in pre-term infants

基于深度学习的早产儿呼吸停止检测

Dineo Serame, Lionel Tarassenko, Mauricio Villarroel

机构 * Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford(牛津大学工程科学系生物医学工程研究所)

AI总结 本研究利用深度学习模型,基于阻抗呼吸描记、心电图和光电容积描记信号,检测早产儿呼吸暂停相关呼吸停止事件,发现信号模态比架构复杂度更重要,单模态IP模型性能最优。

Comments 14 pages main text, 8 figures. Submitted to IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23177 2026-06-23 cs.CV cs.AI 新提交

Interpretable Probabilistic Medical Image Segmentation via Gaussian Process with Explicit Modelling of Annotation Bias and Variability

通过显式建模标注偏差和变异性的高斯过程可解释概率医学图像分割

Qi Li, Yuliang Huang, Shaheer U. Saeed, Qianye Yang, Vasilis Stavrinides, Zachary M. C. Baum, Dean C. Barratt, J. Alison Noble, Tom Vercauteren, Yipeng Hu

机构 * University College London(伦敦大学学院) Queen Mary University of London(伦敦玛丽女王大学) University of Oxford(牛津大学) University College London Hospitals NHS Trust(伦敦大学学院医院NHS信托基金会) Imperial College Healthcare(帝国理工学院医疗保健)

AI总结 提出基于随机变分高斯过程的logit空间概率分割框架,显式分解为图像相关参考分布和标注者特定扰动,提高不确定性校准并保持分割精度。

Comments Accepted at MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21947 2026-06-23 cs.CV cs.AI 新提交

ScalePredictor: Instance-aware Scale Learning for Accurate Quantization of Vision Transformers

ScalePredictor:面向视觉Transformer精确量化的实例感知尺度学习

Changjun Li, Runqing Jiang, Lian Xu, Ye Zhang, Qingyong Hu, Yulan Guo

机构 * School of Electronics and Communication Engineering, Shenzhen Campus, Sun Yat-sen University(中山大学深圳校区电子与通信工程学院) Computer Science and Software Engineering, The University of Western Australia(西澳大学计算机科学与软件工程学院) University of Oxford(牛津大学)

AI总结 提出动态量化框架ScalePredictor,通过浅层激活分布范围与深层最优尺度的隐藏关联,利用多项式尺度投影模块高效生成所有量化尺度,在ImageNet上取得更优精度-效率权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21672 2026-06-23 cs.RO cs.AI cs.LG 新提交

Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models

使用基于接地潜在动作世界模型的异构演示模仿学习

Tianyou Wang, Anson Lei, Joe Watson, Ingmar Posner

机构 * University of Oxford(牛津大学)

AI总结 提出GLAM方法,通过共享潜在动作空间对齐异构数据源,学习接地潜在动作世界模型,在数据稀缺时提升模仿学习性能,平均任务成功率提高48%。

Comments 17 pages, 8 figures. Project page: https://viccccciv.github.io/glam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21173 2026-06-23 cs.LG cs.AI 新提交

Inverting the Bellman Equation: From $Q$-Values to World Models

逆推贝尔曼方程:从 $Q$ 值到世界模型

Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens, Jakob N. Foerster, Oliver Richardson

机构 * FLAIR, University of Oxford(FLAIR,牛津大学) Google DeepMind(谷歌DeepMind) Mila, University of Montreal(Mila,蒙特利尔大学)

AI总结 本文证明基于值的智能体在丰富奖励函数上训练时隐式编码世界模型,提出 $P$-learning 从 $Q$ 值提取模型,并给出编码真实转移核的充分条件,实验验证了隐式模型的准确性和泛化能力。

Comments 48 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏