arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Oxford(牛津大学)

2026-05-14 至 2026-05-14 共收录 18
2605.13846 2026-05-14 cs.CL cs.AI

WARDEN: Endangered Indigenous Language Transcription and Translation with 6 Hours of Training Data

WARDEN:利用6小时训练数据进行濒危原住民语言的转录与翻译

Ziheng Zhang, Yunzhong Hou, Naijing Liu, Liang Zheng

机构 * Australian National University(澳大利亚国立大学) University of Oxford(牛津大学)

AI总结 WARDEN通过分阶段模型处理濒危语言转录与翻译,利用印尼语初始化加速训练,结合专家词典提升翻译性能,以6小时数据优于统一模型。

Comments https://github.com/Ziheng-Zhang-AUS/WARDEN

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13829 2026-05-14 cs.CL cs.AI cs.LG

Negation Neglect: When models fail to learn negations in training

否定忽视:当模型在训练中无法学习否定时的失败

Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen, James Chua, Owain Evans

机构 * University of Oxford(牛津大学) University of Toronto(多伦多大学) Warsaw University of Technology(华沙技术大学) NASK National Research Institute(国家研究 institute NASK) Truthful AI Anthropic UC Berkeley(伯克利大学)

AI总结 研究发现,当模型在训练中接触到标记为假的声明时,会错误地认为这些声明为真,且这种现象不仅发生在否定语句中,还扩展到其他认知限定词,影响模型的行为和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13806 2026-05-14 cs.DS cs.CC cs.GT cs.LG math.OC

Min-Max Optimization Requires Exponentially Many Queries

极小-极大优化需要指数级的查询次数

Martino Bernasconi, Matteo Castiglioni, Andrea Celli, Alexandros Hollender

机构 * Bocconi University(博科尼大学) Politecnico di Milano(米兰理工学院) University of Oxford(牛津大学)

AI总结 研究极小-极大优化的查询复杂性,证明任何算法找到近似 stationary 点所需查询次数随 1/ε 和 d 指数增长。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13740 2026-05-14 cs.LG

Learning POMDP World Models from Observations with Language-Model Priors

从观测中学习POMDP世界模型:利用语言模型先验

Valentin Six, Frederik Panse, Mathis Fajeau, Lancelot Da Costa, Mridul Sharma, Alfonso Amayuelas, Tim Z. Xiao, David Hyland, Philipp Hennig, Bernhard Schölkopf

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) IRIIS University of California, Santa Barbara(加州大学圣芭芭拉分校) University of Tübingen(图宾根大学) University of Oxford(牛津大学) ELLIS Institute Tübingen(图宾根ELLIS研究所)

AI总结 本文提出Pinductor,利用语言模型先验从少量观测-动作轨迹学习POMDP模型,实现高效世界模型学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13681 2026-05-14 cs.LG stat.ML

Sampling from Flow Language Models via Marginal-Conditioned Bridges

通过边缘-条件化桥梁采样流语言模型

Iskander Azangulov, Leo Zhang

机构 * Department of Statistics, University of Oxford(牛津大学统计系)

AI总结 本文提出通过边缘-条件化桥梁采样流语言模型,该方法在每一步从因子化后验中采样干净的一热序列,并利用奥本-乌伦贝克桥进行下一步状态采样,从而在不训练的情况下提升质量-多样性平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13542 2026-05-14 cs.AI cs.CL cs.LG cs.MA

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

RealICU: 大语言模型能否理解长期上下文ICU数据?一个超越行为模仿的基准测试

Chengzhi Shen, Weixiang Shen, Tobias Susetzky, Chen, Chen, Jun Li, Yuyuan Liu, Xuepeng Zhang, Zhenyu Gong, Daniel Rueckert, Jiazhen Pan

机构 * Technical University of Munich(慕尼黑技术大学) TUM University Hospital(TUM大学医院) LMU Munich(慕尼黑大学) University of Sheffield(谢菲尔德大学) University of Oxford(牛津大学) Zhongshan Hospital Fudan University(复旦大学中山医院) Sun Yat-sen University Cancer Center(中山大学肿瘤中心) Imperial College London(伦敦帝国学院) Munich Center for Machine Learning(慕尼黑机器学习中心) relAI – Konrad Zuse School of Excellence in Reliable AI(relAI – 卡诺夫茨卓越可靠人工智能学校)

AI总结 RealICU通过真实ICU场景评估大语言模型的长期上下文理解能力,提出四个医生驱动任务,揭示现有模型在临床推荐中的召回与安全性的权衡问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24649 2026-05-14 cs.CV

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

MedOpenClaw 和 MedFlowBench:在完整研究流程中审计医疗代理

Weixiang Shen, Chengzhi Shen, Yanzhu Hu, Che Liu, Junde Wu, Jiayuan Zhu, Xiao Han, Zongyue Li, Jingpei Wu, Min Xu, Daguang Xu, Yueming Jin, Benedikt Wiestler, Daniel Rueckert, Jiazhen Pan

机构 * Technical University of Munich(慕尼黑技术大学) TUM University Hospital(TUM大学医院) LMU Munich(慕尼黑大学) Imperial College London(伦敦帝国理工学院) University of Oxford(牛津大学) Carnegie Mellon University(卡内基梅隆大学) NVIDIA(NVIDIA公司) National University of Singapore(新加坡国立大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

AI总结 本文提出 MedFlowBench 和 MedOpenClaw,用于评估医疗影像代理在完整研究流程中生成可审计证据的能力,发现仅依赖答案评分不够,需结合正确证据验证。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21975 2026-05-14 cs.AI cs.ET

Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models

注意差距: elicitation协议如何影响语言模型中的 stated-revealed偏好差距

Pranav Mahajan, Ihor Kendiukhov, Syed Hussain, Lydia Nottingham

机构 * University of Oxford(牛津大学) Max Planck Institute for Biological Cybernetics(生物信息学Max Planck研究所) University of Tuebingen(图宾根大学) Cardiff University(卡迪夫大学) Cambridge–Boston Alignment Initiative (CBAI)(剑桥-波士顿对齐倡议)

AI总结 研究探讨了elicitation协议对语言模型中stated-revealed偏好差距的影响,发现中性选项和 abstention 的引入能提高相关性,但进一步允许 abstention 又导致相关性下降。

Comments Accepted to ACL 2026 Eval Eval Workshop and 3rd Technical AI Safety Conference (TAIS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09320 2026-05-14 cs.LG cs.AI cs.CR

Exact Verification of Graph Neural Networks with Incremental Constraint Solving

图神经网络的精确验证与增量约束求解

Minghao Liu, Chia-Hsuan Lu, Marta Kwiatkowska

机构 * University of Oxford(牛津大学)

AI总结 本文提出一种精确验证方法,用于图神经网络在属性和结构扰动下的鲁棒性保证,支持三种聚合函数,实验表明其在节点分类和图分类任务中性能优越。

Comments Extended version of the paper accepted at FM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15616 2026-05-14 cs.CV

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

LENS:基于大语言模型的多级多模态推理评估

Ruilin Yao, Bo Zhang, Jirui Huang, Xinwei Long, Yifang Zhang, Tianyu Zou, Yufei Wu, Shichao Su, Yifan Xu, Wenxi Zeng, Zhaoyu Yang, Guoyou Li, Shilan Zhang, Zichan Li, Yaxiong Chen, Shengwu Xiong, Peng Xu, Jiajun Zhang, Bowen Zhou, David Clifton, Luc Van Gool

机构 * Wuhan University of Technology(武汉理工大学) Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室) University of Oxford(牛津大学) INSAIT, Sofia Un. St Kliment Ohridski(索菲亚大学克里门特·欧里迪斯基学院)

AI总结 LENS提出一个包含3.4K图像和60K+问题的多级评估基准,涵盖8个任务和12种日常场景,用于评估大语言模型在多模态推理中的表现。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09760 2026-05-14 cs.RO cs.NE

Neural Associative Skill Memories for safer robotics and modelling human sensorimotor repertoires

神经关联技能记忆用于更安全的机器人及建模人类传感器运动 repertoire

Pranav Mahajan, Mufeng Tang, T. Ed Li, Ioannis Havoutis, Ben Seymour

机构 * University of Oxford(牛津大学) Yale University(耶鲁大学)

AI总结 本文提出神经关联技能记忆框架,通过自监督预测编码统一技能学习与表达,实现故障检测与上下文感知执行,提升机器人安全性和生物传感器运动学习研究。

Journal ref Neural Computation (2026) 38 (1): 1-27

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13307 2026-05-14 cs.CL cs.HC

PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users

PRISM-X:基于人类和模拟用户的人个性化微调实验

Hannah Rose Kirk, Liu Leqi, Fanzhi Zeng, Henry Davidson, Bertie Vidgen, Christopher Summerfield, Scott A. Hale

机构 * University of Oxford(牛津大学) UK AI Security Institute(英国人工智能安全研究所) University of Texas at Austin(德克萨斯大学奥斯汀分校) Mercor Meedan

AI总结 研究通过大规模内侧实验评估个性化语言模型在多轮对话中的表现,发现基于偏好微调的P-DPO方法优于通用模型和个性化提示,但个性化数据训练收益有限,且微调会放大趋炎附势行为,可能带来长期负面影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12814 2026-05-14 cs.SI cs.CL

Linking Extreme Discourse to Structural Polarization in Signed Interaction Networks

将极端言论与结构极化联系起来:在带符号交互网络中

Zhijin Guo, Li Zhang, Tyler Bonnet, Janet B. Pierrehumbert, Xiaowen Dong

机构 * University of Oxford(牛津大学) University College London(伦敦大学学院) Imperial College London(伦敦帝国学院)

AI总结 本文提出一种基于语言的带符号网络框架,通过LLM立场评分生成连续符号边权重,并利用两种互补指标量化结构极化,分析Reddit Brexit讨论中话语信号与结构极化的时间变化关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15743 2026-05-14 cs.LG astro-ph.EP astro-ph.IM

Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models

连接点:为电离层预报模型准备的机器学习数据集

Linnea M. Wolniewicz, Halil S. Kelebek, Simone Mestici, Michael D. Vergalla, Giacomo Acciarini, Bala Poduval, Olga Verkhoglyadova, Madhulika Guhathakurta, Thomas E. Berger, Atılım Güneş Baydin, Frank Soboczenski

机构 * Department of Information and Computer Science(信息与计算机科学系) University of Hawai‘i at Mānoa(夏威夷大学毛纳罗亚分校) Department of Engineering Science(工程科学系) University of Oxford(牛津大学) Università degli Studi di Roma Sapienza(罗马大学) Free Flight Research Lab(自由飞行研究实验室) University of New Hampshire(新罕布什尔大学) European Space Agency (ESA)(欧洲航天局) NASA Jet Propulsion Laboratory(美国宇航局喷气推进实验室) NASA Headquarters(美国宇航局总部) Space Weather Technology, Research, and Education Center(空间天气技术、研究与教育中心) University of Colorado Boulder(科罗拉多大学博尔德分校) Department of Computer Science(计算机科学系) University of York & King’s College London(约克大学及伦敦国王学院)

AI总结 本文提出一个整合多种电离层和日球层数据的机器学习数据集,用于改进电离层预报模型,支持科学探索和实际应用。

Comments 8 pages, 2 figures, 2 tables. Accepted as a poster presentation in the Machine Learning for the Physical Sciences workshop at NeurIPS 2025. Dataset can be found on Zenodo (https://zenodo.org/records/18343833) or GitHub (https://github.com/FrontierDevelopmentLab/2025-HL-Ionosphere-dataset)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00626 2026-05-14 cs.CV cs.AI

Towards Methane Detection Onboard Satellites

向卫星 onboard 的甲烷检测

Maggie Chen, Hala Lamdouar, Luca Marini, Laura Martínez-Ferrer, Chris Bridges, Giacomo Acciarini

机构 * University of Oxford(牛津大学) Delft University of Technology(代尔夫特理工大学) Universitat de València(瓦伦西亚大学) University of Surrey(萨里大学) European Space Agency (ESA)(欧洲航天局)

AI总结 本文提出一种无需预处理的甲烷检测方法,利用未正射校正数据训练模型,与传统方法性能相当,并释放了相关数据集和代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03167 2026-05-14 cs.CL cs.AI cs.LG

Where Do Reasoning Models Refuse?

推理模型拒绝发生在何处?

Kureha Yamaguchi, Benjamin Etheridge, Andy Arditi

机构 * The Alan Turing Institute(艾伦·图灵研究所) University of Oxford(牛津大学) Northeastern University(东北大学)

AI总结 研究探讨推理模型在生成响应前拒绝决策的位置,发现推理链中的初始句子对拒绝决定有显著影响,并通过激活方向分析揭示了拒绝机制。

Comments v1 accepted to the ICML 2025 Workshop on Reliable and Responsible Foundation Models (R2FM). 20 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14375 2026-05-14 cs.LG cs.CL

Causal Fine-Tuning under Latent Confounded Shift

潜在混杂偏移下的因果微调

Jialin Yu, Yuxiang Zhou, Haoxuan Li, Junchi Yu, Mengyue Yang, Yulan He, Nevin L. Zhang, Philip Torr, Ricardo Silva

机构 * University of Oxford, United Kingdom(牛津大学) Queen Mary University of London, United Kingdom(伦敦玛丽女王大学) Peking University, China(北京大学) University of Bristol, United Kingdom(布里斯托大学) King's College London, United Kingdom(伦敦国王学院) Hong Kong University of Science(香港科学大学) University College London, United Kingdom(伦敦大学学院)

AI总结 本文提出Causal Fine-Tuning方法,通过结构因果模型分解表示为稳定高层和敏感低层成分,提升模型在存在潜在混杂偏移时的鲁棒性,实验显示优于其他基线方法。

Comments ICML 2026 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12536 2026-05-14 q-bio.NC cs.AI cs.IT math.IT

Information as Maximum-Caliber Deviation: A bridge between Integrated Information Theory and the Free Energy Principle

信息作为最大 caliber 偏离:集成信息理论与自由能原理之间的桥梁

Alexander Kearney

机构 * University of Oxford(牛津大学)

AI总结 本文提出信息可定义为动态偏离最大 caliber 路径集的量,通过最大 caliber 变分原理重新推导了IIT的现象学计算,为IIT扩展到新动态领域提供了理论桥梁。

Comments 84 pages, 10 figures, 2 tables Extended version of a Master's thesis, Mathematical Institute, University of Oxford

详情

展开后加载摘要…

URL PDF HTML 收藏