arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Oxford(牛津大学)

2026-08-04 至 2026-08-04 共收录 14
2608.02171 2026-08-04 cs.AI 新提交

From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents

从分析到综合:个性化大语言模型智能体中的隐式行为对齐基准测试

Jiajia Song, Bobo Li, Haiwen Yi, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) University of Toronto(多伦多大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) University of Oxford(牛津大学)

AI总结 该研究针对大语言模型智能体的个性化问题,构建了IBA-Bench基准,提出IBA-Agent框架,实验显示其可在九类场景中提升隐式行为对齐效果,但有效个性化仍是重大挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01575 2026-08-04 cs.LG cs.AI 新提交

Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard

基于精确贝叶斯最优标准测量语言模型的上下文内算法推理能力

Hector Zenil, Luan Ozelim

机构 * Oxford Immune Algorithmics(牛津免疫算法学公司) Oxford University Innovation(牛津大学创新公司) London Institute for Healthcare Engineering(伦敦医疗工程研究所) King’s College London(伦敦国王学院)

AI总结 本研究提出F-ICL基准,以精确贝叶斯最优标准测量语言模型的上下文内算法推理能力,发现多数模型分布与最优标准存在差距,该差距与准确率无关且不受模型规模缩小,基准已开放。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01460 2026-08-04 cs.LG 新提交

Conformalized Large Language Models under Configuration Shift

配置偏移下的共形大语言模型

Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov

机构 * University of Stuttgart(斯图加特大学) Robert Bosch GmbH(罗伯特·博世有限公司) University of Oxford(牛津大学) LMU Munich(慕尼黑大学) Tsinghua University(清华大学) University of Southampton(南安普顿大学) University of Oslo(奥斯陆大学)

AI总结 该研究针对配置偏移对共形大语言模型有效性的影响展开系统分析,通过多维度实证研究揭示其削弱覆盖率的问题,提出两种实用缓解措施以恢复覆盖率并保留效率。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01370 2026-08-04 cs.CV 新提交

Understanding Synergistic Interactions among Pathology Foundation Models via Adaptive Fusion

通过自适应融合理解病理基础模型间的协同交互作用

Yuxiang Xiao, Yang Hu, Bin Li, Tianyang Zhang, Zexi Li, Huazhu Fu, Jens Rittscher, Kaixiang Yang

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Computing and Mathematical Sciences, University of Leicester(莱斯特大学计算与数学科学学院) Leicester Cancer Research Centre, University of Leicester(莱斯特大学莱斯特癌症研究中心) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Nuffield Department of Medicine, University of Oxford(牛津大学纳菲尔德医学院) A*STAR, Singapore(新加坡科技研究局)

AI总结 本研究针对病理基础模型存在的表示偏差问题,提出AdaFusion自适应融合框架,在三个公共基准上验证其性能优于单个模型及其他融合方法,还可提供可解释的组织可视化。

Comments 11 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01356 2026-08-04 cs.CV 新提交

Harnessing Adversarial Distillation to Customise Debiased, Disease-Specific Pathology Foundation Models for Breast Cancer

利用对抗蒸馏定制去偏差的、针对乳腺癌的疾病专用病理学基础模型

Zhiwei Chen, Yang Hu, Yuxiang Xiao, Yakun Ju, Tianyang Zhang, Yingxue Xu, Wei Li, Hao Chen, Jens Rittscher, Kaixiang Yang

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Computing and Mathematical Sciences, University of Leicester(莱斯特大学计算与数学科学学院) Leicester Cancer Research Centre, University of Leicester(莱斯特大学莱斯特癌症研究中心) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Nuffield Department of Medicine, University of Oxford(牛津大学纳菲尔德医学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程学系) ZoyMed(佐医医疗(ZoyMed))

AI总结 本研究提出SmartStu框架,通过多教师集成蒸馏与对抗蒸馏等技术,定制出比通用PFMs小30倍以上且性能相当的乳腺癌专用病理学基础模型。

Comments 11 pages, 2 figures, 2 tables. Accepted to MICCAI 2026 (early accept)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00490 2026-08-04 cs.CV 新提交

Image-Space Rule Discovery

图像空间中的规则发现

Misora Sugiyama, Toya Oyama, Hirokatsu Kataoka

机构 * The University of Tokyo(东京大学) National Institute of Advanced Industrial Science and Technology (AIST)(独立行政法人产业技术综合研究所) University of Oxford(牛津大学)

AI总结 本研究提出WISRD基准测试图像编辑模型的图像空间规则发现能力,发现Nano Banana Pro表现最优,当前模型可部分依赖图像内指令,且该模型在4×4数独等任务上有一定性能。

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00355 2026-08-04 cs.CL cs.LG 新提交

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

CurveShift:智能体的进展是标量吗?区分水平与形态

Hanwen Xing, Pengyun Wang, BingXu Meng, Kumail Alhamoud, Xiang Li, Jicheng Wang, Xin Yu, Xinyang Han, Xiaomin Li, Philip Torr, Yuexing Hao

机构 * University of Southern California(南加利福尼亚大学) University of Chicago(芝加哥大学) University of California, Berkeley(加利福尼亚大学伯克利分校) Massachusetts Institute of Technology(麻省理工学院) Stanford University(斯坦福大学) University of California, Davis(加利福尼亚大学戴维斯分校) Pennsylvania State University(宾夕法尼亚州立大学) Harvard University(哈佛大学) University of Oxford(牛津大学)

AI总结 该研究针对大型语言模型进展的标量总结问题,通过LiveCodeBench基准分离混淆,发现2024年9月后发布的模型在竞赛编程难任务上有超出预期的增益,发布了相关数据集与代码。

Comments 25 pages, 4 figures, 7 tables. Data and code: https://github.com/harvenstar/CurveShift

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00198 2026-08-04 cs.LG 新提交

AutoCause: A Python framework that automates expert decisions in environmental time-series causal discovery

AutoCause:一个自动化环境时间序列因果发现领域专家决策的Python框架

Marco Ruiz, Miguel Arana-Catania, David R. Ardila, Rodrigo Ventura

机构 * ISR-Lisbon, Instituto Superior Técnico(里斯本信号与系统研究所,里斯本高等理工学院) Digital Scholarship at Oxford, University of Oxford(牛津大学牛津数字学术中心) Jet Propulsion Lab., Caltech(加州理工学院喷气推进实验室)

AI总结 AutoCause是一款开源Python框架,通过封装四种因果发现方法、添加参考模型并分级链接,实现环境时间序列因果发现中专家决策的自动化,提升分析的可审计性与可重复性。

Comments 33 pages, 10 figures. Submitted to Environmental Modelling & Software

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07078 2026-08-04 q-fin.TR cs.AI cs.CE

Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?

基于LLM的金融投资策略能否长期跑赢市场?

Weixian Waylon Li, Hyeonjun Kim, Mihai Cucuringu, Tiejun Ma

机构 * AIAI, School of Informatics The University of Edinburgh Edinburgh United Kingdom Global Finance Research Center Sungkyunkwan University Seoul Republic of Korea Dept. of Statistics \& OMI University of California, Los Angeles University of Oxford United States The University of Edinburgh Sungkyunkwan University University of California, Los Angeles University of Oxford

AI总结 提出FINSABER回测框架,在更长时间和更大股票池上评估基于LLM的择时策略,发现其优势在长期和广泛截面下显著下降,且在牛熊市中表现不佳。

Comments KDD 2026, Datasets & Benchmarks Track (Oral) Corrected the FinAgent results and added FinAgent (GPT-4o-mini) in Table 2; conclusions unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08876 2026-08-04 cs.LG 版本更新

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

OTora:一种用于LLM代理推理层面拒绝服务攻击的统一红队框架

Xinyu Li, Ronghui Mu, Lin Li, Tianjin Huang, Gaojie Jin

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Department of Computer Science, University of Oxford(牛津大学计算机科学系) Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)

AI总结 OTora是首个统一的两阶段红队框架,用于实现推理层面拒绝服务攻击,通过优化对抗触发器和生成代理感知的推理负载,提升推理token数量和延迟,同时保持任务准确性。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27269 2026-08-04 cs.AI 版本更新

Unifying biomedical knowledge in a modern multimodal graph

OptimusKG:统一生物医学知识的现代多模态图

Lucas Vittor, Ayush Noori, Iñaki Arango, Joaquín Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Edison Scientific Inc.(Edison科学公司) Oxford Suzhou Centre for Advanced Research, University of Oxford(牛津大学苏格兰研究中心) Broad Institute of MIT and Harvard(MIT与哈佛大学Broad研究所) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究所) Harvard Data Science Initiative(哈佛大学数据科学计划)

AI总结 OptimusKG通过整合结构化和半结构化资源,构建了多模态生物医学标记属性图,统一了分子、解剖、临床和环境领域的知识,验证了其在生物医学领域的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02277 2026-08-04 cs.CR cs.AI 版本更新

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

量化前沿大语言模型突破容器沙盒的能力

Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis, Stuart Jennings, Freddy Tuxworth, Tolga H. Dur, Sam Deverett, John Wilkinson, Jason Gwartz, Harry Coppock

机构 * University of Oxford(牛津大学)

AI总结 研究大语言模型突破容器沙盒的能力,通过引入SANDBOXESCAPEBENCH基准,利用嵌套沙盒架构和CTF评估,涵盖多种逃逸机制,发现添加漏洞时大语言模型能识别利用,证明此类评估对保障沙盒封装高性能模型的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01210 2026-08-04 cs.AI 版本更新

Knowledge Graph Augmented Large Language Models for Disease Prediction

基于知识图谱的大型语言模型用于疾病预测

Ruiyu Wang, Tuan Vinh, Ran Xu, Yuyin Zhou, Jiaying Lu, Francisco Pasquel, Mohammed K Ali, Carl Yang

机构 * Department of Computer Science, Emory University(埃默里大学计算机科学系) Division of Medical Sciences, Oxford University(牛津大学医学系) Department of Computer Science and Engineering, UCSC(UCSC计算机科学与工程系) Nell Hodgson Woodruff School of Nursing, Emory University(埃默里大学内尔·霍根·伍德鲁夫护理学院) School of Medicine, Emory University(埃默里大学医学院) Rollins School of Public Health, Emory University(埃默里大学罗林斯公共卫生学院)

AI总结 本文提出基于知识图谱的大型语言模型框架,用于提升疾病预测的准确性与解释性,通过构建时间一致的推理依据,在多个数据集上取得优于经典基线的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01925 2026-08-04 cs.CL 版本更新

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

用奖励模型增强大语言模型推理能力:一项分析性综述

Qiyuan Liu, Hao Xu, Xuhong Chen, Wei Chen, Yee Whye Teh, Ning Miao

机构 * Department of Data Science and Hong Kong Institute of AI for Science, City University of Hong Kong(数据科学系和香港人工智能科学研究所,香港城市大学) Li Auto Inc., China(中国利汽车公司) Department of Statistics, University of Oxford(统计系,牛津大学)

AI总结 该综述系统介绍奖励模型(RMs)的概念、训练与评估方法,梳理其在大语言模型(LLMs)推理中的三类核心应用,并探讨RMs相关开放问题,为其有效部署提供见解。

Comments Accepted for publication in Artificial Intelligence Review

详情

展开后加载摘要…

URL PDF HTML 收藏