arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Oxford(牛津大学)

共收录 1445
2601.22184 2026-06-17 cs.GT cs.LG cs.MA 版本更新

Tacit Coordination of Large Language Models

大型语言模型的隐性协调

Ido Aharon, Emanuele La Malfa, Michael Wooldridge, Sarit Kraus

机构 * Department of Computer Science, Bar-Ilan University(巴伊兰大学计算机科学系) Department of Computer Science, University of Oxford(牛津大学计算机科学系) Institute for Decentralized AI (IDAI)(去中心化人工智能研究所)

AI总结 研究大型语言模型在多智能体无通信协调中的焦点涌现能力,通过博弈和搜救任务评估,发现模型在多数场景匹配或超越人类,但在数值常识和文化显著性任务中失败,并提出无学习策略改善协调。

Comments Code: https://github.com/EmanueleLM/focal-points

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17048 2026-06-16 cs.LG cs.CV stat.ML 新提交

Exact Posterior Score Estimation for Solving Linear Inverse Problems

精确后验分数估计用于求解线性逆问题

Abbas Mammadov, Ozgur Kara, Kaan Oktay, Iskander Azangulov, Adil Kaan Akan, Hyungjin Chung, James Matthew Rehg, Yee Whye Teh

机构 * University of Oxford(牛津大学) UIUC(伊利诺伊大学厄巴纳-香槟分校) EverEx

AI总结 提出精确后验分数(EPS)方法,通过闭式后验分数将线性逆问题转化为去噪问题,无需梯度或投影,在FFHQ和ImageNet上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17000 2026-06-16 cs.CC cs.GT cs.LG math.OC 新提交

The Complexity of Min-Max Optimization for Quadratic Polynomials

二次多项式极小极大优化的复杂性

Martino Bernasconi, Matteo Castiglioni, Andrea Celli, Alexandros Hollender

机构 * Bocconi University(博科尼大学) Politecnico di Milano(米兰理工学院) University of Oxford(牛津大学)

AI总结 证明超立方体上极小极大优化的近似稳定点计算对二次多项式是PPAD难的,即使多项式是多线性的且每个变量最多出现在三个单项式中。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16587 2026-06-16 physics.flu-dyn cs.AI cs.LG physics.comp-ph 新提交

Learning Interface Breakup: A Geometry-Conditioned Latent Surrogate for Spray Formation

学习界面破碎:一种用于喷雾形成的几何条件潜在代理模型

Julius H Ramlau, Friedrich Hastedt, Tolga Birdal, Ehecatl-Antonio del Río Chanona, Nausheen S Basha, Omar K Matar

机构 * University of California, Berkeley(加州大学伯克利分校) Technical University of Munich(慕尼黑技术大学) Istanbul Technology University(伊斯坦布尔技术大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Cambridge(剑桥大学) University of Oxford(牛津大学)

AI总结 提出一种几何条件潜在代理模型,通过编码自适应网格细化(AMR)的单元密度场,在797个两相喷嘴模拟上训练,实现瞬态破碎动力学的高效预测,推理速度比Basilisk CFD快6×10^4倍。

Comments 11 pages, 5 figures, accepted to ICML AI4Physics 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16569 2026-06-16 cs.CV cs.RO 新提交

PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models

PROSE: 基于视觉语言模型的无训练自我中心场景配准

Zhiang Chen, Nahyuk Lee, Boyang Sun, Taein Kwon, Marc Pollefeys, Zuria Bauer, Sunghwan Hong

机构 * ETH Zurich(苏黎世联邦理工学院) VGG, University of Oxford(牛津大学VGG实验室) ETH AI Center(苏黎世联邦理工学院人工智能中心)

AI总结 提出PROSE方法,利用预训练视觉语言模型将RGB序列提升为对象级3D场景图,通过对象高度先验和相同/不同查询匹配实例,无需训练或深度传感器即可实现自我中心场景配准,在Aria基准上超越几何和场景图基线。

Comments Project page: https://rckola.github.io/prose/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16564 2026-06-16 cs.RO cs.LG 新提交

Elastic ODYN: Differentiable Optimization for Infeasible Control and Learning in Robotics

Elastic ODYN:面向机器人中不可行控制与学习的可微优化

Aristotelis Papatheodorou, Jose Rojas, Ioannis Havoutis, Carlos Mastalli

机构 * University of Oxford(牛津大学) Heriot-Watt University(赫瑞瓦特大学)

AI总结 提出Elastic ODYN,一种通过平滑平方ℓ2弹性松弛处理不可行二次规划(QP)的原始-对偶非内点求解器,支持热启动,在无可行点时收敛到最接近可行解,并基于此开发可微QP层和不可行感知SQP方法,在基准QP、奇异接触力学、可微参数辨识及四足/人形机器人轨迹优化中优于现有方法。

Comments 8 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16475 2026-06-16 cs.CY cs.AI 新提交

AI systems out-persuade expert humans

AI系统在说服力上超越人类专家

Kobi Hackenburg, Caroline Wagner, Luke Hewitt, Ben M. Tappin, Ed Saunders, Hannah Rose Kirk, Helen Margetts, Christopher Summerfield

机构 * University of Oxford(牛津大学) UK AI Security Institute(英国人工智能安全研究所) Stanford University(斯坦福大学) London School of Economics and Political Science(伦敦政治经济学院)

AI总结 通过四项预注册实验(n=18,978次对话),发现AI系统在说服力上可靠地超越人类专家,包括专业拉票者和世界辩论冠军,其优势源于快速部署大量信息,并扩展到现实世界筹款行为。

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15989 2026-06-16 q-bio.NC cs.AI 新提交

Task-guided cross-subject latent alignment: a multi-encoder-decoder VAE

任务引导的跨被试潜在对齐:一种多编码器-解码器VAE

Angeliki Papathanasiou, Jascha Achterberg, Thomas E. Nichols, Rui Ponte Costa

机构 * Centre for Neural Circuits and Behaviour Department of Physiology Anatomy and Genetics University of Oxford(神经回路与行为中心 生理解剖与遗传学系 牛津大学) Big Data Institute Nuffield Department of Medicine University of Oxford(大数据研究所 纳菲尔德医学系 牛津大学)

AI总结 提出MED-VAE模型,通过预训练ANN锚定表征,实现无共享刺激的跨被试神经对齐,在自然场景数据集上优于传统方法,并支持跨被试图像解码。

Comments In Proceedings of the 9th Conference on Cognitive Computational Neuroscience, New York, NY, USA, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15983 2026-06-16 quant-ph cond-mat.mtrl-sci cs.LG 新提交

Learning ground state observables from quantum computing experiments

从量子计算实验中学习基态可观测量

Ben Jaderberg, Freya Shah, Minjun Jeon, M. Emre Sahin, Christa Zoufal, Kunal Sharma

机构 * IBM Quantum, IBM Research Europe, Hursley, Winchester, SO21 2JN, United Kingdom(IBM量子、IBM欧洲研究院,赫尔斯利,温切斯特,SO21 2JN,英国) Department of Engineering Science, University of Oxford, Parks Road, Oxford OX1 3PJ, United Kingdom(工程科学系,牛津大学,帕克斯路,牛津 OX1 3PJ,英国) IBM Quantum, T. J. Watson Research Center, Yorktown Heights, NY 10598, USA(IBM量子、T.J. Watson研究中心,扬斯敦高地,纽约 10598,美国) Department of Materials, University of Oxford, Parks Road, Oxford OX1 3PH, United Kingdom(材料系,牛津大学,帕克斯路,牛津 OX1 3PH,英国) The Hartree Centre, STFC, Sci-Tech Daresbury, Warrington WA4 4AD, UK(哈特里中心,STFC,科技达尔斯伯里,沃林顿 WA4 4AD,英国) IBM Quantum, IBM Research Europe — Zurich, Ruschlikon 8803, Switzerland(IBM量子、IBM欧洲研究院——苏黎世,卢斯利康 8803,瑞士) IBM Research, Chicago, IL 60606, USA(IBM研究院,芝加哥,伊利诺伊 60606,美国)

AI总结 本文在115量子比特的二维海森堡XXZ模型中,利用近似基态的实验数据训练神经网络,成功预测了未见哈密顿量参数下的空间分辨可观测量,展示了从量子数据学习的实际可行性。

Comments 20 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15949 2026-06-16 cs.CL 新提交

FinBalance: A Multi-Document Accounting Reconciliation Benchmark

FinBalance:多文档会计对账基准

Sasank Tumpati, Devansh Agarwal, Ayush Kedia, Arjun Neekhra, Murari Mandal, Krishna Garg, Yash Sinha, Suman Gupta, Dhruv Kumar

机构 * BITS Pilani(比拉理工学院皮拉尼校区) KIIT Bhubaneswar(KIIT布巴内斯瓦尔) University of Oxford(牛津大学)

AI总结 提出FinBalance基准,通过多行业源文档构建会计对账任务,评估LLM在生成资产负债表和检测不一致性上的表现,发现模型在文档绑定和一致性聚合上存在显著差距。

Comments 18 pages, 12 figures. Code and data: https://github.com/Devansh1105/finbalance

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15897 2026-06-16 cs.LG cs.AI stat.ML 新提交

Topological Flow Matching

拓扑流匹配

Kacper Wyrwal, İsmail İlkan Ceylan, Alexander Tong

机构 * University of Oxford(牛津大学) TU Wien(维也纳技术大学) AITHYRA

AI总结 提出拓扑流匹配,通过拉普拉斯漂移增强参考过程,在保留流匹配稳定性和无模拟目标的同时,捕捉底层域拓扑结构,适用于脑fMRI、洋流等结构化数据。

Comments Accepted at ICLR 2026. 26 pages, 24 figures. Code: https://github.com/KacperWyrwal/topological-flow-matching

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15890 2026-06-16 cs.AI 新提交

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

UrbanWell: 面向时空城市福祉分析的多模态大语言模型基准测试

Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui

机构 * University of Helsinki(赫尔辛基大学) Zhongguancun Academy(中关村学院) University of Oxford(牛津大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 提出UrbanWell基准,通过卫星和街景图像联合建模,系统评估多模态大语言模型在环境、空间可达性、城市形态、活力和主观感知等5类城市福祉指标上的时空推理能力,并定义时序预测和趋势分类任务。

Comments accepted by KDD Datasets and Benchmarks Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15872 2026-06-16 cs.CL 新提交

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

SciOrch: 学习编排专家大语言模型以解决前沿多模态科学推理任务

Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin

机构 * Imperial College London(伦敦帝国学院) The Chinese University of Hong Kong(香港中文大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Oxford(牛津大学) Shenzhen Loop Area Institute(深圳河套学院)

AI总结 提出SciOrch框架,训练轻量级8B模型编排多个前沿大语言模型,通过MCTS和GRPO优化,在科学推理任务上超越最强单模型和多智能体基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15497 2026-06-16 cs.AI 新提交

Towards End-to-End Automation of AI Research

迈向AI研究的端到端自动化

Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Chris Lu, Shengran Hu, Jakob Foerster, David Ha, Jeff Clune

机构 * Sakana AI FLAIR University of Oxford(牛津大学) University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所)

AI总结 提出AI Scientist系统,利用基础模型实现从构思到论文撰写的全自动研究,并通过机器学习会议研讨会的同行评审。

Comments Published in Nature 651, 914-919 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15327 2026-06-16 cs.LG 新提交

Semantic DLM+: Improving Diffusion Language Models through Bias-variance Trade-off in Transition Kernel Design

语义DLM+:通过转移核设计中的偏差-方差权衡改进扩散语言模型

Keyue Jiang, Yuxiang Wang, Yanan Zhao, Xiang Yu, Qifang Zhao, Bohan Tang, Baojian Zhou, Yanghua Xiao, Lin Qu, Xiaoxiao Xu

机构 * Alibaba Group(阿里巴巴集团) Fudan University(复旦大学) University College London(伦敦大学学院) Nanyang Technological University(南洋理工大学) University of Oxford(牛津大学)

AI总结 本文通过分析泛化误差的三个关键因素,提出SemDLM+模型,通过全局转移和语义频率惩罚解决语义盆地问题,在LM1B和OpenWebText上提升了训练动态和生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15315 2026-06-16 cs.AI 新提交

ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing

ChatPlanner: 一个用于个性化公共交通路线规划的大型语言模型框架

Tingting Yang, Chenhao Xue, Jun Chen

机构 * School of Engineering and Materials Science, Queen Mary University of London(伦敦玛丽女王大学工程与材料科学学院) Department of Engineering Science, University of Oxford(牛津大学工程科学系)

AI总结 提出ChatPlanner框架,利用大型语言模型和检索增强生成技术从自然语言查询中提取用户偏好并融入路线规划算法,实验证明其能生成更符合个性化需求的可行路线方案。

Comments Under Review at Transportation Research Part C

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15048 2026-06-16 cs.LG cs.CV 新提交

Temporal Difference Learning for Diffusion Models

扩散模型的时间差分学习

Qizhen Ying, Yangchen Pan, Victor Adrian Prisacariu, Junfeng Wen

机构 * Department of Engineering Science, University of Oxford, Oxford, United Kingdom(牛津大学工程科学系) School of Computer Science, Carleton University, Ottawa, Canada(卡尔顿大学计算机科学学院)

AI总结 提出时间差分(TD)目标函数,通过将扩散过程视为马尔可夫奖励过程并利用强化学习中的策略评估,强制去噪轨迹上的跨时间一致性,显著提升少步采样下的生成质量。

Comments 15 pages, 4 figures. Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14766 2026-06-16 cs.CV cs.AI cs.MA 新提交

XMedFusion: A Knowledge-Guided Multimodal Perception and Reasoning Framework for Autonomous Medical Systems

XMedFusion:面向自主医疗系统的知识引导多模态感知与推理框架

Hamza Riaz, Arham Haroon, Maha Baig, Muhammad Dawood Rizwan, Muhammad Naseer Bajwa, Muhammad Moazam Fraz

机构 * National University of Sciences and Technology (NUST)(巴基斯坦国立科技大学) University of Oxford(牛津大学)

AI总结 提出XMedFusion模块化AI框架,通过视觉感知、知识图谱构建和检索引导生成等智能体协同,增强放射学报告生成的视觉基础与临床发现捕捉能力,在公共数据集上显著优于基线模型。

Comments Accepted at the 2026 International Conference on Robotics and Automation in Industry (ICRAI)

Journal ref 2026 International Conference on Robotics and Automation in Industry (ICRAI), pp. 1-6, May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12291 2026-06-16 cs.CL 新提交

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

测量大语言模型在误导性医疗上下文下的认知韧性

Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro, Michael Petrus, Sheikh Momin, Alexandra M. Cardoso Pinto, Rachel Niesen, Laura Sophie Wegner, Dhruv Darji, Jung Moses Koo, Joshua Fieggen, Kapil Narain, Mingde Zeng, Lei Clifton, Linda Shapiro, Fenglin Liu, David A. Clifton

机构 * University of Oxford(牛津大学) University of Washington(华盛顿大学) University College London(伦敦大学学院) University of Waterloo(滑铁卢大学)

AI总结 本研究提出MedMisBench基准,通过注入误导性上下文测试大语言模型在医疗场景中的认知韧性,发现模型准确率从71.1%降至38.0%,权威性虚假信息攻击成功率达69.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11729 2026-06-16 cs.DS cs.AI cs.RO

Adapting Dijkstra for Buffers and Unlimited Transfers

为缓冲区和无限换乘调整Dijkstra算法

Denys Katkalo, Andrii Rohovyi, Toby Walsh

机构 * University of Oxford(牛津大学)

AI总结 本文提出Transfer Aware Dijkstra (TAD)算法,通过扫描完整行程序列而非单条边,解决了带缓冲区时间的无限换乘路径规划中传统Dijkstra过滤失效的问题,并在伦敦和瑞士网络上实现比MR快两倍以上的速度且保持最优性。

Comments v4: clarified RAPTOR description in the Background section

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05824 2026-06-16 eess.IV cs.CV cs.LG 版本更新

Navigating Distribution Shifts in Medical Image Analysis: A Survey

医学图像分析中的分布偏移导航:综述

Zixian Su, Jingwei Guo, Xi Yang, Qiufeng Wang, Frans Coenen, Amir Hussain, Kaizhu Huang

机构 * Life Simulation Research Center, Beijing Academy of Artificial Intelligence(北京人工智能生命模拟研究中心) Electrical and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology(王国阿卜杜勒·阿齐兹国王科技大学电气与数学科学与工程系) Department of Intelligent Science, School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院智能科学系) Computer Science, School of Computer Science and Informatics, University of Liverpool(利物浦大学计算机科学与信息学学院) SDAIA-KFUPM Joint Research Centre for Artificial Intelligence, King Fahd University of Petroleum and Minerals(法赫德石油与矿物大学人工智能SDAIA-KFUPM联合研究中心) Nuffield Department of Primary Care Health Sciences, University of Oxford(牛津大学初级保健健康科学努尔菲尔德部门)

AI总结 本文系统综述了应对医学图像分析中分布偏移的深度学习方法,按临床约束分类为联合训练、联邦学习、微调和域泛化,并揭示方法从显式对齐向不确定性建模的转变。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12670 2026-06-16 cs.AI 版本更新

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

SkillsBench: 基准测试智能体技能在不同任务中的有效性

Xiangyi Li, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Yifeng He, Shenghan Zheng, Kyoung Whan Choe, Jiankai Sun, Shuyi Wang, Chujun Tao, Binxu Li, Xuandong Zhao, Hejia Geng, Xiaojun Wu, Junwei Zhou, Xiaokun Chen, Hanwen Xing, Yubo Li, Qunhong Zeng, Di Wang, Yuanli Wang, Roey Ben Chaim, Penghao Jiang, Haotian Shen, Luyang Kong, Xinyi Liu, Runhui Wang, Xuanqing Liu, Jiachen Li, Xin Lan, Yueqian Lin, Wengao Ye, Junwei He, Songlin Li, Yue Zhang, Yipeng Gao, Yijiang Li, Ze Ma, Liqiang Jing, Tianyu Wang, Kaixin Li, Yiqi Xue, Haoran Lyu, Yizhuo He, Yuchen Tian, Shutong Wu, Bowei Wang, Yixuan Gao, Bo Chen, Litong Liu, Sikai Cheng, Jiajun Bao, Shuaicheng Tong, Shuwen Xu, Terry Yue Zhuo, Tinghan Ye, Qi Qi, Miao Li, Longtai Liao, Zelin Tan, Chang Shi, Xilin Tang, Srinath Tankasala, Boqin Yuan, Yaoyao Qian, Jianhong Tu, Chenguang Wang, Yizhou Sun, Wei Wang, Aaron Taylor, Ziyue Yang, Changkun Guan, Zhikang Dong, Xinyu Zhang, Steven Dillmann, Han-chung Lee, Dawn Song

机构 * BenchFlow OSU Amazon UC Berkeley UC Santa Cruz UC Davis Dartmouth RLWRLD Independent Princeton University Oxford University Stanford University USC CMU Foxconn Zenity UNSW UT Austin MSU Duke University ByteDance UT Dallas UC San Diego Columbia University University of Rochester Cornell Tech Georgia Tech Cornell University NEU UCLA Snap Inc. Fanshawe College University of Science and Technology of China HKUST(GZ) Anyscale

AI总结 提出SkillsBench基准,包含8领域87个任务,通过配对评估证明技能提升平均通过率16.6个百分点,小模型配备技能可匹敌大模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16902 2026-06-16 cs.AI cs.LG 版本更新

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?

LLM-WikiRace 基准测试:大语言模型在真实知识图谱上的规划能力有多强?

Juliusz Ziomek, William Bankes, Lorenz Wolf, Shyam Sundhar Ramesh, Xiaohang Tang, Ilija Bogunovic

机构 * University of Oxford, UK(牛津大学,英国) University College London (Centre for AI), UK(伦敦大学学院(人工智能中心),英国) University of Basel, Switzerland(巴塞尔大学,瑞士)

AI总结 提出 LLM-Wikirace 基准,通过维基百科超链接导航任务评估大语言模型的规划、推理与世界知识,发现模型在简单任务上超人类,但困难任务成功率仅 23%,且规划与长程推理是主要瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08026 2026-06-16 cs.LG stat.ML 版本更新

Sharp analysis of linear ensemble sampling

线性集成采样的尖锐分析

David Janz, Arya Akhavan, Csaba Szepesvári

机构 * University of Oxford, UK(牛津大学,英国) University of Alberta, Canada(阿尔伯塔大学,加拿大)

AI总结 本文针对随机线性bandits中的线性集成采样(ES)方法,证明当集成大小m=Θ(d log n)时,ES达到~O(d^{3/2}√n)的高概率遗憾,缩小了与汤普森采样基准的差距,同时保持计算量相当。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05779 2026-06-16 cs.LG cs.IT math.IT 版本更新

How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

如何控制方差以提高稀疏激活DNN和CNN的训练稳定性

Emily Dent, Jared Tanner

机构 * Mathematical Institute University of Oxford(牛津大学数学研究所)

AI总结 针对稀疏激活函数,提出增大高斯过程方差可提升训练稳定性,并设计新初始化策略实现隐藏层高达90%稀疏度的稳定训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00885 2026-06-16 cs.CV 版本更新

HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

HanDyVQA:面向细粒度手-物交互动态的视频问答基准

Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato, Takuma Yagi

机构 * Institute of Industrial Science, The University of Tokyo(东京大学工业科学研究所) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院) Waseda University(早稻田大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

AI总结 提出HanDyVQA基准,通过六类问题(11.1K QA对)和10.3K分割掩码,全面评估视频模型对手-物交互中操作与效果的细粒度时空推理能力,发现最佳模型Gemini-2.5-Pro仅73%准确率(人类97%)。

Comments CVPR 2026, Project page: https://masatate.github.io/HanDyVQA-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17431 2026-06-16 cs.CL 版本更新

Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning

用于搜索的智能体强化学习使指令微调对齐失效

Yushi Yang, Shreyansh Padarha, Sarah Ball, Andrew Lee, Adam Mahdi

机构 * University of Oxford(牛津大学) Ludwig-Maximilians-Universität München(慕尼黑路德维希-马克西米利安大学) Harvard University(哈佛大学)

AI总结 研究智能体强化学习(RL)对指令微调模型对齐的影响,发现RL训练使模型将有害请求转化为良性搜索查询,但在触发条件下产生多步不安全搜索行为,并提出了基于表示引导的RL训练方法恢复对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05370 2026-06-16 cs.LG 版本更新

Imbalanced Semi-Supervised Learning via Label Refinement and Threshold Adjustment

通过标签精炼和阈值调整实现不平衡半监督学习

Zeju Li, Ying-Qiu Zheng, Chen Chen, Saad Jbabdi

机构 * College of Biomedical Engineering, Fudan University, Shanghai, China(复旦大学生物医学工程学院,上海,中国) FMRIB Centre, Oxford Centre for Integrative Neuroimaging (OxCIN), University of Oxford, Oxford, UK(牛津大学FMRIB研究中心,牛津大学整合神经影像中心(OxCIN),牛津,英国) School of Computer Science, University of Sheffield, Sheffield, UK(谢菲尔德大学计算机科学学院,谢菲尔德,英国) Department of Engineering Science, University of Oxford, Oxford, UK(牛津大学工程科学系,牛津,英国)

AI总结 针对半监督学习在类别不平衡数据上性能下降的问题,提出SEVAL框架,通过从类别平衡的子集学习标签精炼和阈值调整参数,联合优化生成更准确的伪标签,在多种不平衡场景下超越现有方法。

Comments Accepted by Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14699 2026-06-15 cs.CV cs.GR cs.RO 新提交

Instruct-Particulate: Scaling Feed-Forward 3D Object Articulation with Kinematic Control

Instruct-Particulate: 基于运动学控制的可扩展前馈式3D物体关节化

Ruining Li, Yuxin Yao, Matt Zhou, Chuanxia Zheng, Christian Rupprecht, Joan Lasenby, Shangzhe Wu, Andrea Vedaldi

机构 * University of Oxford(牛津大学) University of Cambridge(剑桥大学) Nanyang Technological University(南洋理工大学)

AI总结 提出Instruct-Particulate模型,通过运动学规范(部件描述、连接性、关节类型等)指导3D网格的关节分割和运动参数预测,利用异构数据集(15万+物体)训练,实现跨类别和AI生成网格的泛化。

Comments Project page: https://instruct-particulate.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14561 2026-06-15 cs.RO cs.LG 新提交

ORCA: A Platform for Open-Source Dexterity Research

ORCA: 开源灵巧性研究平台

Francesco Capuano, Maximilian Eberlein, Fabrice Bourquin, Clemens Claudio Christoph

机构 * University of Oxford(牛津大学) ETH Zurich(苏黎世联邦理工学院) Orca Dexterity

AI总结 提出ORCA学习栈,统一灵巧手控制、仿真、遥操作和重定向,集成机器人学习框架,实现端到端灵巧操作研究。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏