arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Toronto(多伦多大学)

共收录 1020
2608.13482 2026-08-14 cs.LG cs.AI cs.CL 新提交

Synthetic Persona Pretraining: Alignment from Token Zero

合成角色预训练:从零开始的对齐

Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West

机构 * EPFL(洛桑联邦理工学院) University of Toronto(多伦多大学) Northeastern University(东北大学) SJTU(上海交通大学) Saarland University(萨尔大学) Hereon(亥姆霍兹极地与海洋研究中心) TUHH(汉堡工业大学) Ontocord AI(Ontocord人工智能公司) TUC(德累斯顿工业大学) DFKI(德国人工智能研究中心)

AI总结 该研究提出合成角色预训练(SPP),在预训练token零阶段植入助手角色,通过标注反思、预训练及角色绑定,提升模型价值构成遵循度与鲁棒性,降低对齐错误率,证明预训练时角色干预是对齐的有效方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12611 2026-08-14 cs.CV cs.LG 新提交

From Visual Widgets to UI Code: Efficient Tool-Grounded Generation

从视觉控件到UI代码:高效的基于工具的生成

Houston H. Zhang, Tao Zhang, Li Gu, Linfeng Ye, Yuanhao Yu, Xinxin Zuo, Yang Wang, Zhixiang Chi

机构 * McMaster University(麦克马斯特大学) University of Toronto(多伦多大学) Concordia University(康考迪亚大学)

AI总结 本研究提出轻量级工具框架WidgetGen,在6个多模态模型和1000个控件上,其视觉重建指标优于直接提示和Widget2Code,且重建的图像-代码对可提升Qwen系列模型性能

Comments ECCV2026 (MUCG Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17897 2026-08-14 cs.CV cs.AI cs.LG cs.RO 版本更新

RadarGen: Automotive Radar Point Cloud Generation from Cameras

RadarGen:从摄像头生成汽车雷达点云

Tomer Borreda, Fangqiang Ding, Sanja Fidler, Shengyu Huang, Or Litany

机构 * Technion(技术学院) MIT(麻省理工学院) NVIDIA(英伟达) University of Toronto(多伦多大学) Vector Institute(向量研究所)

AI总结 RadarGen通过扩散模型从摄像头图像生成逼真的雷达点云,结合BEV对齐的深度、语义和运动线索,提升雷达生成的物理合理性与多模态模拟能力。

Comments ECCV 2026. Project page: https://radargen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11434 2026-08-13 cs.AI cs.CL cs.CV 新提交

Benchmarking LLM Judges for Mobile Agent Evaluation

面向移动智能体评估的LLM评判基准测试

Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang

机构 * Mila – Québec AI Institute(米拉-魁北克人工智能研究所) Concordia University(康考迪亚大学) University of Toronto(多伦多大学) Shanghai University(上海大学) McMaster University(麦克马斯特大学)

AI总结 该研究推出MobileJudgeBench基准,评估6种LLM评判器方法在移动智能体轨迹上的可靠性,发现简单基线评判器具竞争力、基准质量指标可预测评判器效用,且不同LLM后端故障特征相反。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20963 2026-08-13 cs.RO

A Robotic Testing Platform for Pipelined Discovery of Resilient Soft Actuators

一种用于流水线发现鲁棒性软执行器的机器人测试平台

Ang Li, Alexander Yin, Alexander White, Sahib Sandhu, Matthew Francoeur, Victor Jimenez-Santiago, Van Remenar, Codrin Tugui, Mihai Duduta

机构 * Department of Mechanical and Industrial Engineering, University of Toronto(多伦多大学机械与工业工程系) Institute of Materials Science, University of Connecticut(康涅狄格大学材料科学研究所) School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut(康涅狄格大学机械、航空航天与制造工程学院) Material Science and Engineering, University of Connecticut(康涅狄格大学材料科学与工程系) Inorganic Polymers Department, Petru Poni Institute of Macromolecular Chemistry(彼得·波尼宏分子化学研究所无机聚合物部门)

AI总结 本文提出了一种机器人测试平台,用于流水线发现鲁棒性软执行器的最优参数组合,显著提升其操作寿命和负载能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10346 2026-08-12 cs.CV cs.AI 新提交

Towards Unified Dynamic Face Landmark Detection

面向统一动态人脸关键点检测

Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang, Parham Aarabi, Igor Gilitschenski

机构 * University of Toronto(多伦多大学) ModiFace(ModiFace公司)

AI总结 该研究针对人脸关键点检测需为不同N点数据集独立训练模型、仅能输出固定数量关键点的局限,提出FPALP概念与统一动态FLD方法,实现单模型适配多数据集、动态输出指定数量关键点,性能优于部分现有SOTA方法。

Comments 9 pages, 6 figures in Main Paper. 13 pages, 3 figures in Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00135 2026-08-12 cs.LG cs.AI 版本更新

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

论智能体工具调用与强化学习训练的有效性与效率

Tong Liu, Cheng Qian, Matej Cief, Yuan He, Daniele Dan, Nikolaos Aletras, Gabriella Kazai

机构 * University of California, Berkeley(加州大学伯克利分校) University of Cambridge(剑桥大学) University of Toronto(多伦多大学)

AI总结 本文系统分析工具调用评估中的实现选择对结果敏感性的影响,并针对强化学习训练中的计算浪费提出两种加速技术。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12231 2026-08-12 cs.LG 版本更新

Temporal Straightening for Latent Planning

时间拉直用于隐式规划

Ying Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, Mengye Ren

机构 * New York University(纽约大学) Brown University(布朗大学) University of Toronto(多伦多大学)

AI总结 受人类视觉处理中感知拉直假说启发,提出时间拉直方法,通过曲率正则化联合学习JEPA世界模型的编码器和预测器,改善隐式规划中的表示学习,使梯度规划更稳定并提高目标到达任务成功率。

Comments ICML2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13319 2026-08-12 cs.AI cs.HC 版本更新

Situation Graph Prediction for User Perspective Modeling

情境图预测:用户建模的结构化视角推断

Jisung Shin, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama

机构 * Flybits Labs, Creative AI Hub(Flybits实验室、创意人工智能中心) University of Toronto(多伦多大学) Toronto Metropolitan University(多伦多 Metropolitan 大学) MIT Media Lab(MIT媒体实验室)

AI总结 情境图预测通过结构化视角推断提升用户建模能力,揭示潜在状态推断比表层提取更困难。

Comments Accepted to PILA 2026: Workshop on Personal Intelligence in the Agentic AI Era, at KDD 2026, 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09277 2026-08-11 cs.AI cs.PL 新提交

P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation

P³:用于验证代码生成的程序与证明联合规划

Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li, Kaiyu Yang

机构 * Apodex Princeton University(普林斯顿大学) Caltech(加州理工学院) University of Toronto(多伦多大学)

AI总结 P³是一种用于验证代码生成的程序与证明联合规划的LLM智能体工作流,在三个基准上的求解率优于基线,还降低了API成本与运行时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08381 2026-08-11 cs.CV eess.SP 新提交

DoRF++: Spherical Representation Learning over Doppler Radiance Fields for Robust Wi-Fi Sensing

DoRF++:面向鲁棒Wi-Fi感知的多普勒辐射场球面表示学习

Navid Hasanzadeh, Shahrokh Valaee

机构 * University of Toronto(多伦多大学)

AI总结 本文针对Wi-Fi感知的跨用户泛化难题,提出DoRF++模型,将NeRF概念引入Wi-Fi感知,结合球面Transformer实现手势识别,在单多天线AP场景下的困难手势识别中性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26159 2026-08-11 cs.AI cs.CY cs.LG 版本更新

When benchmark inferences do not compose: Projectibility in AI evaluation

当基准推理无法组合:AI评估中的可投射性

Brett Reynolds

机构 * Humber Polytechnic(汉伯理工学院) University of Toronto(多伦多大学)

AI总结 本文针对AI评估中基准推理无法组合的问题,提出非组合原则,结合古德曼的竞争延伸问题与基于论证的有效性框架,通过案例和模拟开发可投射性审计以诊断基准到应用论证的衔接缺陷。

Comments 34 pages, 2 figures, 5 tables. v2 substantially revises Secs. 5-8 and the conclusion, adds a measured instance of factor-structure instability, and corrects a claim in Sec. 3.3 that endpoint alignment suffices for composition. Supersedes the withdrawn arXiv:2510.15236. Code: https://github.com/BrettRey/benchmark-inference-composition

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04412 2026-08-11 cs.AI 版本更新

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

大语言模型作为导师:不可验证强化学习中的策略感知提示适应

Yujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song

机构 * KAIST(韩国科学技术院) Upstage University of Toronto(多伦多大学) Carnegie Mellon University(卡内基梅隆大学) NVIDIA(英伟达)

AI总结 针对不可验证强化学习中训练提示静态导致的问题,提出LLM-as-a-Tutor框架,让大语言模型从评判扩展为导师,通过对比策略展开检测无挑战性提示并添加约束,提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29064 2026-08-11 cs.CL cs.CV cs.HC cs.MA 版本更新

Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

分析多模态大语言模型代理在城市感知中生成解释的角色效应

Neemias da Silva, Matt Ratto, Myriam Delgado, Rodrigo Minetto, Daniel Silver, Thiago H Silva

机构 * Universidade Tecnologica Federal do Parana(巴西南里奥格兰德联邦技术大学) University of Toronto(多伦多大学)

AI总结 通过对比不同角色提示和无角色设置下多模态大语言模型生成的文本,发现标题描述趋同,但理由描述随社会经济和政治属性系统变化,感知标签无显著差异。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23341 2026-08-11 cs.CR cs.AI 版本更新

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

评估部署于智能电网操作中的LLM jailbreaking漏洞:与NERC标准的基准测试

Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber, Amr Youssef, Deepa Kundur

机构 * ECE Department, University of Toronto(多伦多大学电子工程系) CIISE, Concordia University(麦吉尔大学CIISE)

AI总结 本文评估了智能电网操作中部署LLM的jailbreaking漏洞,通过与NERC标准的基准测试,发现DeepInception方法攻击成功率最高,Claude 3.5 Haiku完全免疫,Gemini 2.0 Flash-Lite最易受攻击。

Comments \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25921 2026-08-11 cs.CL cs.CR 版本更新

One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

逐词进行:增量完成分解打破LLM安全

Samee Arif, Naihao Deng, Zhijing Jin, Rada Mihalcea

机构 * University of Michigan(密歇根大学) University of Toronto(多伦多大学)

AI总结 本文提出增量完成分解(ICD)策略,通过逐词生成恶意请求相关词来突破LLM安全机制,评估多种变体在多个基准测试中表现优异,并理论解释其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17632 2026-08-11 cs.LG cs.AI 版本更新

SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer

SMAC: 基于分数匹配的演员-评论家用于鲁棒的离线到在线迁移

Nathan Samuel de Lara, Florian Shkurti

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所)

AI总结 SMAC通过正则化Q函数,使演员-评论家在离线到在线RL迁移中保持性能,有效避免性能下降。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21891 2026-08-11 cs.CL cs.AI cs.LG stat.ME stat.ML 版本更新

Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation

嵌入信任:语义各向同性预测长文本生成中的非事实性

Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner

机构 * New York University(纽约大学) University of Toronto(多伦多大学)

AI总结 该研究提出通过语义各向同性(单位球面上归一化文本嵌入的均匀程度)评估LLMs生成长文本的可信度,其方法无需标注数据等,在多领域仅用少量样本预测非事实性的表现优于现有信号。

Comments Published in Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03094 2026-08-11 cs.CV astro-ph.IM cs.LG physics.ao-ph 版本更新

NeuralDMD: Interpretable Neural Representation of Dynamics from Sparse and Noisy Measurements

NeuralDMD:基于稀疏且含噪测量的可解释动力学神经表示

Ali SaraerToosi, Renbo Tu, Esther Y. H. Lin, Kamyar Azizzadenesheli, Aviad Levis

机构 * University of Toronto(多伦多大学) NVIDIA Corporation(NVIDIA公司)

AI总结 NeuralDMD是结合神经隐式表示与DMD的可解释未训练重建框架,可从稀疏含噪测量中直接重建预测时空动力学,在天气、黑洞观测等任务上优于基线,线性场景下稳定,非线性场景仍有应用潜力。

Comments 53 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20532 2026-08-11 cs.LG stat.ME stat.ML 交叉投稿

One-shot Robust Federated Learning of Independent Component Analysis

独立成分分析的单轮鲁棒联邦学习

Dian Jin, Xin Bing, Yuqian Zhang

机构 * Department of Electrical and Computer Engineering, Rutgers University, New Brunswick(罗格斯大学电气与计算机工程系) Department of Statistical Sciences, University of Toronto(多伦多大学统计学系)

AI总结 针对联邦独立成分分析问题,提出基于k-means聚类与几何中位数的单轮鲁棒聚合算法,在异构场景下通过仿真验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07408 2026-08-10 cs.CV cs.LG 新提交

Addressable Memory for Video World Models

视频世界模型的可寻址内存

Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljoša Ošep

机构 * NVIDIA(英伟达) Princeton University(普林斯顿大学) University of Toronto(多伦多大学) Vector Institute(矢量研究院)

AI总结 针对交互式视频世界模型超出训练时序后内存寻址失效及压缩缓存破坏内存的问题,提出无训练框架 WorldTrace,含两种压缩方法,在新基准 LoopBench 上分别提升时序一致性 15.5%、情景回忆 19.5%。

Comments Project page: https://research.nvidia.com/labs/sil/projects/WorldTrace/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07385 2026-08-10 cs.LG cs.AI eess.AS eess.SP stat.ML 新提交

Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations

全模态分解自编码器学习全栈可穿戴解耦表示

Ioannis Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis, Dimitrios Hatzinakos

机构 * Khalifa University(哈利法大学) University of Toronto(多伦多大学) MIT Media Lab(麻省理工学院媒体实验室) Aristotle University of Thessaloniki(亚里士多德大学)

AI总结 该研究针对现有多模态可穿戴模型的不足,提出OmniDecVAEs框架,在30模态的HAR任务中,提升了识别准确率与数据合成质量,可用于边缘可穿戴与医疗领域。

Comments 15 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06956 2026-08-10 cs.LG physics.chem-ph 新提交

How Molecular Generative Models Organize Molecular Identity

分子生成模型如何组织分子身份

Raul Ortega-Ochoa, Tejs Vegge, Jens S. Bakander, Luis Mantilla Calderon, Alan Aspuru-Guzik, Tonio Buonassisi

机构 * Toyota Research Institute(丰田研究所) Technical University of Denmark(丹麦技术大学) CAPeX Pioneer Center for Accelerating P2X Materials Discovery(CAPeX P2X材料发现加速先锋中心) University of Toronto(多伦多大学) Vector Institute for Artificial Intelligence(向量人工智能研究所) NVIDIA(英伟达) Massachusetts Institute of Technology(麻省理工学院) Acceleration Consortium(加速联盟)

AI总结 该研究通过明确分子身份并将其通过生成过程拉回,揭示了三种分子生成架构的内部储备呈分段常数区域,其组织受多种因素影响,需先表征内部组织才能将生成空间视为可化学导航。

Comments 22 pages (14 main text + 8 supporting information)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15506 2026-08-10 cs.LG cs.AI 版本更新

Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains

寻求SOTA:时间序列预测必须采用领域特定评估以消除虚幻收益

Raeid Saqur, Christoph Bergmeir, Blanka Horvath, Daniel Schmidt, Frank Rudzicz, Terry Lyons

机构 * Dept. of Computer Science, University of Toronto(多伦多大学计算机科学系) Dept. of Data Science & AI, Monash University, Australia(澳大利亚墨尔本大学数据科学与人工智能系) Faculty of Computer Science, Dalhousie University(达尔豪斯大学计算机科学学院) Dept. of Mathematics, University of Oxford(牛津大学数学系) Oxford-Man Institute for Quantitative Finance(牛津-曼彻斯特量化金融研究所) Vector Institute(向量研究所) Department of Computer Science and AI, University of Granada, Spain(西班牙格拉纳达大学计算机科学与人工智能系) DaSCI, Andalucía, Spain(安达卢西亚DaSCI)

AI总结 本文指出当前时间序列预测评估方法掩盖了真实进展,呼吁引入更广泛的非平稳性数据集并要求深度学习模型包含经典基线,以确保报告的改进反映真正的科学进步。

Comments v2 clarifies Transformer temporal-order claims, strengthens benchmark-selection and metric guidance, corrects point-forecast targets for MSE/MAE, improves aggregation/reporting recommendations, adds living-benchmark protocols, revises weather/evaluation wording, makes author emails clickable, and adds five supporting references

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05410 2026-08-07 cs.RO 新提交

Sliding Sensors: Configurable Confidence in State Estimation for Continuum Robots

滑动传感器:连续体机器人状态估计的可配置置信度

Ella Walsh, Spencer Teetaert, Eric Diller, Timothy D. Barfoot, Jessica Burgner-Kahrs

机构 * University of Toronto Robotics Institute(多伦多大学机器人研究所)

AI总结 本研究提出可机械重构的滑动传感器设计,通过调整其在连续体机器人内的位置优化状态估计置信度,可降低全身形状估计误差,为连续体机器人的安全交互提供支撑。

Comments Accepted as an extended abstract at 2026 IEEE 9th International Conference on Soft Robotics (RoboSoft). * Equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28617 2026-08-07 cs.AI cs.CL cs.CY cs.HC 版本更新

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

AISPA:面向大语言模型应用的以用户为中心的系统提示审计框架

Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei

机构 * Stanford University(斯坦福大学) CMU(卡内基梅隆大学) UT Austin(德克萨斯大学奥斯汀分校) University of Toronto(多伦多大学) UCSB(加利福尼亚大学圣巴巴拉分校) WashU(华盛顿大学) OSU(俄亥俄州立大学) UCSC(加利福尼亚大学圣克鲁兹分校) Northwestern University(西北大学) UIUC(伊利诺伊大学厄巴纳-香槟分校) KAUST(阿卜杜拉国王科技大学) MIT(麻省理工学院) University of Oxford(牛津大学) Institute for Decentralized AI(去中心化人工智能研究所)

AI总结 本文提出以用户为中心的AISPA框架,审计88款商业AI产品的3249条系统提示指令,发现其设计差异大、保护指令范围浅、长度增长但仍存问题指令,凸显系统提示需更高透明度与监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03699 2026-08-07 cs.RO cs.SY eess.SY 版本更新

Lost in Time? Continuous Symmetry and Identifiability in Aided Inertial Navigation with Unknown Measurement Delays

迷失在时间中?具有未知测量延迟的辅助惯性导航中的连续对称性和可识别性

Jonathan Kelly, Phone Thiha Kyaw, Mattew Giamou

机构 * Space & Terrestrial Autonomous Robotic Systems (STARS) Laboratory at the University of Toronto Institute for Aerospace Studies (UTIAS)(多伦多大学航天研究所空间与地面自主机器人系统(STARS)实验室) Autonomous Robotics & Convex Optimization (ARCO) Laboratory in the Department of Computing and Software, McMaster University(麦克马斯特大学计算与软件系自主机器人与凸优化(ARCO)实验室)

AI总结 研究辅助导航中单个辅助传感器测量相对于惯性测量流有未知但恒定延迟时系统的可识别性,利用特殊伽利略群刻画无信息轨迹并与延迟测量模型连续对称性相关,揭示可识别性失败轨迹类别及与线性化分析联系。

Comments Accepted to the IEEE International Conference on Multisensor Fusion and Integration (MFI), Pilsen, Czechia, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16763 2026-08-07 cs.AI 版本更新

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

当AI基准测试达到平台期:基准饱和的系统性研究

Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman

机构 * University of California, Berkeley(加州大学伯克利分校) University of Toronto(多伦多大学) University of Washington(华盛顿大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Michigan(密歇根大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本研究定义并分析了60个语言模型基准的饱和现象,发现近半数基准出现饱和,且专家策划而非公开测试数据影响抗饱和能力,为延长基准寿命提供了设计建议。

Comments Published at ICML 2026 (Forty-Third International Conference on Machine Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05139 2026-08-06 cs.CL cs.LG 新提交

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

面向技能原生的大语言模型:用于基准测试和训练长程推理的技能熵

Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora

机构 * Princeton University(普林斯顿大学) Carnegie Mellon University(卡内基梅隆大学) University of Toronto(多伦多大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stanford University(斯坦福大学) University of Oxford(牛津大学)

AI总结 该研究针对现有基准无法评估LLM跨技能长程推理能力的问题,提出Skill Entropy(技能熵)并构建Skill²-Bench基准,还开发Skill-Entropy RL训练框架,显著提升了Qwen3模型在该基准上的表现。

Comments https://github.com/Gen-Verse/Skill-Entropy-RL

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04147 2026-08-06 cs.LG cs.AI cs.CV 新提交

LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling

LiNC:基于逐样本置信度与高斯混合模型的轻量级噪声校正

Abhishek Moturu, Babak Taati, Anna Goldenberg

机构 * University of Toronto(多伦多大学) The Hospital for Sick Children(病童医院) UHN KITE Research Institute(大学健康网络KITE研究所) Vector Institute(向量研究所) Institute of Biomedical Engineering(生物医学工程研究所) Rehabilitation Sciences Institute(康复科学研究所) Department of Laboratory Medicine and Pathobiology(检验医学与病理生物学系)

AI总结 针对医学成像数据集的标签噪声问题,本文提出LiNC方法,通过逐样本置信度参数结合高斯混合模型校正标签,在MedMNISTv2数据集上取得稳定的准确率提升且开销极小。

详情

展开后加载摘要…

URL PDF HTML 收藏