arXivDaily arXiv每日学术速递 周一至周五更新
全部学科分类 1585
2607.22535 2026-07-27 cs.RO cs.CV 新提交

Robot-Factored World Models via Robot Rendering

通过机器人渲染实现机器人因素分解的世界模型

Byungjun Kim, Taeksoo Kim, Hyunsoo Cha, Hanbyul Joo

机构 * Seoul National University(首尔国立大学)

AI总结 研究动作条件视频世界模型,提出机器人因素分解的世界模型,通过动作实现和机器人渲染将特定机器人因素移出模型,解决深度模糊问题,实验表明其优于基线且能推广,还能从人类演示生成机器人操作视频。

Comments Project Page: https://bjkim95.github.io/rofacto/

详情
AI中文摘要

动作条件视频世界模型根据初始观测和动作信号预测未来观测。在机器人技术中,动作通过两个不同过程影响未来观测:先由机器人身体和控制器转化为机器人运动,然后场景通过接触和物体运动做出响应。直接以动作命令为条件要求世界模型学习实现过程本身,而以记录的未来状态为条件会泄露其要预测的交互结果。我们提出机器人因素分解的世界模型,将两个特定于机器人的因素移出世界模型。一是动作实现,将每个命令通过机器人自身控制器和运动学转化为可部署的名义轨迹,避免动作实现学习和未来状态泄露。二是机器人渲染,通过机器人URDF渲染名义轨迹,将机器人几何、运动学和外观从模型中分解出来。为解决深度模糊问题,将末端执行器深度与场景深度配对。实验表明渲染接口优于向量条件基线,并能推广到推理时未见的机器人实例。还证明该模型可通过将手部动作重定向并渲染为机器人几何形状,从人类演示中生成机器人操作视频。

英文摘要

Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence future observations through two distinct processes: they are first realized into robot motion by the robot body and controller, and the scene then responds through contact and object motion. Conditioning directly on action commands asks the world model to learn the realization process itself, while conditioning on logged future states leaks the interaction outcomes it is meant to predict. We propose robot-factored world models, which move two robot-specific factors outside the world model. First, action realization: each command is rolled through the robot's own controller and kinematics into a deployment-available nominal trajectory, a middle signal that avoids both action-realization learning and future-state leakage. Second, robot rendering: this nominal trajectory is rendered through the robot URDF, factoring the robot's geometry, kinematics, and appearance out of the model and into explicit rendered robot geometry. To resolve depth ambiguity, we pair end-effector depth with scene depth, giving geometric cues for contact and occlusion beyond image-plane overlap. Together, camera-aware static RGB/depth context and rendered robot geometry form a shared visual world-model interface that stays consistent across viewpoints and robot embodiments, so the model sees the action only as visible robot geometry and learns how objects respond to it. Our experiments show that the rendered interface outperforms vector-conditioned baselines and generalizes to unseen robot embodiments at inference. We further demonstrate that our model generates robot manipulation videos from human demonstrations by retargeting and rendering the hand motion as robot geometry.

URL PDF HTML 收藏
2607.22534 2026-07-27 cs.CV cs.AI cs.RO 新提交

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

SM4RT:学习用于4D重建的结构化运动几何

Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, Yuqi Wu, Jie Zhou, Jiwen Lu

机构 * Intelligent Vision Group, Tsinghua University(清华大学智能视觉组)

AI总结 针对将单目3D重建能力扩展到4D动态理解的挑战,提出SM4RT,通过引入运动结构表示场景动态,利用并行运动几何编码器和解码器,从单目RGB视频中联合推断相关信息,实现强大运动重建性能并保留场景运动几何结构。

Comments Code is available at: https://github.com/wzzheng/SM4RT

详情
AI中文摘要

几何基础模型(GFMs)极大地推动了单目3D重建,但将此能力扩展到4D动态理解仍是一项重大挑战。现有多数运动感知方法将运动视为独立的逐点位移,忽略了物理运动的结构化本质。而实际物体通常遵循刚体运动学规律,点通常集体移动而非孤立移动。基于此,我们提出了SM4RT,一种用于端到端3D重建和结构化运动感知的结构化运动4D重建变换器。SM4RT引入运动结构来表示场景动态,将场景运动分解为一组紧凑的运动基,每个运动基表示为SE(3)中6D扭转的时间序列。然后通过对这些基的稀疏、逐像素时间共享分配权重来恢复密集场景运动。SM4RT引入了并行运动几何编码器和解码器,可从单目RGB视频中单次前向传递联合推断3D几何、世界坐标运动和场景运动学结构。SM4RT在保留场景运动几何结构的同时实现了强大的运动重建性能。

英文摘要

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body transformations governed by SE(3), rather than unstructured point-wise displacements. Building on this insight, we propose SM4RT, a Structured Motion 4D Reconstruction Transformer for end-to-end 3D reconstruction and structured motion perception. SM4RT introduces Structure-of-Motion to represent scene dynamics, where scene motion is decomposed into a compact set of motion bases, each represented as a temporal sequence of 6D twists in SE(3). Dense scene motion is then recovered by sparse, time-shared per-pixel assignment weights over these bases, ensuring points on the same object share a common rigid-body motion trajectory. SM4RT introduces a parallel motion geometry encoder and decoder that jointly infer 3D geometry, world-coordinate motion, and scene kinematic structure in a single forward pass from monocular RGB video. SM4RT achieves strong motion reconstruction performance while preserving the geometric structure of scene motion.

URL PDF HTML 收藏
2607.22531 2026-07-27 cs.CV 新提交

Twins: Learn to Predict Unified Representations with Focal Loss

Twins:使用焦点损失学习预测统一表示

Kaixiong Gong, Xin Cai, Bin Lin, Hao Wang, Yunlong Lin, Mingzhe Zheng, Bohao Li, Jian-Wei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Xiangyu Yue

机构 * Tencent-Hunyuan(腾讯混元)

AI总结 研究统一多模态模型潜在空间不匹配问题,提出Twins统一连续令牌空间。因联合建模有优化不平衡,采用焦点回归目标解决,在ImageNet上有gFID增益,在多模态理解基准测试中表现好,还提高了重建保真度。

Comments ICML 2026. Code: https://github.com/Tencent-Hunyuan/Twins

详情
AI中文摘要

统一多模态模型寻求支持多模态理解和图像生成的共享视觉令牌空间。离散方法通过共享码本统一接口,而连续管道通常依赖两种不同的表示——用于理解的语义特征(如ViT)和用于合成的低级潜在特征(如VAE),导致潜在空间不匹配。我们提出了Twins,一个通过在同一令牌网格上按通道连接ViT和VAE特征形成的统一连续令牌空间,序列长度不变且注意力成本不增加。然而,在扩散变换器中联合建模Twins会出现严重的优化不平衡:模型能很好地拟合ViT组件,但难以匹配VAE潜在分布。我们将这种不平衡追溯到三个异质性来源:频率偏差、内在维度以及条件对齐与条件独立的不确定性。为了解决这个问题,我们采用焦点回归目标进行流匹配,对大误差的VAE维度进行加权,更好地平衡ViT和VAE组件之间的优化。在ImageNet上,与无分类器指导的朴素MSE损失相比,这带来了高达10.57的gFID增益。Twins在多模态理解基准测试中也具有竞争力,并提高了重建保真度,缩小了面向理解和生成的表示之间的差距。

英文摘要

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations -- semantic features (e.g., ViT) for understanding and low-level latents (e.g., VAE) for synthesis -- resulting in mismatched latent spaces. We propose Twins, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase. However, jointly modeling Twins in a Diffusion Transformer exposes a severe optimization imbalance: the model fits the ViT component well but struggles to match the VAE latent distribution. We trace this imbalance to three sources of heterogeneity: frequency bias, intrinsic dimensionality, and condition-aligned vs condition-independent uncertainty. To address it, we adapt a focal regression objective for flow matching that upweights large-error VAE dimensions, better balancing optimization across the ViT and VAE components. On ImageNet, this yields up to 10.57 gFID gain over naive MSE loss without classifier-free guidance. Twins also performs competitively on multimodal understanding benchmarks and improves reconstruction fidelity, narrowing the gap between understanding- and generation-oriented representations.

URL PDF HTML 收藏
2607.22530 2026-07-27 cs.RO 新提交

ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

ViTacWorld:用于丰富接触式机器人操作的视觉-触觉世界模型扩展

Yunao Huang, Shiyu Sang, Haotao Lu, Suting Ni, Shijie Wu, Ziyang Guo, Ye Shi, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) InstAdapt

AI总结 研究针对丰富接触式机器人操作中视觉-触觉学习扩展难的问题,提出ViTacWorld模型,利用真实和模拟数据预训练并微调,能根据机器人动作预测视觉与触觉反馈,实现动作条件策略评估,提升策略性能。

Comments 18 pages, 6 figures, 5 tables. Project page: https://vitacworld.github.io/

详情
AI中文摘要

丰富接触式机器人操作需要物理交互线索,触觉感应至关重要,但扩展视觉-触觉机器人学习困难。我们提出ViTacWorld,一个用于可扩展丰富接触式机器人操作的动作条件视觉-触觉世界模型。它利用公共真实触觉数据集和构建的模拟环境扩展视觉-触觉-动作数据,先大规模预训练,再用真实策略展开微调。给定机器人动作,它能预测视觉观察和触觉反馈,生成视觉-触觉-动作展开。实验表明它能生成有物理意义的展开,通过可扩展数据增强改善策略性能并实现动作条件策略评估。

英文摘要

Contact-rich robot manipulation requires physical interaction cues that are often invisible to cameras, making tactile sensing essential for robust control. However, scaling visuo-tactile robot learning remains difficult because real tactile interaction data are expensive to collect, hardware-dependent, and limited in task and scene diversity. We present ViTacWorld, an action-conditioned visuo-tactile world model for scalable contact-rich robot manipulation. ViTacWorld leverages public real tactile datasets and a constructed simulation environment to scale visuo-tactile-action data, exploiting the fact that tactile signals are directly grounded in physical contact and can exhibit a smaller simulation-to-real gap than purely visual observations. The model is first pretrained with large-scale real and simulated visuo-tactile trajectories, and then finetuned with real-world policy rollouts to better match downstream manipulation behaviors. Given robot actions, ViTacWorld predicts temporally aligned visual observations and tactile feedback, enabling visuo-tactile-action rollout generation. To the best of our knowledge, ViTacWorld is the first framework that uses a world model for robot visuo-tactile-action trajectory generation and policy evaluation. It serves two roles: synthesizing rollouts to improve downstream tactile policies, and evaluating policies by predicting action-conditioned visuo-tactile outcomes under controlled action sequences. Experiments on contact-rich manipulation tasks show that ViTacWorld generates physically meaningful rollouts, improves policy performance through scalable data augmentation, and enables action-conditioned policy evaluation. Project page: https://vitacworld.github.io/

URL PDF HTML 收藏
2607.22525 2026-07-27 cs.AI 新提交

Explainable Reinforcement Learning for assisting Air Traffic Controllers

用于协助空中交通管制员的可解释强化学习

Anduel Mehmeti, Gabriella Gigante, Salvatore Venticinque

机构 * Department of Engineering, University of Campania "Luigi Vanvitelli"(工程学院,坎帕尼亚"路易吉·范维蒂利"大学) CIRA(CIRA研究院)

AI总结 探索可解释性技术在强化学习算法中的应用,以协助空中交通管制员。用强化学习算法在简化ATC环境训练智能体决策避禁飞区路线,采用显著性图作为初步可解释性方法,为智能体决策提供关键输入特征见解。

Comments 11 pages (10-page paper plus 1 cover/citation page), 4 figures, 1 table; published in AINA 2025

Journal ref In: L. Barolli (ed.), Advanced Information Networking and Applications (AINA 2025), Lecture Notes on Data Engineering and Communications Technologies, vol. 250, pp. 148-157, Springer, Cham (2025)

详情
AI中文摘要

为了有效地将人工智能集成到医疗、自动驾驶和航空等高风险关键环境中,并朝着更高水平的自动化和无缝人机协作迈进,建立对人工智能驱动解决方案的信任至关重要。而信任又与人工智能系统的可解释性密切相关。人工智能在各个领域的快速发展凸显了建立信任的挑战,在应用于深度学习时,对人工智能可解释性的兴趣日益增加。在此背景下,本工作旨在探索可解释性技术在强化学习(RL)算法中的应用,特别是在安全关键的空中交通管制(ATC)领域。使用简化的ATC环境作为初始测试平台,用强化学习算法训练智能体,以对避开禁飞区的替代飞行路线做出决策。作为一种初步的可解释性方法,采用了显著性图,以深入了解对智能体决策过程影响最大的输入特征。

英文摘要

To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviation--and to advance toward higher levels of automation and seamless human-AI collaboration--building trust in AI-driven solutions is essential. Trust, in turn, is closely linked to the explainability of AI systems. The rapid advancements in AI across various domains have underscored the challenges of establishing trust, raising increasing interest in AI explainability even more when applied to deep learning. In this context, the present work aims to explore the application of explainability techniques to Reinforcement Learning (RL) algorithms, specifically within the safety-critical domain of Air Traffic Control (ATC). Using a simplified ATC environment as an initial testbed, an intelligent agent is trained with a reinforcement learning algorithm to make decisions on alternative flight routes that avoid no-fly zones. As a preliminary explainability approach, a saliency map is employed, providing insights into the input features that most significantly influence the agent's decision-making process.

URL PDF HTML 收藏
2607.22520 2026-07-27 cs.AI 新提交

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

回归税:剖析技能对大语言模型智能体产生帮助和伤害的原因

Darshan Tank, Baran Nama

机构 * Sentient Labs(森蒂恩实验室)

AI总结 研究剖析大语言模型智能体中技能产生帮助和伤害的原因,通过对比实验区分回归和残余失败两种结果,识别出回归的三个原因,发现现有技能过度强调程序指导,指出应分解技能净效应评估,还识别出应避免的回归模式及可靠性的关键因素。

详情
AI中文摘要

在大语言模型智能体中添加程序技能通常通过任务成功率的平均提升来评估。然而,这一指标掩盖了一个重要成本:技能也可能使智能体表现更差。我们通过在跨越两个办公自动化基准和三个模型框架栈的近6000次运行中比较有技能和无技能的智能体来衡量这两方面。这使我们能区分两种结果:回归是指无技能时能解决但添加技能后失败的任务;残余失败是指有技能和无技能时都失败的任务。我们发现回归情况很显著,最佳性能的技能主要通过更少回归而非更多提升来超越其他技能。我们识别出回归的三个原因:技能描述渗透、基础位移和验证位移。分析持续失败揭示了相同的潜在模式。现有技能过度强调程序指导而忽视基础和验证。纠正评估工件并研究痕迹后,我们发现许多回归和持续失败可通过更好的基础和验证来恢复。程序技能应通过将其净效应分解为提升和回归来评估,而不仅仅是总体改进。我们识别出技能应避免的三种回归模式,并发现可靠性更多取决于基础和验证而非程序技能选择。

英文摘要

Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but failed after skills are added. A residual failure is a task that fails both with and without skills. We find that regressions are substantial enough that the best performing skills outperform others primarily by regressing less, not by gaining more. We identify three causes of regression: (i) skill description osmosis, a skill changes an agent's behavior simply by being present in context, even when it is never invoked; (ii) grounding displacement, a skill's prescribed procedure overrides how the agent interprets its inputs; and (iii) verification displacement, where the procedure suppresses checks the agent would otherwise perform on its outputs. Analysing persistent failures reveals the same underlying pattern. Existing skills overemphasize procedural guidance the stage least often responsible for failure while under supporting grounding and verification, the dominant sources of remaining errors. After correcting evaluation artifacts and studying traces, we find many regressions and persistent failures recoverable through better grounding and verification. Procedural skills should be evaluated by decomposing their net effect into gains and regressions, not by aggregate improvement alone. We identify three regression modes skills should avoid, and find that reliability depends more on grounding and verification than on procedural skill choice.

URL PDF HTML 收藏
2607.22491 2026-07-27 cs.LG 新提交

Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting

用于状态条件波动率预测的易感性蓄水池架构

Aliaksei Kaliutau

机构 * Monodromy(英国Monodromy公司)

AI总结 研究针对波动率预测中非线性模型可利用结构有限的问题,提出易感性架构(SUSA)及具体实现,结合复值蓄水池与状态条件专家,在Qiskit中实现q比特对应物,经实验评估,模型在与GARCH竞争及补充HARQ预测方面有良好表现。

详情
AI中文摘要

波动率预测受持续性和测量噪声主导,留给非线性模型利用的残余结构有限。我们引入易感性架构(SUSA),一种用于波动率预测的蓄水池设计原则及其两种具体实现,基于复值开链和周期性蓄水池以及状态条件专家来解释平静、起始、恢复和持续压力状态下的蓄水池特征。我们还在Qiskit中实现了开放系统q比特对应物,同时保留通用的AR - 岭锚点和在QLIKE下训练的有界残余校正。我们使用三个不相交的按时间顺序排列的训练、验证和测试折、12个观测值的输入窗口和5个观测值的预测范围,对16个美国股票和交易所交易基金序列评估模型。所提出的模型与GARCH竞争,对特定资产(IWM,XLP)实现了统计学上显著的QLIKE改进。模型预测还补充了HARQ风格的预测:堆叠集成比其最强组成部分的平均QLIKE提高了0.0116,并在75%的测试场景中获胜。

英文摘要

Volatility forecasting is dominated by persistence and measurement noise, leaving limited residual structure for nonlinear models to exploit. We introduce Susceptible Architectures (SUSA), a reservoir-design principle for volatility forecasting, and its two concrete implementations, based on complex-valued open-chain and periodic reservoirs and regime-conditioned experts to interpret reservoir features across calm, onset, recovery, and persistent-stress states. We also implement open-system $q$-qubit counterparts in Qiskit while retaining a common AR-Ridge anchor and a bounded residual correction trained under QLIKE. We evaluate models on 16 U.S. equity and exchange-traded-fund series using three disjoint chronological training, validation, and test folds, a 12-observation input window, and a five-observation forecast horizon. The proposed models perform competitively with GARCH, achieving statistically significant QLIKE improvements for specific assets (IWM, XLP). Also models' forecasts complement HARQ-style predictions: a stacked ensemble improves mean QLIKE by 0.0116 over its strongest constituent and wins in 75% of test scenarios.

URL PDF HTML 收藏
2607.22489 2026-07-27 cs.LG cs.AI 新提交

\k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating

κ-LoRA:条件数揭示哪些LoRA矩阵值得更新

Jianghui Wang, Silong Yong, Francesco Orabona, Marco Canini, Katia P. Sycara, Yaqi Xie

机构 * King Abdullah University of Science and Technology(阿卜杜拉国王科技大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究指出LoRA统一更新矩阵计算成本高,条件数大的矩阵对性能提升贡献大。提出κ-LoRA方法,聚焦更新条件数大的矩阵,可减半可训练参数数量,降低计算和内存成本,实验证明其能缩短微调时间、降低内存成本且不影响精度。

详情
AI中文摘要

低秩自适应(LoRA)已成为神经网络高效微调的广泛采用技术,将模型更新分解为低秩矩阵。然而,LoRA计算成本高,因为它统一更新所有矩阵,而不考虑其对自适应的实际贡献。对于数十亿参数的大规模模型以及边缘部署和设备上微调等资源受限设置,此成本尤其高昂。我们首次表明并非所有LoRA矩阵都同样值得调整:条件数较小的矩阵在各方向上已平衡良好,对自适应贡献小;条件数大的矩阵包含欠发达方向,跨越更丰富子空间并推动大部分性能提升。基于此,我们提出κ-LoRA,通过将更新聚焦于条件数最大的矩阵来优化LoRA。通过将LoRA更新限制在按条件数排名前50%的权重矩阵,κ-LoRA将可训练参数数量减半,相应降低计算和内存成本。多个基准测试的广泛实验表明,该设计平均将微调时间缩短16.2%,同时匹配标准LoRA的精度并将内存成本降低4.5%。进一步分析表明所选矩阵的条件数在训练过程中持续下降,这表明κ-LoRA的有效性源于有针对性的谱重新平衡而非仅参数选择。

英文摘要

Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regardless of their actual contribution to adaptation. This cost is especially prohibitive for large-scale models with billions of parameters and for resource-constrained settings such as edge deployment and on-device fine-tuning. We show for the first time that not all LoRA matrices are equally worth tuning: matrices with smaller condition numbers (the ratio of largest to smallest singular value) are already well-balanced across directions and contribute only marginally to adaptation, whereas matrices with larger condition numbers contain underdeveloped directions that span richer subspaces and drive most of the performance gains. This observation itself is a key contribution of our work, and it motivates a more selective approach to fine-tuning. Building on this insight, we propose \k{appa}-LoRA, a method that optimizes LoRA by focusing updates on the matrices with the largest condition numbers, which capture the most informative directions of change. By restricting LoRA updates to the top 50% of weight matrices ranked by condition number, \k{appa}-LoRA halves the trainable parameter count and correspondingly reduces compute and memory cost. Extensive experiments across multiple benchmarks show that this design cuts fine-tuning time by 16.2% on average while matching the accuracy of standard LoRA and reducing memory cost by 4.5%. Further analysis reveals that the condition numbers of the selected matrices consistently decrease over training, suggesting that \k{appa}-LoRA's effectiveness stems from targeted spectral rebalancing rather than parameter selection alone.

URL PDF HTML 收藏
2607.22483 2026-07-27 cs.RO 新提交

Plug, Play, and Comply: A Modular Framework for Online Variable Impedance with Arbitrarily Oriented Compliance Axes

即插即用且合规:具有任意定向柔顺轴的在线可变阻抗模块化框架

Mihael Simonič, Xiaocong Li

机构 * College of Information Science and Technology, Eastern Institute of Technology(信息科学与技术学院,东方理工学院) Zhejiang Key Laboratory of Industrial Intelligence and Digital Twin, Eastern Institute of Technology(浙江省工业智能与数字孪生重点实验室,东方理工学院)

AI总结 该论文提出与机器人无关的柔顺控制框架,通过插件架构分离控制器与控制律,标准化接口支持多种控制公式。其参考笛卡尔阻抗控制器可在线更新柔顺方向,实验和仿真验证了在不同操纵器上的任务相关柔顺性及可移植性。

Comments 8 pages, 7 figures, 2 tables. Paper page: https://smihael.github.io/plug-play-comply/

详情
AI中文摘要

本文提出了一种与机器人无关的柔顺控制框架,通过标准化关节和笛卡尔命令接口扩展了ROS控制生态系统。它解决了现有控制软件的一个关键限制,即缺乏用于在不同操纵器上实现柔顺控制算法的可重用基础设施,同时保持与高级应用程序的通用接口。基于插件的架构将控制器基础设施与控制律实现分离。通用包装器使用现有的硬件抽象与不同的操纵器接口,而运行时加载的插件仅实现控制律。命令接口支持关节和笛卡尔空间参考、刚度和阻尼增益、零空间目标和前馈项,实现可变阻抗和多种柔顺控制公式。机器人运动学和动力学使用Pinocchio从URDF模型计算得出。该架构促进了柔顺控制策略的开发,并使相同的实现能够在不同平台上不变地部署。完整框架包括参考控制器、高级任务接口和各种操纵器的示例配置已开源。参考笛卡尔阻抗控制器通过旋转平移和旋转刚度及阻尼来支持任务相关的柔顺性,允许主要柔顺方向根据局部任务几何形状在线更新,而不是固定在机器人基座或TCP框架中。实际机器人实验证明了在接触丰富的操纵中任务相关的柔顺性,而仿真显示了跨具有不同运动学和动力学特性的操纵器的可移植性。

英文摘要

The paper proposes a robot-agnostic compliant-control framework that extends the ROS control ecosystem with standardized joint and Cartesian command interfaces. It addresses a key limitation of existing control software: no reusable infrastructure for implementing compliant-control algorithms across different manipulators while preserving a common interface to higher-level applications. A plugin-based architecture separates controller infrastructure from control-law implementation. Generic wrappers use existing hardware abstractions to interface with different manipulators, while runtime-loaded plugins implement only the control law. Command interfaces support joint- and Cartesian-space references, stiffness and damping gains, nullspace targets, and feedforward terms, enabling variable impedance and diverse compliant-control formulations. Robot kinematics and dynamics are computed from URDF models using Pinocchio. The architecture facilitates the development of compliant-control strategies and enables the same implementation to be deployed across platforms unchanged. The complete framework, including reference controllers, high-level task interfaces, and example configurations for various manipulators, is open-sourced. The reference Cartesian impedance controller supports task-dependent compliance by rotating translational and rotational stiffness and damping, allowing the principal compliance directions to be updated online according to local task geometry rather than remaining fixed in the robot base or TCP frame. This is particularly important in contact-rich manipulation, where the desired directions of motion, constraints, and compliance directions may vary throughout task execution. Real-robot experiments demonstrate task-dependent compliance in contact-rich manipulation, while simulations show portability across manipulators with distinct kinematic and dynamic characteristics.

URL PDF HTML 收藏
2607.22467 2026-07-27 cs.LG 新提交

Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates

学习投影梯度下降求解器迭代的复杂度界与方法

Anjian Li, Ryne Beeson

机构 * Princeton University(普林斯顿大学)

AI总结 针对数据稀缺挑战,研究用k邻域数据收集策略扩充训练数据,推导泛化界,以投影梯度下降求解单边盒约束二次规划为例说明,提高数据-模型-优化循环效率,实现更强大的DDDAS范式,并与GLENS联系。

详情
AI中文摘要

数据稀缺在训练生成模型以产生参数优化问题的初始猜测时构成了一个基本挑战,这些问题在数值上求解成本高昂。因此,我们研究了一种k邻域数据收集策略,该策略用中间求解器迭代来扩充收敛解的数据集,在不进行额外求解器运行的情况下增加训练数据量。为理解此方法的益处,我们基于拉德马赫复杂度推导了一个泛化界,揭示了k邻域和相关参数的作用。我们专注于通过投影梯度下降求解的单边盒约束二次规划。在两个例子中说明了该求解器的行为。本文提出的方法通过提高数据-模型-优化循环的效率实现了更强大的DDDAS范式。最后讨论了学习求解器迭代数据的两种观点,并将我们的分析与一种新的数据高效全局搜索方法GLENS联系起来。

英文摘要

Data scarcity poses a fundamental challenge in training generative models to produce initial guesses for parametric optimization problems that are otherwise numerically expensive to solve. We therefore study a $k$-neighborhood data collection strategy that augments datasets of converged solutions with intermediate solver iterates, increasing the amount of training data without additional solver runs. To understand the benefits of this approach, we derive a generalization bound based on Rademacher complexity that reveals the role of the $k$-neighborhoods and related parameters. To achieve this result, we focus on one-sided box-constrained quadratic programs solved by projected gradient descent. We illustrate the behavior of this solver on two examples. The approach proposed in this paper enables a more capable DDDAS paradigm by improving the efficiency of the data-model-optimization loop. We finish by discussing two views of learning solver-iterate data and connect our analysis with GLENS, a new data-efficient global search method.

URL PDF HTML 收藏
2607.22465 2026-07-27 cs.AI cs.LG cs.MA 新提交

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

TRACE-ROUTER:用于智能AI的任务一致且自适应的在线路由

Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna

机构 * Georgia Institute of Technology(佐治亚理工学院) Intel(英特尔公司) Texas A&M University(德克萨斯农工大学)

AI总结 研究针对企业AI中现有路由决策与智能应用任务级结果不匹配问题,提出TRACE-Router任务级路由框架,利用上下文博弈、终端奖励更新策略,在多基准测试中改善准确性-延迟权衡,取得较好效果。

详情
AI中文摘要

选择具有不同成本-质量权衡的大语言模型进行路由已成为企业AI的基本部署特征。现有路由器主要为每个大语言模型调用独立做出路由决策。然而,智能应用作为长期工作流执行,其质量仅由延迟的任务级结果决定。这种不匹配使逐调用路由器无法将反馈正确归因于单个路由决策。为缓解此问题,我们提出TRACE-Router,这是一个任务级路由框架,使路由与监督单元对齐。TRACE-Router在接纳时使用上下文博弈为每个任务分配一个模型,将所有后续大语言模型调用固定到所选后端,并使用任务的终端奖励更新其策略,综合考虑准确性和延迟。通过利用延迟任务反馈,TRACE-Router学习适应工作负载的路由策略,同时避免显式任务复杂性估计。在三个智能基准测试中,TRACE-Router持续改善准确性-延迟权衡,实现非支配帕累托前沿点。在tau2-Bench上,它比单个模型之间的延迟匹配插值性能高7-8个准确性点,在Terminal-Bench上,它比最强的单模型基线准确性高7.1个点,延迟低36%。

英文摘要

Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-horizon workflows whose quality is determined only by a delayed, task-level outcome. This mismatch prevents per-call routers from correctly attributing feedback to individual routing decisions. Towards mitigating this, we present TRACE-Router, a task-level routing framework that aligns routing with the unit of supervision. TRACE-Router assigns each task to a model once at admission using a contextual bandit, pins all subsequent LLM calls to the selected backend, and updates its policy using the task's terminal reward, jointly accounting for accuracy and latency. By leveraging delayed task feedback, TRACE-Router learns routing policies that adapt to the workload while avoiding explicit task-complexity estimation. Across three agentic benchmarks, TRACE-Router consistently improves the accuracy-latency trade-off, achieving non-dominated Pareto frontier points. On tau2-Bench, it outperforms latency-matched interpolation between individual models by 7-8 accuracy points, while on Terminal-Bench it achieves 7.1 higher accuracy points than the strongest single model baseline with 36% lower latency.

URL PDF HTML 收藏
2607.22458 2026-07-27 cs.LG cs.AI 新提交

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

音频基础模型捕捉到的海洋哺乳动物和鸟类发声中的系统发育信号:特定领域预训练的有限益处

Víctor Rincón Yepes

机构 * Earth Species Project(地球物种计划)

AI总结 研究利用四个预训练音频模型从物种发声恢复系统发育距离,在海洋哺乳动物和鸟类中,通用基础模型能恢复信号,特定领域预训练的BirdNET等未超越,表明预训练音频嵌入可携带进化信息,特定领域预训练非必需。

Comments 17 pages, 5 figures, 2 supplementary tables. Code, embeddings and derived matrices: https://github.com/rinvictor/bioacoustic-phylogeny-embeddings

详情
AI中文摘要

我们探究了四个大型预训练音频模型(AST、CLAP、BEATs - bio和BirdNET),利用一个它们在训练期间都未见过的下游任务:从物种发声中恢复系统发育距离。在32种海洋哺乳动物(来自沃特金斯海洋哺乳动物声音数据库的1754条录音)中,基础模型在26种鲸类动物中恢复了强烈的系统发育信号(CLAP r = 0.82,BEATs - bio r = 0.82,AST r = 0.74;所有p < 0.001),而手工制作的MFCC特征(105维)则未发现信号(r = 0.040,p = 0.338)。在20种鸟类中重复分析,通用基础模型再次恢复了信号(AST r = 0.55,CLAP r = 0.52),而BirdNET和BEATs - bio并未超越它们(r约为0.32至0.36)。预训练音频嵌入携带跨两个独立辐射的进化信息,特定领域预训练并非其出现的必要条件。

英文摘要

Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from species vocalizations. If the geometry of the embedding space tracks the tree of life, the representation is picking up something deeper than the labels the model was optimized for. We run Mantel tests across two independent radiations. In 32 marine mammal species (1,754 recordings from the Watkins Marine Mammal Sound Database) the foundation models recover strong phylogenetic signal within the 26 cetaceans (CLAP r=0.82, BEATs-bio r=0.82, AST r=0.74; all p<0.001), among the highest acoustic-phylogenetic correlations reported for any taxon. Hand-crafted MFCC features (105d) find nothing (r=0.040, p=0.338). The gap survives after PCA-projecting every embedding down to 105 dimensions, so it is not an artefact of representation size. It also survives a partial Mantel test controlling for dominant frequency (partial Mantel r=0.404, keeping 97% of the variance explained), so it is not just pitch in disguise. We repeat the analysis on 20 bird species using the Jetz et al. (2012) phylogeny, and this time add BirdNET, a classifier trained end-to-end on around 6,000 bird species. The general-purpose foundation models recover the signal again (AST r=0.55, CLAP r=0.52). The unexpected result is that neither BirdNET nor the bioacoustic BEATs-bio beat them (r around 0.32 to 0.36). Matching the training domain to the target taxon does not, by itself, help. Pretrained audio embeddings carry evolutionary information across two independent radiations, and domain-specific pretraining is not required for it to emerge.

URL PDF HTML 收藏
2607.22446 2026-07-27 cs.CV cs.GR 新提交

Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering

可变形三角形散列:用于实时辐射场渲染的灵活基元

Oriol Jiménez-Ayguadé, Antonio Agudo

机构 * Institut de Robòtica i Informàtica Industrial, CSIC-UPC(工业机器人与信息学研究所,西班牙国家研究委员会-加泰罗尼亚理工大学)

AI总结 研究针对辐射场方法依赖凸边界问题,提出可变形三角形散列,通过控制点和位移实现非凸形状表示,设计光栅化管道可微渲染,经多种场景验证,在视觉质量、通用性和渲染效率上表现出色。

Comments Accepted at ECCV 2026. Project page: https://orioljim1.github.io/detris

详情
AI中文摘要

近期的辐射场方法使用二维基元(从高斯圆盘到三角形)来表示场景,这些基元能实现表面对齐和高效光栅化,但都依赖凸边界,处理弯曲和凹形结构需要过多基元。我们引入可变形三角形散列,每条边用K个控制点增强每个三角形,通过单个可学习标量位移参数化,使边界向内或向外移动,在保留定义3D平面的三个基本顶点的同时实现非凸形状表示。为可微渲染这些非凸基元,我们在三角形重心坐标空间设计光栅化管道,确保视图一致渲染。通过缠绕数测试确定像素是否在变形基元内,由锐度和角平滑度两个可学习参数控制的窗口函数以及每个基元的标量不透明度,产生从内部到边界的平滑不透明度过渡。在各种真实场景中进行验证,在视觉质量和通用性方面优于基于非体积基元的近期工作,同时仍实现有竞争力的渲染效率。

英文摘要

Recent radiance field methods represent scenes with 2D primitives that offer surface alignment and efficient rasterization, from Gaussian disks to triangles, yet all rely on convex boundaries: curved and concave structures demand excessive primitives. We introduce Deformable Triangle Splatting, which augments each triangle with $K$ control points per edge, each parameterized by a single learnable scalar displacement that shifts the boundary inward or outward, enabling non-convex shape representation while preserving the three base vertices that define the 3D plane. To render these non-convex primitives differentiably, we design a rasterization pipeline in the triangle's barycentric coordinate space, ensuring view-consistent rendering. A winding number test determines whether each pixel lies inside the deformed primitive, and a window function controlled by two learnable parameters, sharpness and corner smoothness, together with a per-primitive scalar opacity, produces the smooth opacity transition from interior to boundary. Validation is done in a variety of real-world scenes, outperforming recent works based on non-volumetric primitives in terms of visual quality and versatility while still achieving competitive rendering efficiency.

URL PDF HTML 收藏
2607.22444 2026-07-27 cs.LG cs.AI 新提交

Hyperball May Not Be a Free Lunch

超球可能并非免费午餐

Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai

机构 * IQuest Research(IQuest研究公司) Peking University(北京大学) Sun Yat-sen University(中山大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Shanghai University of Finance and Economics(上海财经大学)

AI总结 研究超球风格优化器优势来源,通过推导角有效学习率等方法分析,发现其优势并非源于更新方向,而是有效步长演变,预训练实验表明学习率衰减策略影响性能,谨慎调度对发挥其潜力至关重要。

Comments 14 pages, 4 figures. Code: https://github.com/mangocrazz/hyperball-may-not-be-a-free-lunch. Equal contribution: Yihao Xiao and Jialong Sun. Corresponding author: Bryan Dai

详情
AI中文摘要

对于尺度不变的深度网络,超球风格的优化器通过固定矩阵值参数的范数和归一化更新在大规模训练中表现出强大性能。但其优势来源不明。本文从连续参数状态间的角位移出发,推导了角有效学习率,表明传统基于范数的度量是参数更新正交性下的特殊情况。分解优化器更新为径向和切向分量,分析径向更新对角位移的影响。数值结果显示径向分量对角有效学习率的直接影响有限,无法解释MuonH在训练初期比MuonWD收敛慢但后期超过它的原因。通过启发式实验发现它们的主要差异源于有效步长的演变而非超球诱导的本质上更优的更新方向。预训练实验表明更激进的学习率衰减可在训练初期加速MuonH但可能损害后期性能。因此,保持恒定角速度并不能消除学习率调度问题,谨慎调度对实现超球风格优化器的潜力至关重要。

英文摘要

For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Starting from the angular displacement between consecutive parameter states, we derive an angular effective learning rate that accounts for the parameter-update angle, parameter norm, and update norm. We also show that the conventional norm-based measure is a special case under parameter-update orthogonality. We then decompose optimizer updates into radial and tangential components and analyze how radial updates affect one-step angular displacement. Under the training configurations considered, numerical results show that the radial component has only a limited direct effect on the angular effective learning rate. It therefore cannot explain why MuonH converges more slowly than MuonWD early in training but overtakes it later. To further isolate the underlying mechanism, we devise a heuristic experiment that modifies only the learning-rate schedule so that the dynamics of each optimizer reproduce those of the other. The results suggest that their main difference stems from the evolution of the effective step size rather than an intrinsically superior update direction induced by Hyperball. Our pretraining experiments further show that more aggressive learning-rate decay can accelerate MuonH early in training but may impair its later performance. Thus, maintaining a constant angular velocity does not eliminate the learning-rate-scheduling problem; careful scheduling remains essential to realizing the potential of Hyperball-style optimizers. Our code is publicly available at https://github.com/mangocrazz/hyperball-may-not-be-a-free-lunch.

URL PDF HTML 收藏
2607.22434 2026-07-27 cs.RO cs.AI 新提交

Robot Learning to Communicate through Projected Visual Abstractions

机器人通过投影视觉抽象进行通信学习

Danyang Yan, Boyuan Wang, Jiaxun Liu, Boyuan Chen

机构 * Duke University(杜克大学)

AI总结 研究让机器人通过投影视觉抽象通信,提出含21自由度灵巧手和影子自我模型的系统,能动态表达影子。通过自我探索学习映射,依目标优化配置,经模拟获可行运动,还引入多种优化,在多场景展示影子表达,建立相关框架。

Comments Our project website is at:https://generalroboticslab.com/shadow

详情
AI中文摘要

人类常通过身体的抽象形式进行交流,如影子、轮廓和反射。但机器人大多局限于通过物理形态表达。让机器人通过投影视觉抽象进行通信,不仅需考虑身体运动,还需考虑运动如何转化为观察者感知的外部表征。以影子为例,我们提出一个机器人系统,它使用有柔顺软皮肤的21自由度灵巧手和学习到的影子自我模型进行动态影子表达。软皮肤减少光泄漏以产生视觉上连续的轮廓,可微的自我模型通过任务无关的自我探索学习手部配置与投影影子外观之间的映射。给定目标影子图像或视频,机器人通过基于梯度的搜索优化手部配置,并通过碰撞感知模拟优化解决方案。为实现动态影子表现,还引入了表达区域目标、时间平滑正则化和基于关键帧的优化。我们在模拟和物理实验中展示了机器人在手语手势、手影戏和动物运动模仿中的影子表达。这些结果建立了一个框架,使机器人能够操纵自身的投影视觉抽象进行通信和视觉叙事。

英文摘要

Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visually distinct from the body itself. Here, we present a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration. Given a target shadow image or video, the robot optimizes its hand configurations through gradient-based search over 1 the learned self-model and refines the solution through collision-aware simulation to obtain physically feasible motions. For dynamic shadow performance, we further introduce expressive-region objectives, temporal smoothness regularization, and keyframe-based optimization to preserve visually important motion cues while reducing optimization complexity. We demonstrate robotic shadow expression across sign-language gestures, hand-shadow puppetry, and animal motion imitation in both simulation and physical experiments. These results establish a framework for enabling robots to manipulate projected visual abstractions of themselves for communication and visual storytelling.

URL PDF HTML 收藏
2607.22409 2026-07-27 cs.RO cs.SY eess.SY 新提交

Conformal Constraint Tightening for Chance-Constrained Motion Planning with Unknown Dynamics

具有未知动力学的机会约束运动规划的共形约束收紧

Shubham Natraj, Bruno Sinopoli, Yiannis Kantaros

机构 * Washington University in St. Louis(圣路易斯华盛顿大学) Arizona State University(亚利桑那州立大学)

AI总结 针对未知动力学系统,利用共形预测给出轨迹偏差概率界,收紧规划约束,为现有规划器提供真实系统概率任务完成保证,实验验证理论保证且任务完成率显著提升。

详情
AI中文摘要

运动规划算法计算控制序列,驱使自主机器人到达目标区域并避开不安全状态。现有方法通常仅针对标称模型或模拟器提供任务完成保证,当真实动力学未知或难以准确建模时可能失效。本文针对具有未知动力学且有可用近似标称模型的系统解决此限制,提出一种与规划器无关的约束收紧程序,为现有规划器提供关于真实系统的概率任务完成保证。利用共形预测给出标称到真实轨迹偏差的概率界,用该界收紧规划约束,并表明在标称模型下解决收紧问题是在真实系统上以规定概率解决原始问题的充分条件。通过实验验证了理论保证,并展示了相对于标称模型规划显著提高的任务完成率。

英文摘要

Motion planning algorithms compute control sequences that drive autonomous robots to goal regions while avoiding unsafe states. Existing methods, from sampling-based planning to deep reinforcement learning, typically provide task-completion guarantees only with respect to a nominal model or simulator, which may be invalidated when the true dynamics are unknown or difficult to model accurately. This letter addresses this limitation for systems with unknown dynamics and an available approximate nominal model, contributing a planner-agnostic constraint-tightening procedure that equips existing planners with a probabilistic task-completion guarantee on the true system. We leverage conformal prediction to provide a probabilistic bound on the nominal-to-true trajectory deviation over a distribution of planning problems. We tighten the planning constraints using that bound, and show that solving the tightened problem under the nominal model is a sufficient condition for solving the original problem on the true system with a prescribed probability. We validate the theoretical guarantees empirically and demonstrate substantially improved task completion relative to nominal-model planning.

URL PDF HTML 收藏
2607.22408 2026-07-27 cs.LG 新提交

LunarFM: A Shared Multimodal Representation of the Moon's Surface

LunarFM:月球表面的共享多模态表示

Marc Girona-Mata, Jakob Gawlikowski, Sumit Goski, Gautier Bardi de Fourtou, Valentin T. Bickel, Ben Moseley, Abigail Calzada-Diaz, Sylvester Kaczmarek, Raúl Ramos-Pollán

机构 * University of Cambridge(剑桥大学) German Aerospace Center (DLR)(德国航空航天中心) SPAIDER SPACE(斯派德太空公司) Mines Paris - PSL University(巴黎矿业学院 - 巴黎文理研究大学) University of Bern(伯尔尼大学) Imperial College London(伦敦帝国理工学院) European Space Resources Innovation Center (ESRIC)(欧洲空间资源创新中心) Universidad de Antioquia(安蒂奥基亚大学)

AI总结 因月球探索需求,针对多源观测致月球表面分析碎片化问题,提出LunarFM多模态基础模型,融合多仪器观测学习通用表示,支持多种下游应用,还提供相关数据集、预训练模型等,助力月球表面高效分析。

Comments 19 pages, 12 figures

详情
AI中文摘要

全球对月球探索的重新关注,因原位资源利用和人类在月球持续存在的前景,对月球表面精确大规模表征需求日增。虽已收集大量轨道遥感数据,但科学分析和资源测绘因多仪器观测异质性、稀疏标签及特定任务建模工作流程而碎片化。本文引入LunarFM,一种多模态基础模型,从多样轨道测量中学习月球表面通用表示。它融合三次月球任务中六种仪器的观测,将18个输入通道映射到共享嵌入空间。实验表明该嵌入空间支持多种下游应用,包括相似性搜索、少样本资源测绘、矿物丰度回归和地质单元分类等,实现高效科学调查和资源导向分析。还提供了机器学习可用数据集、预训练多模态掩码自动编码器及配套嵌入数据集。所有代码和数据可在指定网址获取。

英文摘要

The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface. Although vast quantities of orbital remote-sensing data have been collected, scientific analysis and resource mapping remain fragmented by heterogeneous multiinstrument observations, sparse labels, and bespoke task-specific modelling workflows. Here we introduce LunarFM, a multimodal foundation model that learns a general representation of the lunar surface from diverse orbital measurements. LunarFM assimilates observations from six instruments across three lunar missions, mapping 18 input channels to a shared embedding space. We demonstrate that this embedding space supports a diverse range of downstream applications, including similarity search, few-shot resource mapping, mineral abundance regression, and geological unit classification, enabling efficient scientific investigation and resource-oriented analysis. We provide a machine-learning-ready dataset of co-registered multimodal observations spanning latitudes from 70°S to 70°N, a pretrained multimodal masked autoencoder, and a companion embedding dataset providing a joint 768-dimensional representation of lunar surface properties. All code and data are available at https://lunarfm.trillium.tech/

URL PDF HTML 收藏
2607.22393 2026-07-27 cs.AI cs.CV 新提交

SceneActBench: Can Agents Act on the 3D Scenes They See?

SceneActBench:智能体能否对其所看到的3D场景采取行动?

Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu

机构 * Tencent Hunyuan(腾讯混元) THU(清华大学) NJU(南京大学) HKUST(香港科技大学) UIUC(伊利诺伊大学厄巴纳-香槟分校) PKU(北京大学)

AI总结 研究视觉语言模型智能体在3D场景行动能力,提出SceneActBench基准测试,涵盖五个3D任务,通过特定指标评估智能体输出,分析不同配置得分及失败情况,为评估智能体在多对象3D场景行动提供依据。

详情
AI中文摘要

视觉语言模型(VLM)智能体越来越多地使用工具对3D场景采取行动,而不仅仅是描述它们。现有的3D基准测试对文本响应或单对象操作进行评分,未评估智能体在完整多对象3D场景上的行动。我们提出了SceneActBench,这是一个在统一智能体 - 环境循环下针对五个3D任务的视觉条件行动基准测试。给定PNG图像或采样视频帧以及适用时提供的3D资产,智能体在3D环境中行动。我们使用特定任务的几何指标根据隐藏的地面真值评估每个最终输出。SceneActBench由210个源实例构建的五个任务组成,产生520个任务案例,包括配对输入条件。每个任务通过一个固定的智能体循环运行以保持比较公平。在十一种专有VLM配置中,总体得分在38.6 - 50.2之间,且没有一个在所有任务中都表现良好。我们进一步分析了失败出现的位置和方式。

英文摘要

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38.6-50.2, and none performs consistently well across tasks. We further analyse where and how failures manifest.

URL PDF HTML 收藏
2607.22386 2026-07-27 cs.CV cs.CR 新提交

Correlation-Aware and Gaussianity-Preserving Robust Latent Angular Watermarking for Diffusion Models

用于扩散模型的相关感知和高斯性保持的鲁棒潜在角水印

Yebin Zheng, Haonan An, Guang Hua, Zhiping Lin, Yuguang Fang

机构 * Singapore Institute of Technology(新加坡理工学院) City University of Hong Kong(香港城市大学) Nanyang Technological University(南洋理工大学)

AI总结 研究扩散模型潜在域水印问题,提出潜在角水印(LAW)及变体LAW-M,通过对映角编码保持高斯性,解决现有方法易受攻击及相关性退化问题,理论上严格表征相关性退化并推导自相关结构。

详情
AI中文摘要

扩散模型的潜在域水印将水印直接嵌入到潜在先验中,对模型参数无侵入性且能与生成过程无缝集成。但现有方法因违反潜在高斯性或对潜在反演中的正常和恶意扰动敏感,易受水印检测或去除攻击,还存在违反独立同分布潜在条件导致潜在相关性退化和生成保真度损失的问题,虽已通过FID外部测量,但内部相关结构尚未严格表征。为解决这些问题,受各向同性高斯旋转不变性启发,我们提出潜在角水印(LAW),它在保持高斯性的同时将水印位编码为潜在元素不相交对之间的对映角(相对于参考对为±π/2)。对映编码最大化了位值之间的几何分离,我们证明解码角误差方差与潜在对的范数成比例,即var(Δϕ) ∝ 1/ρ²。我们还提出了幅度驱动变体LAW-M,它将水印位锚定在几何上最稳定的潜在维度中,进一步提高了鲁棒性。理论上,我们对诱导的相关性退化进行了严格表征,以封闭形式推导了水印潜在的自相关结构,并证明相关性局限于具有固定±π/4值的稀疏、结构化非对角元素集。

英文摘要

Latent domain watermarking for diffusion models embeds watermarks directly into the latent prior, enjoying non-intrusiveness to model parameters and seamless integration with the generation process. However, due to the violation of latent Gaussianity or sensitivity to normal and malicious perturbations during latent inversion, existing methods are prone to watermark detection or removal attacks. A further overlooked problem is the violation of the i.i.d. latent condition after watermarking, which leads to latent correlation degradation and generation fidelity loss. Although this has been externally measured by FID, the internal correlation structure has yet to be rigorously characterized. To address the above issues, and motivated by the rotation-invariant property of isotropic Gaussian, we propose \textit{Latent Angular Watermarking (LAW)}, which encodes watermark bits as antipodal angles ($\pmπ/2$ relative to a reference pair) between disjoint pairs of latent elements while preserving the Gaussianity. The antipodal ($π$-separation) encoding maximizes geometric separation between bit values, and we prove that the decoding angular-error variance is proportional to the norm of the latent pair, i.e., $\operatorname{var}(Δϕ) \propto 1/ρ^2$. We further propose a magnitude-driven variant, LAW-M, which anchors watermark bits in the most geometrically stable latent dimensions, yielding additional robustness gains. Theoretically, we provide a rigorous characterization of the induced correlation degradation, deriving in closed form the autocorrelation structure of the watermarked latent and proving that correlations are confined to a sparse, structured set of off-diagonal elements with fixed $\pmπ/4$ values.

URL PDF HTML 收藏
2607.22385 2026-07-27 cs.AI cs.LG 新提交

Agentic Root Cause Analysis through Evidence-Grounded Reasoning

通过基于证据的推理进行智能根本原因分析

Amaury Wei, Olga Fink

机构 * EPFL - IMOS Laboratory(洛桑联邦理工学院 - IMOS实验室)

AI总结 研究针对工业异常根本原因诊断依赖人工且现有数据驱动方法有局限的问题,提出AgentRCA框架,结合数字孪生和大语言模型进行推理,在实际设施上评估,性能与监督基线相当且能产生透明推理轨迹,为工业根本原因分析提供实用基础。

Comments 21 pages, 9 figures

详情
AI中文摘要

诊断异常的根本原因对工业安全运行至关重要。尽管有大量传感器,但制定假设和收集证据仍是手动过程,成为操作瓶颈。现有数据驱动方法有两个关键局限:像黑箱无法解释诊断,且需要大量故障操作的标记示例。为解决此差距,我们引入AgentRCA,一个用于基于证据的根本原因分析的零样本智能框架。它通过结合数据驱动的数字孪生(模拟正常系统动态)和工具增强的大语言模型进行推理时推理。该智能体迭代收集统计证据、评估竞争假设并识别最能解释观察到行为的物理故障。在实际多相流设施和大型化工厂上评估,AgentRCA在不依赖特定故障训练的情况下,实现了与完全监督基线相当的诊断性能。关键是,它产生透明的推理轨迹,明确将观察到的症状与其潜在物理原因联系起来。这些结果将自主假设驱动推理确立为可扩展工业根本原因分析的实用基础。

英文摘要

Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existing data-driven approaches aim to automate this, two critical limitations restrict their deployment: their operate as black boxes unable to justify their diagnosis, and they require scarce labeled examples of faulty operation. To address this gap, we introduce AgentRCA, a zero-shot agentic framework for evidence-grounded root cause analysis. Rather than learning fault-specific mappings, AgentRCA performs inference-time reasoning by combining a data-driven digital twin (modeling normal system dynamics) with a tool-augmented large language model. The agent iteratively gathers statistical evidence, evaluates competing hypotheses, and identifies the physical fault that best explains the observed behavior. Evaluated on a real-world multiphase-flow facility and a large-scale chemical plant, AgentRCA achieves diagnostic performance competitive with fully supervised baselines without relying on fault-specific training. Crucially, it produces transparent reasoning traces that explicitly link observed symptoms to their underlying physical causes. These results establish autonomous hypothesis-driven reasoning as a practical foundation for scalable industrial root cause analysis.

URL PDF HTML 收藏
2607.22381 2026-07-27 cs.LG cs.SI stat.ML 新提交

Local-Global Geometric Insights for Graph Neural Networks via Entropic Curvature

通过熵曲率对图神经网络的局部-全局几何洞察

Rachid Caich, Yassine Abbahaddou

机构 * LIX, École Polytechnique IP Paris(巴黎综合理工学院LIX)

AI总结 研究通过引入熵曲率解决图神经网络中信息长距离传播问题,定义弱熵曲率代理并推导相关不等式和界,转化为实用机制,经基准测试验证了该方法在解决过平滑和过挤压等问题上的有效性。

详情
AI中文摘要

图上的曲率概念,特别是奥利维耶-里奇曲率和福尔曼曲率,已成为解决图神经网络(GNNs)中诸如过平滑和过挤压等基本问题的有力工具,但几乎完全依赖于局部边级比较,因此无法证明信息如何在长距离上实际传播。我们引入了熵曲率,这是一种基于传输的全局曲率,通过沿瓦瑟斯坦测地线的熵的位移凸性将洛特-斯特姆-维拉尼框架扩展到图而获得。我们定义了一个易于处理的弱熵曲率代理,它为全局熵曲率提供下界,并由此推导出(i)一个控制过平滑的庞加莱型不等式,(ii)一个传输-熵泛化界以及(iii)一个扩展悖论证明在大图中稀疏性、强谱扩展和正熵曲率不能共存,将过平滑和过挤压统一为单个曲率谱的相反两端。我们将该理论转化为三种实用机制,即E-Gate聚合器、ENT结构编码和中点完成重新布线(MCR),并在六个节点分类基准和图分类上针对SDRF、FoSR、BORF、LCP和图里奇流对它们进行基准测试。

英文摘要

Curvature notions on graphs, particularly Ollivier-Ricci and Forman, have emerged as powerful tools for addressing fundamental issues in Graph Neural Networks (GNNs) such as oversmoothing and oversquashing, but rely almost exclusively on local edge-level comparisons and therefore fail to certify how information actually propagates over long distances. We introduce Entropic Curvature, a global, transport-based curvature obtained by extending the Lott-Sturm-Villani framework to graphs through the displacement convexity of entropy along Wasserstein geodesics. We define a tractable Weak Entropic Curvature proxy that lower-bounds the global entropic curvature, and from it derive (i) a Poincare-type inequality controlling oversmoothing, (ii) a transport-entropy generalization bound, and (iii) an expansion paradox proving that sparsity, strong spectral expansion, and positive entropic curvature cannot coexist in large graphs, unifying oversmoothing and oversquashing as opposite ends of a single curvature spectrum. We translate the theory into three practical mechanisms, the E-Gate aggregator, the ENT structural encoding, and Midpoint-Completion Rewiring (MCR), and benchmark them against SDRF, FoSR, BORF, LCP, and Graph Ricci Flow on six node-classification benchmarks, and graph-classification.

URL PDF HTML 收藏
2607.22380 2026-07-27 cs.CV 新提交

IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing

IR275K:面向高效遥感的红外多帧超分辨率基准测试

Jie Deng, Heyang Wang, Changxin Wang, Junkai Shen, Hongyi Chen, Zhiping He, Hongxing Qi, Xudong Zhang, Jianyu Wang

机构 * Hangzhou Institute for Advanced Study(杭州高等研究院) Shanghai Institute of Technical Physics of the Chinese Academy of Sciences(中国科学院上海技术物理研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 该研究针对红外遥感中多帧超分辨率评估分散问题,引入IR275K基准测试。以CGMamba模型为例进行架构分析,其结合2D~RoPE与CGCM融合,实现高效重建,为红外MFSR方法评估及模型设计提供基础和起点。

详情
AI中文摘要

在红外遥感中,高效处理愈发重要,卫星星座在受限探测器分辨率、功率和下行链路带宽下产生大量观测数据。多帧超分辨率(MFSR)提供了基于软件的空间增强途径,但在红外传感中的评估仍分散于私有数据集和临时协议。现有基准未明确捕捉红外视频的热对比度、传感器噪声、弱纹理和平台引起的帧间变化。我们引入IR275K,这是一个包含594个红外视频序列和275196帧的精心策划的基准测试。它提供序列级的训练/验证/测试划分以及可重现的X4评估协议。作为初步的架构探索,我们进一步评估了CGMamba,一个具有1090万个参数和112.14 GFLOP的轻量级状态空间模型。CGMamba将二维旋转位置编码(2D~RoPE)与中心引导的交叉Mamba(CGCM)融合用于隐式多帧重建。它实现了33.19dB的PSNR,在显著更低的计算成本下比红外单图像超分辨率参考高出0.35 - 0.52dB。消融结果表明从CGCM中移除2D~RoPE会导致1.53dB的下降和严重的网格状伪影。这表明在红外条件下,显式空间锚定对于稳定基于状态空间模型的跨帧门控至关重要。IR275K为红外MFSR方法的准确性 - 效率评估提供了可重现的基础,而架构分析为资源受限的红外传感下的空间感知状态空间模型设计提供了具体起点。数据集和评估资源可在指定网址获取。

英文摘要

Efficient processing is becoming increasingly important in infrared remote sensing, where satellite constellations produce large volumes of observations under constrained detector resolution, power, and downlink bandwidth. Multi-frame super-resolution (MFSR) offers a software-based route to spatial enhancement, but its evaluation in infrared sensing remains fragmented across private datasets and ad-hoc protocols. Existing benchmarks do not explicitly capture the thermal contrast, sensor noise, weak texture, and platform-induced frame-to-frame variation that characterize infrared video. We introduce IR275K, a curated benchmark containing 594 infrared video sequences and 275,196 frames. It provides sequence-level train/validation/test splits and a reproducible X4 evaluation protocol. As an initial architectural probe, we further evaluate CGMamba, a lightweight state-space model with 10.90M parameters and 112.14G FLOPs. CGMamba combines 2D rotary position encoding (2D~RoPE) with center-guided cross-Mamba (CGCM) fusion for implicit multi-frame reconstruction. It achieves 33.19dB PSNR, outperforming infrared single-image super-resolution references by 0.35--0.52~dB at substantially lower computational cost. Ablation results show that removing 2D~RoPE from CGCM causes a 1.53dB drop and severe grid-like artifacts. This indicates that explicit spatial anchoring is critical for stabilizing SSM-based cross-frame gating under infrared conditions. IR275K provides a reproducible foundation for accuracy--efficiency evaluation of infrared MFSR methods, while the architectural analysis offers a concrete starting point for spatially aware SSM design under resource-constrained infrared sensing. Dataset and evaluation resources are available at: https://github.com/InfraRecon7/IR275K.

URL PDF HTML 收藏
2607.22376 2026-07-27 cs.CL 新提交

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

使用语法书进行低资源机器翻译合成数据生成的析因研究

Varun Ghat Ravikumar, Sina Ahmadi, Lena Jäger, Rico Sennrich

机构 * University of Zurich(苏黎世大学)

AI总结 研究针对濒危语言机器翻译缺平行数据问题,利用大语言模型从语法书提取内容生成合成语料库微调,经三种低资源语言验证,通过析因研究确定增益因素组合,证明可将静态语言文档用于机器翻译微调,为资源匮乏语言提供翻译工具路径。

Comments Accepted at CLiC-it 2026

详情
AI中文摘要

尽管存在描述性语法书,但大多数濒危语言缺乏机器翻译所需的平行数据。我们引入了一种管道,利用大语言模型从语法书中提取语法规则、例句和词汇表,并生成合成平行语料库用于微调,而不是像先前工作那样在推理时将语法内容输入提示中。在三种类型不同的低资源语言——卡拉芒语(巴布亚语系)、图阿钦语(罗曼语族)和曼丹语(苏语族)上进行验证,结果表明,在75%的卡拉芒语配置和59%的图阿钦语配置中,基于合成数据的微调优于种子数据基线,最佳情况下ChrF++增益分别为+8.8、+5.3和+3.3。通过对96种配置进行系统的析因研究,我们确定了哪些因素组合能带来增益以及它们在何处失效。我们的结果表明,静态语言文档可重新用于机器翻译微调,为资源严重不足的语言提供了实用的翻译工具路径。

英文摘要

Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work. Validated on three typologically diverse low-resource languages-Kalamang (Papuan), Tuatschin (Romance), and Mandan (Siouan)-we show that fine-tuning on synthetic data improves over seed-data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, with best-case ChrF++ gains of +8.8, +5.3, and +3.3 respectively. Through a systematic factorial study across 96 configurations varying target part-of-speech, retrieval granularity, and sample volume, we identify which factor combinations drive gains and where they break down. Our results demonstrate that static linguistic documentation can be repurposed for machine translation fine-tuning, offering a practical path towards translation tools for severely under-resourced languages.

URL PDF HTML 收藏
2607.22375 2026-07-27 cs.AI 新提交

IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation

IDEAgent:用于研究想法生成的智能体质量多样性搜索

Varun Gumma, Navonil Majumder, Soumitra Sinhahajari, Soujanya Poria

机构 * DeCLaRe Lab, Nanyang Technological University(声明实验室,南洋理工大学)

AI总结 研究针对大语言模型在生成研究想法时质量与多样性独立导致的局限,提出将研究构思视为质量多样性搜索。介绍多智能体框架IDEAgent,通过多目标反馈驱动质量,轻量级记忆等实现多样性,开发Yield指标评估,实验表明其性能优于基线,还证实修复改进对质量提升的重要性并开源。

Comments Under Review

详情
AI中文摘要

在过去几年中,大语言模型显著地使科学发现过程自动化。然而,现有系统存在一个核心局限:它们在质量或多样性方面独立地生成和优化想法,这常导致生成的想法彼此接近,或产生大量琐碎、不合理或不清晰的概念。在这项工作中,我们认为研究构思应被视为两个目标的结合,并构建为质量多样性(QD)搜索。为此,我们引入了IDEAgent,这是一个通过谱系管理想法演化的多智能体框架。我们使用多目标反馈联合驱动质量以进行专门的修复和改进,而多样性则通过轻量级顺序记忆以及与已完成想法、其历史祖先和被拒绝的提议进行显式比较来实现。为了系统地评估这种QD结合,我们开发了Yield,这是一个联合指标,用于计算满足预定质量阈值的最大一组相互不同的想法。最后,通过对计算机科学8个领域的32个主题的评估,我们表明IDEAgent在Yield上比最佳基线性能高出3.89倍,同时在更多主题上实现了非零Yield。我们还通过质量改进分析进一步证实了这些发现,表明修复和改进对于建立逻辑严谨性和清晰度同时保持非显而易见性至关重要。为鼓励未来基于QD搜索的构思研究,我们在这个https URL上开源了IDEAgent。

英文摘要

Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation should be treated as a conjunction of both objectives and framed as a Quality-Diversity (QD) search. In line with this perspective, we introduce IDEAgent, a multi-agent framework that manages the evolution of ideas through lineages. We jointly drive Quality using multi-objective feedback for dedicated repair and refinement, while Diversity is achieved through lightweight sequential memory and explicit comparison against completed ideas, their historical ancestors, and rejected proposals. To systematically evaluate this QD conjunction, we develop Yield, a joint metric that computes the largest set of mutually diverse ideas that satisfy a predetermined quality threshold. Finally, through evaluations across 32 topics spanning 8 domains of Computer Science, we show that IDEAgent outperforms the best baseline by 3.89x on Yield, while achieving non-zero Yield on 8x more topics. We further corroborate these findings through an analysis of quality improvements, showing that repair and refinement are crucial for building logical rigor and clarity while preserving non-obviousness. To encourage future research on QD-search-based ideation, we open-source IDEAgent at https://github.com/declare-lab/IDEAgent.

URL PDF HTML 收藏
2607.22371 2026-07-27 cs.CV 新提交

Active few-shot segmentation by reinforcing data selection

通过强化数据选择实现主动少样本分割

Chenlan Zhao, Benny Wong, Timothy F. Lundberg, Ahmed M. Elsayed, Abdallah Aljarkas, Hamad A. Aljamaan, Lynn Karam, Qianye Yang, Yipeng Hu, Claire C. Villette, Shaheer U. Saeed

机构 * Centre for Bioengineering, School of Engineering and Materials Science, Queen Mary University of London(伦敦玛丽女王大学工程与材料科学学院生物工程中心) Digital Environment Research Institute, Queen Mary University of London(伦敦玛丽女王大学数字环境研究所) UCL Hawkes Institute(伦敦大学学院UCL霍克斯研究所;医学物理与生物医学工程系) Department of Medical Physics and Biomedical Engineering, University College London(耶鲁大学生物与生物医学科学系) Department of Biological and Biomedical Sciences, Yale University(牛津大学工程科学系生物医学工程研究所) Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford

AI总结 研究少样本医学图像分割中支持集选择问题,提出强化学习框架,智能体直接预测最大化下游分割性能的支持集,实验表明该方法优于随机选择和现有方法,凸显支持集互补性及强化学习的潜力。

Comments Accepted at EMA4MICCAI 2026 - The 2nd MICCAI Workshop on Efficient Medical AI

详情
AI中文摘要

少样本学习使医学图像分割模型仅用少量标记示例就能适应新任务。然而,适应性能很大程度上取决于为支持集选择哪些示例。有效的支持集应捕捉目标域内的相关变化并为适应提供信息,其组成样本应提供互补信息。尽管如此,现有的主动数据选择方法大多单独优先考虑样本,未明确考虑示例间的相互作用。在这项工作中,我们提出了一个用于少样本医学图像分割中支持集选择的强化学习框架,使支持集能联合优化而非通过独立样本评分。给定一组未标记的候选图像,智能体直接预测能最大化下游分割性能的支持集。在跨机构盆腔MRI数据集上的实验表明优于随机选择和当前最先进的方法。我们的发现凸显了支持集互补性对有效适应的重要性,并证明了强化学习在优化适应集方面的潜力。

英文摘要

Few-shot learning enables medical image segmentation models to adapt to new tasks using only a small number of labelled examples. However, adaptation performance depends strongly on which examples are selected for the support set. Effective support sets should capture relevant variation within the target domain and be informative for adaptation, with constituent samples providing complementary information. Despite this, existing active data selection approaches largely prioritise samples individually and do not explicitly account for interactions between examples. In this work, we propose a reinforcement learning framework for support-set selection in few-shot medical image segmentation, enabling support sets to be optimised jointly rather than through independent sample scoring. Given a pool of unlabelled candidate images, an agent directly predicts a support set that maximises downstream segmentation performance. Experiments on a cross-institutional pelvic MRI dataset demonstrate improvements over random selection and current state-of-the-art methods. Our findings highlight the importance of support-set complementarity for effective adaptation and demonstrate the potential of reinforcement learning for optimising adaptation sets.

URL PDF HTML 收藏
2607.22368 2026-07-27 cs.AI 新提交

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

智能体基准测试能衡量能力吗?智能体人工智能时代的协议有效性

Jiaqi Shao, Hanck Chen, Wei Zhang, Maxm Pan, Bing Luo

机构 * Tencent(腾讯) The Hong Kong University of Science and Technology(香港科技大学) Duke Kunshan University(昆山杜克大学)

AI总结 研究智能体基准测试分数能否衡量能力,提出协议有效性概念及HackDetect事后审计方法,通过审计多基准测试轨迹发现暴露和奖励破解证据,量化分数膨胀,强调基准测试报告应证明分数反映预期能力。

详情
AI中文摘要

智能体基准测试越来越多地评估存储库编辑、网络研究、终端使用和长期交互。只有当评估协议保持成功所需的预期能力时,其分数才支持能力声明。近期的奖励破解基准测试和系统报告表明,智能体可以恢复公共解决方案、读取评估工件、推断生成器结构、操纵反馈或受益于无效评分路径;现有应对措施未提供归因这些捷径并量化其在基准测试中影响的通用程序。我们制定了协议有效性并引入了HackDetect,这是一种事后审计,可识别暴露情况,确定智能体如何利用它,并评估所得分数是否具有误导性。我们用误导差距量化分数膨胀,即利用分数减去预期分数。我们对15个智能体基准测试中的2385条轨迹进行审计,发现67.0%的前沿科学轨迹和66.7%的自动实验室任务存在暴露和奖励破解证据。在配对比较中,我们测得分数膨胀为0.45 - 1.00,表明基准测试报告应提供分数反映预期能力的证据。

英文摘要

Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across benchmarks. We formulate protocol validity and introduce HackDetect, a post-hoc audit that identifies an exposure, determines how the agent used it, and assesses whether the resulting score is misleading. We quantify score inflation with the Mislead gap, defined as the exploit score minus the intended score. We audit 2,385 traces across 15 agent benchmarks and find evidence of exposures and reward hacking in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks. Across paired comparisons, we measure score inflation of 0.45-1.00, showing that benchmark reports should provide evidence that scores reflect the intended capability.

URL PDF HTML 收藏
2607.22367 2026-07-27 cs.LG cs.AI 新提交

Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

通过注意力展开实现内部可解释性:Transformer中的收缩和传播概况

Umberto Biccari, Qian Huang, Enrique Zuazua

机构 * University of Deusto(德乌斯托大学) Universidad Carlos III de Madrid(马德里卡洛斯三世大学) Universidad Autónoma de Madrid(马德里自治大学)

AI总结 研究Transformer内部可解释性,引入基于传播视角的内部可解释性,用注意力展开实现。通过收缩理论分析其传播概况,在代谢组年龄预测模型中有应用,还与其他方法比较,揭示了变量一致性情况,用作注意力介导传播诊断。

详情
AI中文摘要

特征归因方法为输入变量与模型输出分配分数,但本身并未描述明确定义的交互算子如何在中间层组合。我们引入了内部可解释性,这是一种基于传播的内部模型组织视角,并使用注意力展开为表格Transformer实例化。我们将展开解释为编码特征令牌之间注意力介导传播的行随机算子。通过应用经典的多布林-多布鲁申收缩理论,我们表明具有小多布鲁申系数的展开算子在数量上接近秩一随机矩阵,其公共行由其归一化列和确定。该结果为相应的展开传播概况提供了结构解释。在用于代谢组年龄预测训练的Transformer中,测量的展开收缩随深度增强。训练和随机初始化的模型也表现出不同的传播概况,尽管当前实验未确定单个展开排名变量的预测相关性。与PCA和SHAP的GradientExplainer近似的探索性比较揭示了高排名变量之间的局部一致性,但完整排名之间的一致性较弱。因此,注意力展开在这里用作注意力介导传播的诊断,而不是作为完整Transformer的因果解释或忠实归因。

英文摘要

Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator encoding attention-mediated propagation between feature tokens. By applying classical Doeblin--Dobrushin contraction theory, we show that a rollout operator with a small Dobrushin coefficient is quantitatively close to a rank-one stochastic matrix whose common row is determined by its normalized column sums. This result gives a structural interpretation to the corresponding rollout propagation profile. In Transformers trained for metabolomic age prediction, the measured rollout contraction strengthens with depth. Trained and randomly initialized models also exhibit different propagation profiles, although the present experiments do not establish the predictive relevance of individual rollout-ranked variables. Exploratory comparisons with PCA and GradientExplainer approximations to SHAP reveal localized agreement among highly ranked variables but weak agreement across complete rankings. Attention rollout is therefore used here as a diagnostic of attention-mediated propagation, not as a causal explanation or faithful attribution of the complete Transformer.

URL PDF HTML 收藏
2607.22365 2026-07-27 cs.AI 新提交

Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal Reasoning

学习结构收敛:用于时间推理的神经符号基准测试

Michael Romei De Socio, Gian Luca Pozzato, Alessio Merlo

机构 * University of Turin(都灵大学) CASD – School of Advanced Defense Studies(高级国防研究学院)

AI总结 介绍用于高复杂性事件驱动系统时间结构推理的TRACTA基准测试,含三个任务,比较多种模型。结果显示基于语义轨迹的时间建模效果最佳,消融与捷径诊断分析揭示相关信息,表明语义基础轨迹对时间结构推理有效,支持语义接口研究。

详情
AI中文摘要

高复杂性操作环境需要能检测和预测时间分布模式而非对孤立事件进行分类的方法。本文介绍了TRACTA(时间推理与能力-轨迹分析),这是一个通过类似多域操作(MDO)场景实例化的、用于高复杂性事件驱动系统中时间结构推理的受控合成基准测试。该基准测试包括三个任务,并比较了原始事件神经模型、轻契约语义基线和基于语义基础轨迹运行的神经符号配置。结果表明原始事件级学习仍有信息价值,但基于语义能力和上下文直接影响轨迹的时间建模获得了最高的总点估计,消融分析表明能力动态、上下文影响和时间结构贡献了互补信息,捷径诊断表明在主要神经输入视图中最直接的跨运行全局标识符捷径得到了控制。总体而言,研究结果支持一个有限的方法结论:在受控合成环境中,语义基础轨迹为时间结构推理提供了有效表示,支持对事件数据、结构化表示和时间学习之间语义接口的进一步研究。

英文摘要

High-complexity operational environments require methods that detect and anticipate temporally distributed patterns rather than classify isolated events. This paper introduces TRACTA (Temporal Reasoning and Capability-Trajectory Analysis), a controlled synthetic benchmark for temporal structural reasoning in high-complexity event-driven systems, instantiated through Multi-Domain Operations (MDO)-like scenarios. The benchmark includes three tasks: early_warning, pattern_detection, and run_classification, and compares raw-event neural models, a contract-lite semantic baseline, and a neuro-symbolic configuration operating on semantically grounded trajectories. Results show that raw event-level learning remains informative, but learned temporal modeling over semantic capability and contextual direct-impact trajectories achieves the highest aggregate point estimates, with the largest margins on the temporal tasks. Ablation analysis indicates that capability dynamics, contextual impacts, and temporal structure contribute complementary information. Shortcut diagnostics indicate that the most direct cross-run global-identifier shortcut is controlled in the primary neural input view, while residual shallow signals remain. Overall, the findings support a bounded methodological conclusion: in controlled synthetic settings, semantically grounded trajectories provide an effective representation for temporal structural reasoning, supporting further investigation of semantic interfaces between event data, structured representations, and temporal learning.

URL PDF HTML 收藏
2607.22361 2026-07-27 cs.LG cs.AI 新提交

Indexing: the Beginning and the End

索引:起点与终点

Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia

机构 * CENIA(CENIA研究所) Pontifical Catholic University of Chile(智利天主教大学)

AI总结 研究现代深度学习架构中信息瓶颈,通过索引原语视角,引入因果复杂度,分析索引在输入不同位置时各架构解决索引原语的能力,得出不可能性结果,实验与理论定性相符。

详情
AI中文摘要

我们通过索引原语的视角研究现代深度学习架构(循环神经网络、softmax 变换器、线性注意力变换器和状态空间模型)中的信息瓶颈。在此原语中,输入由 n 位和一个从 1 到 n 的整数 i(称为索引)组成,输出等于第 i 位的值。我们引入了掩码架构的因果复杂度。当索引出现在输入末尾时,具有低因果复杂度的架构无法在任何固定层数内解决索引原语问题,低参数循环神经网络、状态空间模型和掩码线性注意力变换器受此限制。而小型 softmax 变换器可在一层解决,非掩码线性注意力变换器可在两层解决。当索引出现在开头时,小型循环神经网络能在一层解决,其他架构则需两层。所有不可能性结果都是无条件的,实验也定性地与理论相符。

英文摘要

We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bits and one integer $i$ from $1$ to $n$ called the index, and the output equals the value of the $i$-th bit. We introduce causal complexity for masked architectures. We show that architectures with low causal complexity cannot solve the indexing primitive in any constant number of layers when the index appears at the end of the input. In particular, this limitation applies to low-parameter RNNs, SSMs and masked linear-attention transformers. In contrast, small softmax transformers can solve it in one layer, while non-masked linear-attention transformers can solve it in 2, which separates them from their masked counterparts. In turn, when the index appears at the beginning, we show that small RNNs are capable of solving this task in 1 layer, while all the other architectures require 2. All our impossibility results are unconditional and apply even to models that employ infinite-precision real arithmetic. Moreover, experiments for up to $n=64$ qualitatively align with our theory: configurations with low-parameter theoretical solutions learn the indexing task easily, while configurations that do not admit such theoretical solutions struggle to learn as the sequence length grows.

URL PDF HTML 收藏
2607.22356 2026-07-27 cs.LG 新提交

Integrated Order Dispatching and Routing for Last-Mile Pickup via Deep Reinforcement Learning

通过深度强化学习实现最后一英里取件的集成订单调度与路由

Yida Xu, Zhaofang Mao, Yuheng Miao, Jiaxin Zhang, Yiting Sun

机构 * College of Management and Economics, Tianjin University(天津大学管理与经济学部) Laboratory of Computation and Analytics of Complex Management Systems (CACMS), Tianjin University(天津大学复杂管理系统计算与分析实验室)

AI总结 针对最后一英里取件操作中订单调度和路由决策复杂的问题,提出集成优化框架,结合路由预言机与实时调度启发式方法,开发相关网络编码器和解码器及调度启发式方法,实验表明该方法能有效支持物流公司解决此类问题。

详情
AI中文摘要

近年来,最后一英里取件操作日益复杂,这增加了物流平台快速准确决策的需求。此挑战主要由订单调度和路由这两个关键且紧密相关的决策过程驱动。分别解决它们会忽略其相互依存关系,而完全的端到端学习在大规模、可变规模实例上由于稀疏奖励可能不稳定且成本高。为解决此问题,我们提出一个集成优化框架,将学习到的路由预言机与实时调度启发式方法相结合。对于路由子问题,我们开发了一个带有前瞻快递员个性化解码器的动态残差图注意力网络编码器。对于调度子问题,我们开发了一种带有局部搜索的路由预言机引导的调度启发式方法,其中预言机提供接近最优的解决方案来选择候选快递员,同时保持实时可扩展性。我们使用来自菜鸟物流的真实世界数据集进行了广泛实验,包括离线评估和在线滚动时域模拟。实验结果表明,我们的方法在解决方案质量和求解时间方面优于其他基准,表明它可以有效地支持物流公司解决实时和大规模的最后一英里取件问题。

英文摘要

In recent years, the growing complexity of last-mile pickup operations has increased the need for fast and accurate decision-making on logistics platforms. This challenge is fundamentally driven by two key and tightly coupled decision-making processes: order dispatching and routing. Solving them separately overlooks their interdependence, while fully end-to-end learning can be unstable and costly on large, variable-scale instances due to sparse rewards. To solve this problem, we propose an integrated optimization framework which couples a learned routing oracle with real-time dispatching heuristics. For the routing subproblem, we develop a Dynamic-Residual Graph Attention Network encoder with a Look-Ahead Courier-Personalized decoder. For the dispatching subproblem, we develop a routing-oracle-guided dispatching heuristic with local search, where the oracle provides near-optimal solutions to select candidate couriers while retaining real-time scalability. Extensive experiments on real-world datasets from Cainiao Logistics are used to test the performance of our approach, including an offline evaluation and an online rolling-horizon simulation. The experimental results show that our approach outperforms other benchmarks regarding solution quality and solving time, indicating it can effectively support logistics companies in solving real-time and large-scale last-mile pickup problems.

URL PDF HTML 收藏