arXivDaily arXiv每日学术速递 周一至周五更新
今天一直被爬虫使劲儿爬,网站有点不稳定
全部学科分类 2218
2608.11205 2026-08-12 cs.CV 新提交

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

AdvFD:通过对抗性Fréchet距离损失提升视觉生成性能

Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang

机构 * Peking University(北京大学) KlingAI(可灵AI)

AI总结 本研究提出AdvFD算法,通过补充静态Fréchet损失的对抗学习表示并引入真实特征白化,解决Fréchet攻击问题,提升了JiT、pMF等骨干网络的单步生成器后训练性能。

Comments Project Page: https://gasaiyu.github.io/AdvFD-page/

详情
AI中文摘要

Fréchet距离近来已成为一种有效的分布级生成器后训练目标,可补充传统的样本级扩散和流匹配损失。然而,直接优化Fréchet目标会导致Fréchet攻击:目标指标持续提升,但视觉质量和其他特征空间中的Fréchet对齐可能停滞或恶化。我们将该失败归因于现有Fréchet损失使用的静态预训练特征空间,这些空间仅提供真实与生成分布差异的不完整且固定的视图。为解决此局限,我们提出对抗性Fréchet距离(AdvFD),它用经校准的对抗学习表示补充FD-Loss中的静态表示目标。AdvFD在原始静态Fréchet目标基础上增加了可学习表示,该表示会以对抗方式最大化真实与生成样本间的Fréchet差异,而生成器则在所得自适应特征空间中最小化该差异。为防止对抗表示通过特征放大 trivial 地增大目标,我们进一步引入真实特征白化,它对其尺度和协方差几何进行归一化,以稳定极小极大优化。大量实验表明,AdvFD在JiT和pMF两种骨干网络及不同模型规模下,均能持续改进单步生成器后训练。

英文摘要

Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization. Extensive experiments show that AdvFD consistently improves one-step generator post-training across both JiT and pMF backbones and across different model scales.

URL PDF HTML 收藏
2608.11203 2026-08-12 cs.CV 新提交

Capturing Uncertainty in Human Motion for Representation Learning in Soccer

捕捉足球运动中人类动作的不确定性以用于表示学习

Yizhou Xu, Lars Bretzner, Tiesheng Wang, Atsuto Maki

机构 * KTH Royal Institute of Technology(皇家理工学院) EA Sports TRACAB

AI总结 本文提出一种用于足球3D骨骼动作的自监督表示学习框架,通过建模未来动作概率分布提升预测准确率,该表示可迁移至多个足球下游应用,展现强跨任务泛化能力。

详情
AI中文摘要

本文提出一种用于理解足球中基于3D骨骼的人类动作的自监督表示学习框架,采用未来动作预测作为学习目标。由于人类动作固有不确定性,考虑多种合理未来动作对捕捉潜在动作动态并学习有效表示至关重要。为此,我们引入一种用于动作预测的条件模块,该模块在3D欧几里得空间中对离散未来动作的概率分布进行建模,通过未来轨迹的显式监督学习多模态特性。在大规模足球运动员跟踪数据上的实验表明,我们的方法大幅提升了动作预测准确率。此外,学习到的表示可有效迁移至多个足球下游应用,展现出强大的跨任务泛化能力。

英文摘要

This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion prediction as the learning objective. Since human motion is inherently uncertain, accounting for multiple plausible futures is essential for capturing the underlying motion dynamics and learning effective representations. To this end, we introduce a conditioning module for motion prediction that models a probabilistic distribution over discretized future motions in 3D Euclidean space, learning multimodality with explicit supervision from future trajectories. Experiments on large-scale soccer player tracking data show that our approach substantially improves motion prediction accuracy. Moreover, the learned representations effectively transfer to multiple soccer downstream applications, demonstrating strong cross-task generalization.

URL PDF HTML 收藏
2608.11201 2026-08-12 cs.CV 新提交

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

VidForensics-M1:用于AI生成视频取证的、具备可验证时间定位的元检测强化学习

Bowei Liu, Zheng Lu, Yuhan Bian, Xinchen Zhang, Xingming Shui, Yuesheng Huang, Xuhuan Li, Zihao Liu, Yifan Yang, Jun Zhou, Xiu Li

机构 * Tsinghua University(清华大学) Peking University(北京大学) Renmin University of China(中国人民大学) Microsoft(微软公司)

AI总结 该研究针对AI生成视频检测泛化性不足的问题,首次将元检测引入该领域,提出结合可验证时间证据的VidForensics-M1模型及相关机制,实现了鲁棒可泛化的检测。

Comments 27 pages, 15 figures

详情
AI中文摘要

近期视频生成模型的进展显著提升了合成视频的真实感,模糊了生成内容与真实内容的边界,引发了对虚假信息传播的担忧。现有基于多模态大语言模型(MLLM)的检测器主要依赖监督微调或标签级强化学习,其中粗粒度的监督限制了其对未见过场景和新兴视频生成器的泛化能力。为克服这些局限,我们首次将元检测引入AI生成视频检测,通过在强化学习中联合优化预测标签与支撑证据,实现可靠的伪造检测。该范式需要可靠的证据信号及将其整合到标签级优化的有效机制。文本理由提供伪造痕迹的语义描述,但其生成与验证依赖外部模型,导致监督易受幻觉和语义偏差影响。相比之下,时间定位提供更客观、可验证的证据,因为伪造构建过程中可精确控制被篡改的时间区间。基于此洞见,我们提出自动化数据构建流水线,通过用边界帧条件视频生成模型替换时间片段,生成配对的真实-伪造视频。此外,我们引入证据引导的奖励再分配机制,该机制根据证据质量在标签正确的响应间重新分配奖励,在保留可靠标签监督的同时,鼓励检测器获取细粒度且可验证的伪造定位能力。大量实验表明,VidForensics-M1可有效利用可验证的时间证据,实现鲁棒且可泛化的AI生成视频检测。

英文摘要

Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and authentic content and raising concerns about misinformation. Existing MLLM-based detectors mainly rely on supervised fine-tuning or label-level reinforcement learning, where coarse supervision limits generalization to unseen scenarios and emerging video generators. To overcome these limitations, we are the first to introduce \textbf{meta-detection} into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning. This paradigm requires reliable evidence signals and effective mechanisms to integrate them into label-level optimization. Textual rationales provide semantic descriptions of forgery artifacts, but their generation and verification depend on external models, making supervision vulnerable to hallucinations and semantic biases. In contrast, temporal grounding provides more objective and verifiable evidence, as manipulated intervals can be precisely controlled during forgery construction. Based on this insight, we propose an automated data construction pipeline that generates paired real-fake videos by replacing temporal segments with boundary-frame-conditioned video generation models. Furthermore, we introduce \textbf{Evidence-Guided Reward Redistribution}, which performs evidence-aware credit assignment by redistributing rewards among label-correct responses according to evidence quality. This preserves reliable label supervision while encouraging detectors to acquire fine-grained and verifiable forgery localization capabilities. Extensive experiments demonstrate that \textbf{VidForensics-M1} effectively leverages verifiable temporal evidence to achieve robust and generalizable AI-generated video detection.

URL PDF HTML 收藏
2608.11200 2026-08-12 cs.CL cs.AI cs.LG 新提交

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

ConVAWG:面向针对妇女和女孩的暴力场景下可控合成对话生成的检索驱动框架

Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola

机构 * University of Sheffield(谢菲尔德大学) Forensic Capability Network(法医能力网络) University of Leeds(利兹大学) University of Warwick(华威大学)

AI总结 本研究针对VAWG场景建模的空白,提出检索驱动框架ConVAWG,生成符合CPS标准的合成VAWG多轮对话,发布6000余个对话事件,经多类评估验证其质量与领域保真度。

详情
AI中文摘要

合成对话生成为研究敏感领域的对话动态提供了途径,在这些领域中,真实数据难以获取、发布或标注。潜在的虐待可能发生在线上或线下:威胁和胁迫可直接出现在消息中,而监视、孤立、跟踪和身体暴力等行为可能被计划、披露或通过对话提及。隐私和法律约束使得发布大规模真实对话数据集变得困难;现有工作大多聚焦于线上虐待的句子级毒性,而在将虐待建模为一种关系性且随时间展开的现象方面存在空白。本研究将针对妇女和女孩的暴力(VAWG)场景建模为多轮对话。我们提出ConVAWG,这是一个用于生成符合CPS标准的合成VAWG聊天对话的检索驱动框架。ConVAWG基于角色种子、英国国家统计局报告的人口统计模式、官方犯罪定义以及检索到的家庭凶杀案审查案例构建场景;将其转换为分层事件时间线;生成多场景角色扮演对话;并对适当的话语应用针对性的激活引导毒性控制。我们发布了涵盖200个场景的6000多个多轮对话事件,带有丰富的场景级、事件级和轮级元数据。广泛的人工评估、大模型作为评判者的评估、消融实验以及下游任务表明,该框架具有出色的对话质量和领域保真度。

英文摘要

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally. Privacy and legal constraints make it difficult the release of large-scale real conversation datasets; existing work has mostly focused on sentence-level toxicity of online abuses, leaving a gap in modelling abuse as a relational and temporally unfolding phenomenon. In this work, we focus on modelling Violence Against Women and Girls (VAWG) scenarios as multi-turn dialogues. We introduce ConVAWG, a retrieval-grounded framework for generating CPS-aligned synthetic VAWG chat dialogues. ConVAWG builds scenarios from persona seeds, demographic patterns reported by the UK Office for National Statistics, official crime definitions, and retrieved Domestic Homicide Review cases; converts them into hierarchical event timelines; generates multi-scene role-play dialogues; and applies targeted activation-steered toxicity control to appropriate utterances. We release over 6,000 multi-turn dialogue events across 200 scenarios with rich scenario-, event-, and turn-level metadata. Extensive human evaluation, LLM-as-Judge assessment, ablations, and downstream tasks show strong dialogue quality and domain fidelity.

URL PDF HTML 收藏
2608.11197 2026-08-12 cs.LG cs.CL 新提交

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

超越特征袋:稀疏自编码器中的集合级不稳定性

Nikolai Bolik, Lennart Stöpler, Artur Andrzejak

机构 * Heidelberg University(海德堡大学)

AI总结 该研究以稀疏自编码器(SAE)隐集合重叠度为相似度度量,发现SAE特征不通过简单特征袋语义组合,其激活集合与人类概念判断存在显著不匹配。

详情
AI中文摘要

Shani等人(2026)表明,大语言模型(LLM)表示大体上能恢复人类的类别边界,但无法反映细粒度的典型性结构。他们的分析采用了密集模型表示上的余弦相似度。我们使用活跃稀疏自编码器(SAE)隐集合的重叠度作为更具可解释性的相似度度量,重新审视其方法。我们首先验证该集合级度量的意义:SAE隐集合可在受控玩具模型中恢复类组合结构,并在自然文本中诱导语义连贯的邻域。将人类概念分析扩展到SAE集合相似度后,我们发现SAE激活集合相比密集嵌入或残差流状态,并未更忠实地恢复人类类别边界或类别内典型性,而是追踪模型内部的相似度结构。为进一步探究该差距,我们在受控语义修改下研究活跃隐集合,发现人类对概念变化的判断与SAE活跃集合的变化存在显著不匹配。我们将此解释为证据:在非理想环境中,SAE特征并非通过简单的特征袋语义组合。

英文摘要

Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.

URL PDF HTML 收藏
2608.11175 2026-08-12 cs.RO cs.SY eess.SY 新提交

Risk-Aware Kinodynamic Motion Planning Under Uncertainty For Safe Navigation on Planetary Environments

行星环境安全导航下考虑不确定性的风险感知动力学运动规划

Sachin Sunil Kelkar, Tanmay Dokania, Yashwanth Kumar Nakka

机构 * Daniel Guggenheim School of Aerospace Engineering(丹尼尔·古根海姆航空航天工程学院) Georgia Institute of Technology(佐治亚理工学院)

AI总结 针对自主空间探索中环境交互与感知系统的不确定性导致的安全规划问题,提出结合AO-RRT采样规划与序列凸规划的风险感知动力学运动规划方法,可将轨迹风险降低约97%。

Comments 3 pages, 4 figures

详情
AI中文摘要

对于自主空间探索而言,机器人智能体需要执行运动规划,而环境交互可能是未知的。学习这类交互(例如轮式机器人的地形力学)会引入不确定性,进而导致危险的运动规划,可能引发危险操作或任务失败。此外,感知系统引发的不确定性会加剧安全运动规划的问题。在本论文中,我们研究具有风险感知的成本最优动力学运动规划问题,分两步解决:第一步,基于采样的规划器AO-RRT生成动态可行、风险感知且渐近成本最优的轨迹;第二步,我们将运动规划建模为非线性优化问题,以AO-RRT轨迹为初始解,使用序列凸规划(SCP)求解。通过使用条件风险价值(CVaR)量化风险,我们在仿真和硬件实验的所有轨迹中证明风险降低了约97%以上。

英文摘要

For autonomous space exploration, robotic agents need to perform motion planning in which environmental interactions may be unknown. Learning these interactions, such as terrain mechanics for wheeled robots, can introduce uncertainties that lead to risky motion plans and potentially hazardous operations or mission failures. Moreover, uncertainties induced by perception-based systems can exacerbate the problem of safe motion planning. In this letter, we address the problem of performing cost-optimal kinodynamic motion planning with risk awareness. We approach this in two steps. First, a sampling-based planner (AO-RRT) generates a dynamically feasible, risk-aware, and asymptotically cost-optimal trajectory. Second, we formulate motion planning as a nonlinear optimization problem and solve it using sequential convex programming (SCP), using the AO-RRT trajectory as an initial solution. By quantifying risk using conditional value-at-risk (CVaR), we demonstrate a reduction in risk by over $\sim$97\% across trajectories in simulation and hardware experiments.

URL PDF HTML 收藏
2608.11174 2026-08-12 cs.RO 新提交

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

VIScore:诊断潜在世界模型中与规划相关的质量

Haiyu Wu, Randall Balestriero, Morgan Levine

机构 * Altos Labs(阿尔托斯实验室) Brown University(布朗大学)

AI总结 该研究针对潜在世界模型规划成功率与潜在空间属性脱节的问题,提出VIScore指标,覆盖编码器、预测器、规划器,在跨任务场景下斯皮尔曼相关性超0.75,可更好解释规划成功率。

详情
AI中文摘要

将潜在空间正则化为各向同性高斯分布,可为世界模型规划提供稳定且信息最大化的空间。然而,潜在空间属性与成功规划之间仍存在脱节。我们通过比较SIGReg和VISReg两种正则化损失函数来研究这一问题,二者具有相同的分布目标但属性不同。与SIGReg相比,VISReg在控制中心、尺度和形状正则化的权重方面具有更高的灵活性,且更大的批量大小可实现更精细的分布近似。我们发现,前者虽在自监督学习(SSL)中有益,但对规划无帮助;而后者可提高分布外(OOD)数据集上的规划成功率。这促使我们深入研究与成功率相关的因素。与仅关注编码潜在空间的现有指标不同,我们提出了真实性-影响力-清醒度分数(VIScore),该指标可量化给定编码特征的预测器的可达性和容量,以及基于搜索的规划器的幻觉。与直线度、物理状态探测和赋能相比,我们证明,由于VIScore的测量覆盖了编码器、预测器和规划器,其比其他指标更能解释成功率,表现为强斯皮尔曼相关性。具体而言,在跨任务成功率池中,VIScore在已见和未见模型及数据集上均始终实现超过0.75的斯皮尔曼相关性。此外,VIScore是唯一在所有测试场景中校准误差低于常数拟合的指标,凸显了这三个方面对规划成功的重要性。我们希望该指标能助力未来世界模型的设计与诊断研究。

英文摘要

Regulating the latent space to an isotropic Gaussian distribution provides a stable and information-maximized landscape for world model planning. However, the latent space property and successful planning remain disconnected. We first study this by comparing SIGReg and VISReg, two regularization loss functions with the same distribution target but different properties. Compared with SIGReg, VISReg has more flexibility in controlling the weights of center, scale, and shape regularization, and a larger batch size brings a finer distribution approximation. We find that the former, despite being beneficial in self-supervised learning (SSL), does not help the planning, whereas the latter improves the planning success on out-of-domain (OOD) datasets. This motivates a deep understanding of the factors that correlate with the success rate. Unlike the previous metrics focusing on the encoded latent only, we propose the Veracity-Influence-Sobriety score (VIScore), a metric that quantifies the reachability and capacity of a predictor given the encoded feature, and the hallucination of the searching-based planner. Compared with straightness, physical-state probing, and empowerment, we show that, with the measurement covering encoder, predictor, and planner, VIScore explains the success rate better than the others, as reflected by a strong Spearman correlation. Specifically, VIScore consistently achieves a Spearman correlation over 0.75 on both seen and unseen models and datasets on the cross-task success rate pool. Moreover, VIScore is the only metric that has a calibration error below the constant fit across all testing scenarios, showcasing the importance of these three aspects in planning success. We hope this metric can help future studies on world model design and diagnosis.

URL PDF HTML 收藏
2608.11171 2026-08-12 cs.CL cs.AI cs.CY 新提交

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

从可解释性到控制:TrustNLP研讨会六年的启示

Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan

机构 * Meta Autodesk(欧特克公司) New Jersey Institute of Technology(新泽西理工学院) Salesforce(salesforce公司) University of California, Los Angeles(加州大学洛杉矶分校) Amazon AGI(亚马逊AGI)

AI总结 该研究基于TrustNLP研讨会六年论文,分析NLP可信领域从可解释性到生成式系统控制的转变,明确各信任维度的发展趋势及与领域整体的关联性,提出结构性见解与研究方向。

Comments 17 pages, 2 figures, 3 tables. Submitted to ACL ARR August 2026 cycle (EACL 2027)

详情
AI中文摘要

可信自然语言处理研讨会(TrustNLP)自2021年起与重要的ACL会议联合举办,历经六届,论文数量从8篇增至41篇,记录了该领域从静态模型的事后可解释性向生成式系统的机制理解与主动控制的转变。我们综合了所有144篇会议论文的见解,依据成熟框架(TrustLLM、DecodingTrust)确立的六个信任维度对其进行分类,观察到其与能力涌现存在共现关系。首批高影响力聊天模型的发布同时激活了所有信任维度,而后续模型代际则将重点转向真实性与安全对齐。分类研究的分析显示,真实性是增长最快的维度(2021-2022年不存在,到2025-2026年占论文的37%),公平性仍是最稳定的主题,可解释性则呈现U型轨迹:因事后方法失去相关性而下降,但在2026年通过机制可解释性再度兴起。同期与ACL、NAACL、EACL及EMNLP(约2000篇论文)的跨会场对比表明,TrustNLP的主题分布与领域平均水平高度吻合。我们确定了四个结构性见解,并为研究界总结了可操作的方向。

英文摘要

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.

URL PDF HTML 收藏
2608.11167 2026-08-12 cs.CV cs.CL cs.LG 新提交

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

多模态代码切换:将视觉对象交织入语言以实现显式对象级对齐

Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

AI总结 针对现有多模态大语言模型的图像级对齐存在指称歧义的问题,提出多模态代码切换(MMCS)范式,构建含77.3万样本的数据集,仅用5万样本即可匹配或超越60万图像-文本对训练的模型,提升了视觉基础与感知能力。

详情
AI中文摘要

现有多模态大语言模型(Multimodal Large Language Models, MLLMs)主要依赖图像-文本对进行模态对齐预训练,将全局图像表示映射为长文本描述。然而这种图像级对齐存在指称歧义:模型难以从全局表示中推断多个视觉对象与文本实体之间的对应关系,导致数据效率低下和语义基础不佳。为解决该问题,我们提出多模态代码切换(MultiModal Code-Switching, MMCS),一种提供显式对象级监督的新型预训练范式。受代码切换的语言现象启发,MMCS通过用对应视觉对象替换文本实体来交织视觉与语言,强化局部视觉-语言基础。我们进一步开发可扩展的数据合成流水线,生成包含77.3万个样本的预训练数据集,具备准确的对象-实体对应关系。实验表明MMCS数据效率极高:仅用5万个样本,其性能可匹配或超越在60万个图像-文本对上训练的模型;此外,MMCS在不同模型规模下均持续提升视觉基础与感知能力。

英文摘要

Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision. Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding. We further develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object-entity correspondences. Experiments show that MMCS is highly data-efficient: with only 50K samples, it matches or surpasses models trained on 600K image-text pairs. Furthermore, MMCS consistently improves visual grounding and perception capabilities across varying model scales.

URL PDF HTML 收藏
2608.11162 2026-08-12 cs.LG 新提交

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

分层经验贝叶斯朴素贝叶斯:极小极大平滑与校准及AODE扩展

Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu

机构 * Van Lang University(范朗大学) Ho Chi Minh City University of Transport(胡志明市交通大学)

AI总结 该研究提出HEB-NB及HEB-AODE,解决NB固定平滑强度致高基数表格数据偏差问题,经31个基准测试,其在概率指标、对数损失及ECE上均有显著改进。

Comments This manuscript has been submitted to the Knowledge-Based Systems

详情
AI中文摘要

朴素贝叶斯(NB)分类器仍是分类数据的标准选择,但其广泛使用的平滑规则(如拉普拉斯、利德斯通、克里切夫斯基-特罗菲莫夫和m估计)均规定固定平滑强度,忽略特征基数、样本量和类别不平衡,在现代高基数表格数据上引入非零偏差。我们提出分层经验贝叶斯朴素贝叶斯(HEB-NB),其中每个类别-特征条件概率由狄利克雷先验平滑,该先验的浓度通过第二类最大似然数据自适应学习,实现跨类别的合理信息共享,同时保留闭式推理。我们进一步引入HEB平均单依赖估计器(HEB-AODE),表明自适应平滑可清晰迁移至NB的结构松弛。理论上,我们为HEB-NB建立非渐近ℓ₁误差界,匹配经验分布极小极大率加非零数据自适应偏差,同时得到匹配拉普拉斯的紧下界,产生与拉普拉斯的有限样本风险级严格分离。我们还通过总变差张量化推导插件式 excess贝叶斯风险界,并得到总体Top-1预期校准误差(ECE)推论。实证上,在31个UCI和OpenML基准中,HEB-NB在概率指标上取得最佳平均弗里德曼秩,在高基数数据集上对数损失降低达22.1%,且HEB-AODE较普通AODE持续改进。结合HEB-NB与互信息加权使Top-1 ECE降低41%-70%,在概率准确性和校准上取得显著提升。

英文摘要

The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data. We propose hierarchical empirical-Bayes Naive Bayes (HEB-NB), in which each class-feature conditional probability is smoothed by a Dirichlet prior whose concentration is learned data-adaptively via Type-II maximum likelihood, enabling principled information sharing across classes while retaining closed-form inference. We further introduce HEB average one-dependence estimators (HEB-AODE), showing that the adaptive smoothing transfers cleanly to structural relaxations of NB. Theoretically, we establish a non-asymptotic $\ell_1$ error bound for HEB-NB matching the empirical-distribution minimax rate plus a vanishing data-adaptive bias, together with a matching Laplace-tight lower bound that yields a finite-sample, risk-level strict separation from Laplace. We further derive a plug-in excess Bayes-risk bound via total-variation tensorization and a population top-1 expected calibration error (ECE) corollary. Empirically, across 31 UCI and OpenML benchmarks, HEB-NB attains the best average Friedman rank on probabilistic metrics, with up to 22.1% log-loss reductions on high-cardinality datasets and consistent improvements of HEB-AODE over vanilla AODE. Combining HEB-NB with mutual-information weighting reduces top-1 ECE by 41%-70%, demonstrating substantial gains in probabilistic accuracy and calibration.

URL PDF HTML 收藏
2608.11154 2026-08-12 cs.LG 新提交

DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains

DACRI:面向关键供应链的决策感知因果干预排名

Shiqi Huang, Jiani He, Dingyan Shang, Yihua Xu, Jize Li, Yan Lyu, Lashimi Muraleedharan Nair

机构 * Independent Researcher(独立研究者)

AI总结 该研究提出CriticalSCM-Bench v1基准,对比LambdaMART与恒定缓冲策略等,明确自适应干预排名在关键供应链场景的适用范围及模型复杂度的价值边界。

Comments Accepted for presentation at the International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME 2026), 15--17 October 2026, Bali, Indonesia. 5 tables; no figures. Benchmark and code: https://github.com/dyshang/dacri-criticalscm-bench

详情
AI中文摘要

检测或归因供应链中断与选择能最大化可恢复净值的干预措施并非同一回事。我们提出CriticalSCM-Bench v1,这是一个包含因果真实值、配对事实/反事实推演以及明确净值目标的受控合成基准。相较于全信息训练选择的静态基准,LambdaMART的中位数标准化净值提升了5.7%至16.2%,在半导体和关键材料原型上有配对统计支持,但在数字基础设施上无此支持。在数字基础设施领域,基于领域知识的恒定缓冲策略表现仍更优,这表明模型复杂度的提升并非总能得到合理证明。在部分和延迟场景中,LambdaMART保留了全夹紧值的33%至75%。压力测试进一步显示,干预保真度、时机、成本及保留的中断事件可改变策略排序,关键材料表现出最弱的分布外保留能力。此外,一项针对540代的受保护解释研究在确定性验证和模板回退后保留了所有固定干预决策,尽管确切措辞仍不稳定。在该受控环境中,研究结果明确了自适应排名能产生价值的场景,以及更简单的结构性策略仍更优的场景。

英文摘要

Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a controlled synthetic benchmark with causal ground truth, paired factual/counterfactual rollouts, and an explicit net-value objective. Relative to a full-information train-selected static benchmark, LambdaMART improves median normalized net value by 5.7--16.2\%, with paired statistical support on the semiconductor and critical-material archetypes but not on digital infrastructure. On digital infrastructure, a domain-informed constant-buffer policy remains stronger, showing that greater model complexity is not uniformly justified. Across partial and delayed settings, LambdaMART retains 33--75\% of full-clamp value. Stress tests further show that intervention fidelity, timing, cost, and held-out disruptions can alter policy ordering. Critical materials show the weakest out-of-distribution retention. Separately, a guarded explanation study over 540 generations preserves every fixed intervention decision after deterministic validation and template fallback, although exact wording remains unstable. Within this controlled setting, the results identify regimes in which adaptive ranking adds value and those in which simpler structural policies remain preferable.

URL PDF HTML 收藏
2608.11150 2026-08-12 cs.CV 新提交

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

CausalSplat:面向3D高斯溅射的综合层级推理

Jiayu Ding, Meilu Song, Yun Chen, Wei Gao, Ge Li

机构 * Peking University(北京大学) North China Electric Power University(华北电力大学) Hunan University(湖南大学)

AI总结 针对3D高斯溅射推理局限,本文提出CausalSplat框架,结合视觉语言模型与3D场景图,构建两个推理基准,在相关任务上实现最优性能与强泛化性。

Comments Accepted to ECCV 2026

详情
AI中文摘要

尽管3D高斯溅射(3DGS)已推动开放词汇场景理解的发展,但现有方法仍局限于显式查询,难以解释具身交互所需的隐式意图、复杂空间约束与常识推理。为填补这一空白,本文提出3D高斯分割推理任务,并构建两个基准数据集Causal-LERF与Causal-ScanNet,系统评估常识、空间、 affordance(功能可供性)及反事实推理能力。评估显示,当前最优方法在这些推理挑战上表现较差。因此,本文提出CausalSplat框架,将视觉语言模型与3D场景图结合,以解耦显式结构感知与隐式逻辑推理。大量实验表明,CausalSplat在本文的推理基准上达到最优性能,同时在标准指称与开放词汇3D分割任务上展现出强泛化性。项目页面:this https URL

英文摘要

While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal-LERF and Causal-ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat

URL PDF HTML 收藏
2608.11143 2026-08-12 cs.LG 新提交

A Recommendation System Approach for Interference-Robust Sensor Subset Selection

面向抗干扰传感器子集选择的推荐系统方法

Kaan Buyukkalayci, Kyle Pak, Merve Karakas, Christina Fragouli

机构 * University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

AI总结 本文针对现有基于RSSI的传感器子集选择方法易受声学干扰的问题,提出结合频带声学特征与双塔MLP架构的推荐系统框架,在户外车辆跟踪任务中较RSSI基线提升约20%精度且保持低计算开销。

详情
AI中文摘要

本文提出一种用于跟踪任务的传感器子集选择方法。现有研究表明,低成本的声学接收信号强度指示器(RSSI)测量值可用于推荐传感器节点子集,这些节点的昂贵传感模态(如摄像头)能实现高跟踪精度。尽管基于RSSI的方法效率较高,但易受声学干扰影响。我们提出一种受推荐系统启发的框架,该框架利用频带声学特征和双塔多层感知器(Two-Tower MLP)架构对候选传感器子集进行高效评分。户外车辆跟踪部署的实验结果显示,与RSSI基线相比,所提方法可将精度提升约20%,同时保持实时选择性传感所需的低计算开销。

英文摘要

This paper develops a method for sensor-subset selection for tracking. Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive sensing modalities, such as cameras, can achieve high tracking accuracy. While efficient, RSSI-based approaches are challenged by acoustic interference. We propose a recommendation-system-inspired framework that instead leverages frequency-band acoustic features and a Two-Tower Multi-Layer Perceptron (MLP) architecture to efficiently score candidate sensor subsets. Experimental results on outdoor vehicle-tracking deployments show that the proposed method can improve accuracy by around 20\% over the RSSI baseline while maintaining the low computational overhead required for real-time selective sensing.

URL PDF HTML 收藏
2608.11136 2026-08-12 cs.AI 新提交

sLTN: Structural Logic Tensor Networks

sLTN:结构逻辑张量网络

Davide Rinaldi, Luciano Serafini

机构 * Nokia Bell Labs(诺基亚贝尔实验室) Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)

AI总结 本文提出LTN的扩展版本sLTN,将结构维度作为一等元素,实现时间、序列等结构约束的逻辑表达,完成其形式化与PyTorch实现,在时序推理示例中验证了框架有效性。

详情
AI中文摘要

逻辑张量网络(LTN)是一种神经符号框架,其中一阶逻辑通过张量运算进行解释,可将逻辑约束与可微学习相融合。但原始LTN公式主要适用于以个体的扁平集合表示的数据,未明确捕捉时间顺序、序列位置或图连接性等结构组织。本文提出sLTN,即LTN的扩展版本,它将结构维度作为语言的一等元素。结构维度表示与特定领域组织相关的命名张量轴,如时间步、序列位置或图节点,可被显式量化、通过结构关系关联,并直接在逻辑层面表达时间、序列和关系约束。本文对sLTN的语法和模糊张量语义进行了形式化,证明在无结构维度时,该框架会退化为原始LTN语义作为特例。此外,本文描述了基于声明式签名、公式解析和张量解释的PyTorch实现,并在代表性的时间和序列推理示例中对该框架进行了说明。本文是sltn库的配套论文,该库可在指定网址获取。

英文摘要

Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tensor operations, enabling logical constraints to be integrated with differentiable learning. However, the original formulation of LTN is primarily suited to data represented as flat collections of individuals, and does not explicitly capture structural organization such as temporal order, sequential position, or graph connectivity. We introduce sLTN, an extension of LTN that makes structural dimensions first-class elements of the language. Structural dimensions represent named tensor axes associated with domain-specific organization, such as time steps, sequence positions, or graph nodes. They can be quantified explicitly, related through structural relations, and used to express temporal, sequential, and relational constraints directly at the logical level. We formalize the syntax and fuzzy tensor semantics of sLTN and show that, in the absence of structural dimensions, the framework recovers the original LTN semantics as a special case. We further describe a PyTorch implementation based on a declarative signature, formula parsing, and tensorial interpretation. The framework is illustrated on representative temporal and sequential reasoning examples. This paper serves as a companion to the sltn library, available at https://github.com/logictensornetworks/sltn.

URL PDF HTML 收藏
2608.11123 2026-08-12 cs.CV cs.LG 新提交

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

AlbumentationsX:适用于图像及相关标注的统一增强流水线

Vladimir Iglovikov

机构 * Albumentations LLC(阿尔布门泰森有限责任公司)

AI总结 AlbumentationsX 是一款统一图像及相关标注的增强流水线库,可避免数据错位,支持自定义变换、保存流水线,适用于 PyTorch 框架的训练流程。

Comments 8 pages, 1 figure. Source code: https://github.com/albumentations-team/AlbumentationsX

详情
AI中文摘要

当图像与其标注接收不同的随机变化时,数据增强可能会破坏训练样本。裁剪操作必须对图像、掩码、边界框、关键点、立体视图、视频帧或体素使用相同的坐标,若代码路径分别选择这些值,会导致数据静默错位。AlbumentationsX 将变换列表、概率、标注设置和随机种子保存在一个 Compose 对象中,每次调用仅选择一次随机值并将其应用于训练样本的所有支持部分。该库会将每个对象的掩码、边界框和标签保持在一起,允许项目添加自定义变换,还可保存流水线定义、展示单次调用的操作结果并重新运行该调用。示例中 Compose 位于文件解码为数组后、PyTorch 将样本分组为批次前,AlbumentationsX 会执行声明的变换,从业者仍需根据任务决定翻转、裁剪、颜色变化等操作是否保留正确标签。

英文摘要

Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. Code paths that choose these values separately can silently misalign the data. AlbumentationsX keeps the transform list, probabilities, annotation settings, and random seed in one Compose object. Each call chooses random values once and applies them to every supported part of the training example. The library keeps each object's mask, box, and label together and lets projects add their own transforms. It can also save the pipeline definition, show what happened in one call, and run that call again. The examples place Compose after files have been decoded into arrays and before PyTorch groups examples into a batch. AlbumentationsX executes the declared transforms. Practitioners still decide whether a flip, crop, color change, or other operation preserves the correct label for their task.

URL PDF HTML 收藏
2608.11110 2026-08-12 cs.CL 新提交

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

行动胜于言语:测量使用工具的智能体的跨语言策略保留率

Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram

机构 * Microsoft Research India(微软研究院印度分部)

AI总结 该研究测量使用工具的智能体的跨语言行动策略保留率,发现前沿模型在贪心解码下保留71%-73%策略,参数低于100亿时保留率崩溃,且存在影响测量的混淆因素。

Comments Accepted in COLM 26

详情
AI中文摘要

当使用工具的智能体在不同语言中被赋予相同任务时,它是否仍会采取相同步骤?多语言评估很少提出这个问题:它们比较最终答案,却忽略了行动。然而这些行动是成果的体现:它们决定成本和延迟,决定系统如何失败,也是其行为唯一可审计的部分。我们以行动策略为测量对象,涉及8个模型、6个并行基准和41种语言(共238万次 rollout)。朴素测量方法失效:原始轨迹相似度与任何可辩护结论之间存在5个混淆因素,每个因素都能翻转结论:短轨迹得分更高,空轨迹得分完美,无关轨迹有超过一半的时间偶然一致,差距受每个模型的可复现性限制,且同一模型在一种语言中被问两次相同问题会给出不同答案,没有基线。我们消除了所有5个混淆因素,每一次修正都使效应更大。差异证明是结构性的,而非采样噪声:它在每个单元中都能在贪心解码下存活,且随着温度升高保持稳定,即使模型的自一致性降低。以自身可复现性归一化后,4个不同的前沿模型在贪心解码下收敛,每个模型在跨语言中保留71%-73%的行动策略,模型身份仅解释5.7%的方差。在参数规模约100亿以下时,保留率会崩溃,而较小模型之间的排序在很大程度上是偶然基线的产物,我们通过排列而非假设来测量该基线。智能体通过英语路由非英语任务;这种转换具有因果重要性,已通过4个模型的预注册预测得到确认,且模型在被要求时不会放弃这种转换。最后,单个轨迹提取正则表达式(而非模型)造成了多语言失败:两个工作示例使一个模型的测量准确率提升了26倍,而其在可读输出上的准确率几乎没有变化。

英文摘要

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.

URL PDF HTML 收藏
2608.11096 2026-08-12 cs.CV 新提交

Every Packet Counts: Dispersing Information for Loss-Resilient Learned Image Compression

每个数据包都重要:面向抗丢包的学习型图像压缩的信息分散方法

Yuhang Wei, Chuqin Zhou, Yibo Shi, Jing Wang, Guo Lu

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Technologies Ltd.(华为技术有限公司) Central Media Technology Institute(中央媒体技术研究院)

AI总结 本文针对学习型图像压缩的抗丢包问题,提出ICR、ICG机制及双层双分支自回归结构,实验显示其在丢包场景下的重建质量和稳定性均优于现有方法,且可泛化到突发丢包场景。

Comments 16 pages, 12 figures, 8 tables. Joint first authors: Yuhang Wei and Chuqin Zhou. Corresponding author: Guo Lu. To appear in Proceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10-14, 2026, Rio de Janeiro, Brazil

详情
AI中文摘要

学习型图像压缩(LIC)已取得令人瞩目的率失真性能,但现有方法对数据包丢失仍高度敏感,而数据包丢失是卫星通信和应急通信中的常见挑战。这种脆弱性源于分组阶段的非均匀信息分布,以及熵编码阶段的顺序解码依赖关系。本文提出一种端到端抗丢包图像压缩方案,同时解决这两个问题:在分组前,引入信道间重分配(ICR)机制以重新分配信道能量,防止关键信息集中在小部分信道中;随后,采用交错信道分组(ICG)策略以步长方式划分潜在信道,将信息分散到各个数据包中,且每个数据包保持在受限大小内;为限制丢包导致的级联误差,采用双层双分支自回归结构以缩短依赖链。大量实验表明,本文方法在重建质量和稳定性上均优于现有方法:在丢包率为20%时,相比LossResilientLIC,其平均峰值信噪比(PSNR)提升1.84 dB,同时PSNR方差降低一个数量级;值得注意的是,仅在均匀随机丢包下训练的模型,可泛化到由Gilbert-Elliott信道建模的突发丢包场景,且优于专门针对此类条件训练的方法。

英文摘要

Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss, a common challenge in satellite and emergency communications. This vulnerability stems from non-uniform information distribution at the packetization stage and sequential decoding dependencies at the entropy coding stage. We propose an end-to-end loss-resilient image compression scheme that addresses both. Before packetization, we introduce an Inter-Channel Redistribution (ICR) mechanism to redistribute channel energy, preventing critical information concentrating in a small subset of channels. Then, an Interleaved Channel Grouping (ICG) strategy partitions latent channels in a strided manner to disperse information across packets, with each packet kept within constrained sizes. To limit cascading errors from lost packets, we adopt a two-layer dual-branch autoregressive structure to shorten the dependency chain. Extensive experiments demonstrate that our method consistently outperforms existing approaches in both reconstruction quality and stability. At 20% packet loss, it achieves an average PSNR gain of 1.84 dB over LossResilientLIC while reducing PSNR variance by an order of magnitude. Notably, trained under uniform random loss only, our model generalizes to bursty loss modeled by the Gilbert-Elliott channel, outperforming methods explicitly trained for such conditions.

URL PDF HTML 收藏
2608.11095 2026-08-12 cs.AI cs.LG cs.SE 新提交

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

CLAUDE.md 为何不断膨胀?智能体编码中的灾难性记忆

Kushal Chakrabarti

机构 * South Park Commons(南公园社区)

AI总结 该研究针对智能体编码中提示文件无限制膨胀的问题,提出通过提示注释消除多余指令,可使现实中智能体指令遵循能力提升多达23.1%。

详情
AI中文摘要

像 http URL 这样的智能体编码 README 文件在真实代码仓库中会无限制地增长,仅在仓库停用或有人彻底重写该文件时才会停止。我们将此归因于不完善的记忆:追加一条指令总是成本低廉,但一旦某条指令的基本原理消失,删除它而不冒正确性回归风险的成本为 O(2^|D|),其中 |D| 是提示中的指令数量。我们将这种偏差命名为灾难性记忆,它是持续学习所围绕的灾难性遗忘的逆过程。首先,我们在 1867 个仓库的 247694 条指令生命周期中对该现象进行了表征:智能体提示会无限制增长,在其生命周期内增长超过三倍(+226%),每次提交净增加 +4.9 条指令;此外,指令越旧,被删除的可能性越小(对数风险为 -0.032/次提交)。然后,我们证明提示注释可以阻止这种增长:对 IFEval 进行反转可得到可验证的世界,其最优提示是已知的,而编码潜在推理的注释可消除 99.3% 的多余指令(从 +211.3% 降至 +1.4%)。最后,将相同的反转应用于 WildIFEval,我们表明提示注释可将现实世界中智能体的指令遵循能力提升多达 23.1%。如果英语是新的代码,我们为何还没有注释?

英文摘要

Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D| instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which continual learning is organized. First, we characterize this phenomenon across 247,694 instruction lifetimes in 1,867 repositories: agentic prompts grow without bound, more than tripling over their lifetime (+226%), gaining +4.9 net instructions every commit; further, the older an instruction gets, the less likely it is to be deleted (log-hazard -0.032/commit). Then, we show that prompt comments can halt the growth: inverting IFEval yields verifiable worlds whose optimal prompts are known, and there comments encoding latent reasoning remove 99.3% of excess instructions (+211.3% to +1.4%). Finally, applying the same inversion to WildIFEval, we show that prompt comments can improve real-world agentic instruction-following by up to 23.1%. If English is the new code, why don't we have comments yet?

URL PDF HTML 收藏
2608.11093 2026-08-12 cs.LG cs.CV 新提交

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

跨视图特征匹配:综述、基准测试与基础模型视角

Songlin Du, Xiaoyong Lu, Zeyu Wu, Xiaobo Lu, Guobao Xiao, Bin Fan, Jiayi Ma, Takeshi Ikenaga

机构 * School of Automation, Southeast University(东南大学自动化学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) Institute of Artificial Intelligence, University of Science and Technology Beijing(北京科技大学人工智能研究院) School of Robotics, Wuhan University(武汉大学机器人学院) Graduate School of Information, Production and Systems, Waseda University(早稻田大学信息生产系统研究科)

AI总结 本综述梳理跨视图特征匹配领域进展,构建结构化分类体系,开展统一基准测试,提炼设计原则,探讨开放挑战,为该领域发展提供全面参考。

Comments This manuscript goes beyond a conventional survey. It proposes a new taxonomy for cross-view feature matching, provides extensive benchmarking under unified datasets and protocols, and offers original analysis from the perspective of vision foundation models. These contributions provide substantive methodological synthesis, empirical findings, and new research insights

详情
AI中文摘要

跨视图特征匹配旨在在视角差异极大的图像间建立可靠的对应关系。过去十年,该领域已从特定任务模型发展为日益统一、可泛化的对应模型,近期视觉基础模型(Vision Foundation Models, VFMs)的出现进一步推动了研究进展。尽管取得了这些进展,但现有研究在问题表述、模型架构、训练范式和评估协议方面仍存在高度多样性,导致难以对该领域形成统一理解。本综述对跨视图特征匹配进行了系统性梳理:首先介绍涵盖特征提取、单类型特征匹配器、多类型特征匹配器、基于VFMs的方法、训练策略与鲁棒估计的结构化分类体系,为分析与比较提供连贯框架;进一步考察近期进展,提炼关键设计原则,强调向统一、可泛化对应模型的转变;还在统一协议下对代表性最先进方法开展了系统性实验基准测试,实现公平且全面的性能对比;此外,探讨了效率、极端条件下鲁棒性、跨域泛化等开放挑战与未来方向。本综述旨在为理解视觉基础模型时代跨视图特征匹配的演变、当前态势及未来发展提供全面且结构化的参考。

英文摘要

Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation protocols, making it difficult to obtain a unified understanding of the field. In this survey, we present a unified review of cross-view feature matching. We first introduce a structured taxonomy covering feature extraction, single-type feature matcher, multi-type feature matcher, VFMs based methods, training strategy and robust estimation, providing a coherent framework for analysis and comparison. We further examine recent advances, distilling key design principles and highlighting the shift toward unified and generalizable correspondence models. We also provide a unified experimental benchmarking of representative state-of-the-art methods under consistent protocols, enabling fair and comprehensive performance comparisons. In addition, we discuss open challenges and future directions, including efficiency, robustness under extreme conditions, and cross-domain generalization. This survey aims to provide a comprehensive and structured reference for understanding the evolution, current landscape, and future development of cross-view feature matching in the era of vision foundation models.

URL PDF HTML 收藏
2608.11080 2026-08-12 cs.AI 新提交

RTSKG: Building a Rail Transit Station Knowledge Graph Dataset

RTSKG:构建轨道交通站点知识图谱数据集

Shutong Zhu, Tianxing Wu, Runfeng Liu, Yuang Gu, Xuan He, Yuan Zhu

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education(教育部新一代人工智能技术及其交叉应用重点实验室) School of Architecture, Southeast University(东南大学建筑学院)

AI总结 针对城市级轨道交通站点任务数据组织中忽视城市实体交互的问题,构建RTSKG数据集并验证其在商铺推荐和客流量预测中的有效性,可支撑城市级轨道交通站点分析。

Comments 21 pages, Accepted by ISWC 2026

详情
AI中文摘要

轨道交通系统在城市交通和经济发展中发挥着至关重要的作用。作为此类系统的关键组成部分,轨道交通站点是重要的交通枢纽,可提升城市可达性并带动周边区域发展。城市级轨道交通站点相关任务(如客流量预测)需要大规模城市数据,但现有研究在数据组织方面往往忽视各类城市实体间的复杂交互。为解决上述问题,本文构建了轨道交通站点知识图谱(Rail Transit Station Knowledge Graph,RTSKG)数据集,该数据集明确建模不同类型城市实体间的空间与语义交互,以助力城市级轨道交通站点相关任务。RTSKG 采用专门设计的统一模式整合了轨道交通站点、道路路段、兴趣点等异构城市实体,可作为关联数据通过指定 URL 获取。对站点区域商铺推荐和知识增强型客流量预测的评估验证了 RTSKG 的有效性,凸显其支持城市级轨道交通站点分析的潜力。

英文摘要

Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems, rail transit stations function as critical transport hubs that enhance urban accessibility and stimulate development in surrounding areas. City-level rail transit station related tasks (e.g., ridership prediction) require large-scale urban data, but current studies often neglect complex interactions among various urban entities in terms of data organization. In this paper, to address the above issue, we build a Rail Transit Station Knowledge Graph (RTSKG) dataset which explicitly models the spatial and semantic interactions among different kinds of urban entities, to benefit city-level rail transit station related tasks. RTSKG integrates heterogeneous urban entities, such as rail transit stations, road segments, and points of interest, with a specially designed unified schema, and is accessible as Linked Data at https://w3id.org/rtskg/. Evaluations on station-area store recommendation and knowledge-enhanced ridership prediction demonstrate the effectiveness of RTSKG, highlighting its potential to support city-level rail transit station analysis.

URL PDF HTML 收藏
2608.11079 2026-08-12 cs.AI 新提交

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

SkillZip:通过发现可重用结构实现自进化智能体的免评估技能压缩

Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li

机构 * Alibaba Group(阿里巴巴集团) Zhejiang University(浙江大学) Duke University(杜克大学)

AI总结 SkillZip 是一种免评估的自进化智能体技能压缩方法,通过提炼重复规则与动作序列实现高效压缩,在性能、泛化性及成本上优于评估引导压缩。

详情
AI中文摘要

自进化智能体通过追加成功的过程与失败修复来积累可重用技能。随着时间推移,同一需求常被在多个分支、示例和警告中重复表述,而常见动作序列被复制而非重用,导致技能注入成本高昂且难以维护。通用提示压缩不适用于此场景,因为技能并非扁平文本:其名称和描述定义适用时机,工作流控制执行,工具与输出契约约束有效性,即使无采样任务激活,罕见异常也可能至关重要。评估引导压缩可测试这些行为,但会引入 rollout(试执行)、成本,并依赖压缩时的评估集。我们提出 SkillZip,一种免评估方法,通过寻找技能最短的忠实结构解释来压缩技能。其核心思路是“一次解释,多次引用”:在适用范围处单次陈述重复规则,将重复动作序列提炼为共享过程,仅保留差异作为显式异常。我们将此思路形式化为基于技能契约与残差的类型化最小描述长度目标,需满足每个提取的触发器、工作流边、工具需求、义务及输出字段的硬覆盖约束。该公式提供简单的共享阈值,通过构造保留唯一罕见规则,支持高效局部更新。SkillZip 具有单次模式(含一次结构化提取调用与确定性优化),以及持续的“写时压缩(Zip-on-Write)”模式,可整合每次自进化补丁,无需重放任务或重新解析完整历史。通过全面实验评估,我们证明 SkillZip 在压缩性能、可泛化性及成本开销方面的有效性与优越性。

英文摘要

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.

URL PDF HTML 收藏
2608.11076 2026-08-12 cs.CV 新提交

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

基于基础模型的高效数据采样(FEEDS):用于泛癌、多示踪剂PET/CT数据集的标签高效训练策略

Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James B. Yu, Indrani Bhattacharya

机构 * Geisel School of Medicine at Dartmouth(达特茅斯盖泽尔医学院) Dartmouth Hitchcock Medical Center(达特茅斯-希区柯克医疗中心) Yale University(耶鲁大学)

AI总结 该研究提出FEEDS策略,利用视觉基础模型嵌入选择信息性和多样性未标注病例,在减少70%标注负担时达到全标注训练性能,解决了PET/CT病灶分割的标签稀缺问题。

Comments Code is publicly available on https://github.com/Image-and-Multimodal-Data-Analytics/FEEDS

详情
AI中文摘要

全身PET/CT成像中的自动化病灶分割可辅助临床医生在不同放射性示踪剂和癌症类型下开展癌症检测、分期及治疗规划。然而,训练能捕捉病灶大小、分布及外观差异的病灶分割模型需要大量标注数据集,其创建既耗时又依赖专业知识。因此,在有限标注PET/CT数据上训练的模型往往缺乏临床应用所需的准确性和泛化能力。我们提出FEEDS(Foundation model-Enabled Efficient Data Sampling,基于基础模型的高效数据采样),这是一种标签和计算高效的学习策略,利用视觉基础模型的嵌入来选择最具信息性和多样性的未标注病例供专家标注。与无监督、半监督及主动学习方法不同,FEEDS是一种仅需有限、代表性训练集的单步训练范式,兼具标签和计算高效性。我们使用AutoPET-III数据集对FEEDS进行训练和验证,在三个保留测试集(AutoPET-III、DeepPSMA及达特茅斯-希区柯克医学中心的内部数据集)上测试其准确性和泛化性,在体素、病灶及解剖区域层面评估临床效用,以评估高危区域的性能和治疗规划效用。FEEDS的表现优于基于随机采样的标注、基于伪标签的半监督学习以及仅用有限标注数据的训练,可在所有三个测试集、FDG和PSMA示踪剂及多种疾病间实现泛化,在减少70%标注负担的情况下达到了100%全标注训练的性能。FEEDS通过提供一种从大型未标注临床库中构建具有代表性和多样性的标注队列的实用方法,解决了自动病灶分割框架中的标签稀缺问题。

英文摘要

Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotracers and cancer types. However, training lesion segmentation models that capture variations in lesion size, distribution, and appearance requires large annotated datasets, whose creation is both time- and expertise-intensive. As a result, models trained on limited labeled PET/CT data often lack the accuracy and generalizability needed for clinical use. We present FEEDS (Foundation model-Enabled Efficient Data Sampling), a label- and compute-efficient learning strategy that uses vision foundation model embeddings to select the most informative and diverse unlabeled cases for expert annotation. Unlike unsupervised, semi-supervised, and active learning approaches, FEEDS is a one-step training paradigm requiring only a limited, representative training set, making it label- and compute-efficient. We train and validate FEEDS using the AutoPET-III dataset. We test its accuracy and generalizability on three held-out sets: AutoPET-III, DeepPSMA, and an internal Dartmouth-Hitchcock Medical Center dataset. We evaluate clinical utility at the voxel, lesion, and anatomic region level to assess performance in high-risk areas and treatment planning utility. FEEDS outperforms random-sampling-based labeling, pseudolabel-based semi-supervised learning, and training with limited labeled data alone. It generalizes across all three test sets, FDG and PSMA tracers, and multiple diseases, matching fully-labeled (100\%) training performance with 70\% less annotation burden. FEEDS addresses the challenge of label scarcity in an automatic lesion segmentation framework by providing a practical approach for constructing representative and diverse annotation queues from large, unannotated clinical repositories.

URL PDF HTML 收藏
2608.11075 2026-08-12 cs.CV 新提交

Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues

帧中静态,事件中动态:将事件相机的特征重新思考为运动线索

Hesam Araghi, Jan van Gemert, Nergis Tomen

机构 * Delft University of Technology(代尔夫特理工大学)

AI总结 本文将事件相机特征视为运动线索,分析了结构张量特征值与时空密度值,结合局部几何特征增强运动估计,在DSEC基准上提升了光流网络的准确性,数据稀缺与低容量模型增益显著。

详情
AI中文摘要

事件相机以高时间分辨率异步捕捉强度变化,这需要为下游任务开发新颖的预处理方法。与静态的强度快照不同,事件数据固有地编码了场景动态和物体运动的信息,这意味着从事件中衍生的特征可能表现出与基于帧的视觉没有直接相似性的行为。在本文中,我们分析了基于事件的角点检测中使用的两个特征——结构张量的特征值和时空密度值——并证明它们是运动线索。我们假设这些特征结合局部几何信息可以增强运动估计任务。为了验证这一点,我们首先从理论上分析了运动角点处结构张量的特征值与运动方向的关系。然后,我们在合成数据集上设计了受控实验,证实用特征值和密度值扩展局部几何特征可提供互补的运动信息,且对纹理和散粒噪声具有鲁棒性。最后,我们将所提出的特征集成到最先进的基于事件的光流网络中,并在真实世界的DSEC基准上进行评估,结果显示添加的特征始终提高了准确性,在数据稀缺场景和低容量模型中获得的增益最大。本文的代码可在此处获取:this https URL。

英文摘要

Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike static intensity snapshots, event data inherently encode information about scene dynamics and object motion, meaning that features derived from events can exhibit behaviors with no direct analogue in frame-based vision. In this paper, we analyze two features used in event-based corner detection---the eigenvalues of the structure tensor and the spatiotemporal density values---and show that they are \emph{motion cues}. We hypothesize that these features, combined with local geometric information, can enhance motion estimation tasks. To validate this, we first theoretically analyze how the eigenvalues of the structure tensor at moving corner points relate to the direction of motion. We then design controlled experiments on a synthetic dataset, confirming that extending local geometric features with eigenvalues and density values provides complementary motion information and is robust to texture and shot noise. Finally, we integrate the proposed features into a state-of-the-art event-based optical flow network and evaluate on the real-world DSEC benchmark, where the added features consistently improve accuracy, with the largest gains in data-scarce scenarios and for lower-capacity models. The code for this paper can be found at: \href{https://github.com/hesamaraghi/static-in-frames-dynamic-in-events}{https://github.com/hesamaraghi/static-in-frames-dynamic-in-events}.

URL PDF HTML 收藏
2608.11074 2026-08-12 cs.CV 新提交

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

CapProbe:通过全场景密集问答评估详细图像描述

Mouxiao Huang, Qiangyu Yan, Borui Jiang, Han Shu

机构 * Huawei Technologies(华为技术有限公司)

AI总结 CapProbe是用于评估VLMs生成的详细图像描述的全场景密集问答基准,通过区域对齐的事实核查,揭示了13个VLMs的覆盖差距、能力-效率权衡及遗漏的失败模式,相关资源将很快发布。

详情
AI中文摘要

评估视觉语言模型(VLMs)生成的详细图像描述,不能仅停留在表层语义相似性层面。基于参考的指标(如CIDEr和SPICE)以及“LLM作为评分者”协议难以验证密集事实断言,而现有的基于问答(QA)的替代方案通常存在探测密度较低、领域覆盖较窄,或未明确建立单个问题与分割图像区域之间对齐关系的问题。我们提出CapProbe,这是一个全场景密集问答基准,将详细图像描述评估转化为区域对齐的事实核查任务。每张图像被分解为覆盖前景和背景元素的粗粒度语义区域;对于每个保留的区域,我们生成涵盖10个语义类别的多项选择题,形成密集的已探测视觉事实清单。在由37个L1领域和219个L2子领域组成的两级分类法指导下,CapProbe包含346张图像、1868个区域和25650个问题,平均每张图像对应74个问答对。语言评分者仅根据描述进行回答;“不确定”选项和有效准确率提供了一种依赖于评分者的代理指标,用于区分未回答的探测项与错误解决的探测项,而基于密度的指标会惩罚冗长但无信息的描述。该协议具有成本效益:通过将无约束的标量评分转化为结构化的多项选择题阅读任务,它减少了开放式评分偏差,同时保持了评分者条件性,并且在固定阅读器下产生相对稳定的模型排名。对13个VLMs的实验显示,模型之间存在较大的覆盖差距、明显的能力-效率权衡,以及稀疏或基于重叠的评估通常会遗漏的失败模式。基准数据、注释和评估代码将很快发布。

英文摘要

Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and LLM-as-scorer protocols struggle to verify dense factual claims, while existing QA-based alternatives generally offer lower probe density, narrower domain coverage, or no explicit alignment between individual questions and segmented image regions. We introduce CapProbe, a full-scene dense QA benchmark that turns detailed caption evaluation into region-aligned factual checking. Each image is decomposed into coarse semantic regions covering both foreground and background elements; for every retained region, we generate multiple-choice questions spanning 10 semantic categories, forming a dense checklist of probed visual facts. Guided by a two-tier taxonomy of 37 L1 domains and 219 L2 sub-domains, CapProbe comprises 346 images, 1,868 regions, and 25,650 questions, averaging 74 QA pairs per image. A language judge answers from the caption alone; an Uncertain option and Effective Accuracy provide a judge-dependent proxy for distinguishing unanswered probes from incorrectly resolved ones, while density-based metrics penalize verbose yet uninformative captions. The protocol is cost-effective: by converting unconstrained scalar scoring into structured MCQ reading, it reduces open-ended scoring bias while remaining judge-conditioned and yields relatively stable model rankings under a fixed reader. Experiments on 13 VLMs show large Coverage gaps across models, a clear competency-efficiency trade-off, and failure modes that sparse or overlap-based evaluation often misses. The benchmark data, annotations, and evaluation code will be released soon.

URL PDF HTML 收藏
2608.11064 2026-08-12 cs.CV cs.AI 新提交

Entropy-Centric Explainable AI for Remote Sensing Image Segmentation

面向遥感图像分割的以熵为核心的可解释人工智能

Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany, Ali J. Ghandour

机构 * Faculty of Engineering, Lebanese University(黎巴嫩大学工程学院) University of Paris-Est Créteil (UPEC)(巴黎东部克雷泰伊大学) EFREI(EFREI学院) National Center for Remote Sensing - CNRS(国家遥感中心 - 法国国家科学研究中心)

AI总结 针对遥感图像分割的黑箱模型透明度问题,提出以熵为核心的可解释人工智能方法及新评估方法,实验验证其优于现有适配的语义分割XAI方法。

详情
AI中文摘要

人工智能(AI)已成为解决关键领域复杂问题的强大方法,但其模型的决策过程引发诸多担忧,主要原因是深度神经网络在性能超越同类模型的同时,特征提取与预测存在模糊性。在遥感等关键领域,需使用黑箱模型分析高分辨率图像,透明度的缺失限制了对这些模型的信任,进而阻碍了其应用。鉴于此,解释和理解AI模型的复杂决策过程变得至关重要,可解释人工智能(XAI)旨在提供决策方式与原因的洞见,弥合这一差距。尽管图像分类任务的解释已取得显著进展,但图像分割领域仍有较大提升空间。在此背景下,本文提出一种以熵为核心的语义分割XAI方法,还提出一种新的XAI评估方法,用于高效衡量该方法突出区域的相关性。实验结果表明,与近期适配的语义分割XAI方法相比,所提XAI方法具有优越性。

英文摘要

Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and prediction. Consequently, in critical domains such as remote sensing, where high-resolution imagery must be analyzed using black-box models, the lack of transparency limits trust in these models and, thus, their adoption. In light of this reality, explaining and understanding the complex decision-making process of AI models has become essential. Explainable AI (XAI) aims to bridge this gap by providing insights into how and why certain decisions are made. While significant progress has been achieved in explaining image classification tasks, image segmentation still offers considerable room for improvement. In this context, this paper proposes an entropy-centric XAI method for semantic segmentation. Moreover, a new XAI evaluation methodology is proposed to efficiently measure the relevance of the regions highlighted by the proposed XAI method. Experimental results demonstrate the superiority of the proposed XAI method compared with recently adapted XAI methods for semantic segmentation.

URL PDF HTML 收藏
2608.11063 2026-08-12 cs.RO 新提交

Deployment Is Not Destiny: Robot Recomposition in the Field with Unseen Software, Hardware, and Compute Payloads

部署并非宿命:面向现场未知软件、硬件与计算负载的机器人重组

Steven Swanbeck, Jonathan Salfity, Jeffery Gunawan, Corrie Van Sice, Mitch Pryor, Robert Blake Anderson

机构 * Texas Robotics(德克萨斯机器人研究所) Walker Department of Mechanical Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校沃克机械工程学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 该研究提出一种机器人运行时重组框架,支持现场快速集成未知软件、硬件与计算负载,将重配置时间缩至数分钟,在灾难响应场景中验证了其灵活性与实用性。

详情
AI中文摘要

大多数机器人的子系统紧密耦合,尽管是其复杂性的自然结果,但会形成整体式设计,在初始部署后难以适应且耗时。为解决这一挑战,我们提出了一种框架及配套抽象,用于运行时重组,使机器人能快速集成此前未知的模块化软件、硬件与计算负载。我们的方法允许非专家用户通过真正的即插即用流程,在现场快速添加新能力。关键在于,新资源不仅对主机机器人立即可用,还会共享给分布式对等节点,使计算受限的系统能访问强大的远程新能力。我们的框架将重新配置时间缩短至数分钟,无需开发者介入,而传统手动集成通常需要数小时的专家工作。我们在两个灾难响应场景中演示了该方法,包括在运行中的核反应堆设施进行放射源定位,以及在黑暗、难以到达的空间开展热引导人员搜索。这些演示表明,现场重组能为动态需求提供及时、灵活且可及的适应,代表着向创建能随所支持的任务、技术和环境快速演进的机器人迈出的关键一步。

英文摘要

The tight coupling of subsystems in most robots, though a natural consequence of their complexity, leads to monolithic designs that are time-consuming and difficult to adapt after initial deployment. To address this challenge, we present a framework and supporting abstractions for recomposition during runtime that enable robots to quickly integrate previously unseen modular software, hardware, and compute payloads. Our approach allows non-expert users to quickly add new capabilities in the field through a true plug-and-play process. Crucially, new resources are not only immediately available to a host robot but are also shared with distributed peers, enabling compute-constrained systems to access powerful new remote capabilities. Our framework reduces reconfiguration time to a matter of minutes with no developer intervention, in stark contrast to the hours of expert effort often required for traditional manual integration. We demonstrate our method in two disaster response scenarios, including radioactive source localization at an operational nuclear reactor facility and a thermal-guided search for people in dark, difficult-to-reach spaces. These demonstrations show how in-field recomposition provides timely, flexible, and accessible adaptation to dynamic requirements, representing a critical step toward creating robots that can quickly evolve alongside the tasks, technologies, and environments they support.

URL PDF HTML 收藏
2608.11061 2026-08-12 cs.LG 新提交

Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

批量大小还是负样本?内存受限推荐系统训练的选择规则

Artyom Sabitov, Daniil Volkov, Alexey Zaytsev

机构 * Moscow Independent Research Institute of Artificial Intelligence(莫斯科独立人工智能研究院) Intellectual data analysis and predictive modeling institute(智能数据分析与预测建模研究院) Applied AI Institute(应用人工智能研究院) Risk department(风险部门)

AI总结 该研究针对内存受限的推荐系统训练,分析固定内存预算下批量大小与负样本数量的权衡,提出优先扩大批量的选择规则,经实验验证可提升收敛速度与推荐质量。

详情
AI中文摘要

大规模神经推荐系统通常采用针对全部物品词汇表的softmax交叉熵损失函数进行训练。对于数量庞大的候选物品K,最终分类层会占据主要内存,处理n个样本的批次时需要存储O(nK)个对数几率和梯度。采样softmax通过将损失函数限制在仅k个远小于K的候选负样本上,将内存成本降至O(nk)。然而,在固定预算B = n k的条件下,究竟应优先选择更大的批量还是更多的负样本仍不明确。我们通过分析内存约束下的采样softmax训练来解决该问题。在标准平滑性和方差假设下,理论证据表明最快收敛来自n ~ B、k ~ 1的分配。因此,可执行的规则是在计算约束下尽可能包含更多样本。我们的理论得到受控合成数据及四个真实序列推荐基准(包括MovieLens-20M)的支持。在相同内存约束下,建议的配置相比不平衡替代方案实现了更快的收敛速度和更优的最终推荐质量。这些发现为推荐系统训练期间的内存配置提供了理论和实证基础。代码、可复现材料及所有生成图表的脚本可在该httpsURL获取。

英文摘要

Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $n$ examples. Sampled softmax reduces this cost by restricting the objective to only $k \ll K$ candidate negative items, resulting in an $O(nk)$ memory. However, for a fixed budget $B = n k$, it remains unclear whether one should prioritize larger batches or the inclusion of more negative items. We address this question by analyzing sampled-softmax training under a fixed memory constraint. Under standard smoothness and variance assumptions, our theoretical evidence suggests that the fastest convergence arises from an $ n \sim B, k \sim 1$ allocation. So, an actionable rule is to include as many objects as possible given computational constraints. Our theory is supported by controlled synthetic and synthetic and four real sequential recommendation benchmarks, including MovieLens-20M. The suggested configuration achieve faster convergence and better final recommendation quality than imbalanced alternatives within the same memory constraint. These findings provide a theoretical and empirical foundation for configuring memory during the training of recommender systems. Code, reproducibility materials, and all scripts for generating figures are available at https://anonymous.4open.science/r/LimitedMemoryRule-BBFB

URL PDF HTML 收藏
2608.11054 2026-08-12 cs.LG 新提交

Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study

面向基因组学应用的不确定性感知深度学习:一项实证研究的见解

Sepideh Saran, Mahsa Ghanbari, Uwe Ohler

机构 * The Berlin Institute for Medical Systems Biology(柏林医学系统生物学研究所) Max Delbrück Center for Molecular Medicine(马克斯·德尔布吕克分子医学中心) Technical University of Berlin(柏林工业大学) Humboldt University of Berlin(柏林洪堡大学)

AI总结 本研究通过对比三种不确定性量化方法,分析其在基因组学两类应用中的表现,明确了贝叶斯神经网络的优势,为基因组学UQ方法的应用提供了指南。

Comments 21 main pages, 42 total pages, 12 main figures, 13 supplementary figures

详情
AI中文摘要

深度学习模型已成为基因组学众多应用中的标准计算工具,但不确定性量化(UQ),尤其是该领域不同不确定性估计的可靠性,却很少受到系统关注。本研究针对基因组学应用开展深度学习模型中UQ的实证分析,在一系列实验中对比了Deep Ensembles、贝叶斯神经网络(Bayesian Neural Networks)和蒙特卡洛失活(Monte Carlo-dropout)方法,评估它们在不同场景下量化不确定性的能力,同时考虑两个基因组应用领域及模态的常见数据集特征:序列-活性模型和单细胞表达分析。我们的系统对比框架为基因组学中UQ方法的适用性与可靠性提供了指南,明确了它们在不同场景下的优势与局限性。研究表明,尽管贝叶斯神经网络存在计算劣势,但在捕捉基因组学中强类别不平衡和分布外数据导致的不确定性方面表现更优;此外,我们还展示了如何利用不确定性分数选择蛋白质-RNA相互作用中的高质量预测。

英文摘要

Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. In a series of experiments, we contrast Deep Ensembles, Bayesian Neural Networks, and Monte Carlo-dropout methods. We assess their ability to quantify uncertainty in different scenarios, accounting for common dataset characteristics in two genomic application areas and modalities: sequence-to-activity models, and single-cell expression analysis. Our systematic comparison framework provides guidelines for the applicability and reliability of UQ methods in genomics, highlighting their strengths and limitations in different scenarios. We show that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages. Moreover, we show how uncertainty scores can be used to select high-quality predictions in protein-RNA interactions.

URL PDF HTML 收藏
2608.11053 2026-08-12 cs.CV cs.AI 新提交

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

非洲真实多植物数据集上深度学习目标检测模型的对比评估

Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen

机构 * Bayero University(巴耶罗大学) Alpen-Adria-Universität Klagenfurt(克拉根福阿尔卑斯-亚得里亚大学) Gombe State University(贡贝州立大学) Aliko Dangote University of Science and Technology(阿里科·丹格特科技大学) Ahmadu Bello University(艾哈迈杜·贝洛大学) Federal University Dutse(杜塞联邦大学)

AI总结 该研究针对非洲真实农业场景的多植物数据集,对比评估YOLO系列、Faster R-CNN、RT-DETR等6种目标检测模型,发现RT-DETR性能最优,YOLO系列训练效率更高,为农业植物检测提供了可靠方案。

详情
AI中文摘要

计算机视觉在农业中的应用在提升作物监测和精准农业方面展现出巨大潜力,但现有许多方法依赖的受控数据集无法充分代表真实农业条件,尤其是在非洲等代表性不足的地区。本研究使用从尼日利亚农场手动采集的真实世界数据集AgriAISeg 1,对六种目标检测模型YOLOv5、YOLOv8、YOLO11、YOLO26、Faster R-CNN和RT-DETR进行对比评估。AgriAISeg包含3382张芝麻、卷心菜和番茄作物的图像,这些图像是在不同环境条件下拍摄的,包括光照变化、遮挡和视角变化。对模型进行训练后,使用精确率、召回率、mAP@0.5和mAP@0.5:0.95评估性能。结果显示,RT-DETR的整体性能最高,精确率为0.768,mAP@0.5:0.95为0.624;YOLOv8和YOLO11也表现出强劲且稳定的性能。相比之下,Faster R-CNN的准确率明显较低,整体mAP@0.5为0.466,表明其在复杂田间条件下的有效性降低。此外,基于YOLO的模型相比Faster R-CNN表现出更优的训练效率。这些发现表明,现代单阶段检测器和基于Transformer的检测器为真实世界农业环境中的植物检测提供了更可靠、高效的解决方案。

英文摘要

The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches rely on controlled datasets that do not adequately represent realworld farming conditions, particularly in underrepresented regions such as Africa. This study presents a comparative evaluation of six object detection models YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR using a real-world dataset, AgriAISeg 1 , collected manually from Nigerian farms. AgriAISeg comprises 3,382 images of sesame, cabbage, and tomato crops captured under varying environmental conditions, including changes in illumination, occlusion, and viewing perspectives. Models were trained, and performance was assessed using precision, recall, mAP@0.5, and mAP@0.5:0.95. The results show that RT-DETR achieved the highest overall performance with a precision of 0.768 and mAP@0.5:0.95 of 0.624, while YOLOv8 and YOLO11 also demonstrated strong and consistent performance. In contrast, Faster R-CNN recorded significantly lower accuracy, with an overall mAP@0.5 of 0.466, indicating reduced effectiveness under complex field conditions. In addition, YOLO-based models exhibited superior training efficiency compared to Faster R-CNN.These findings demonstrate that modern one-stage and transformer-based detectors provide more reliable and efficient solutions for plant detection in realworld agricultural environments.

URL PDF HTML 收藏
2608.11052 2026-08-12 cs.LG stat.ML 新提交

Efficient Hypergradient Descent for Inverse Reinforcement Learning

用于逆强化学习的高效超梯度下降

Nikita Sevriukov, Anna Barabanova, Uliana Gagarina, Karina Ivanova, Sofiia Kasaeva, Ilya Levin, Marina Sheshukova

机构 * HSE University(高等经济大学)

AI总结 该研究针对逆强化学习双层优化的计算挑战,利用策略费舍尔信息矩阵的特性设计结构化超梯度,通过流式频谱草图近似逆费舍尔向量乘积,在控制环境中实现了高效且性能良好的IRL方法。

详情
AI中文摘要

逆强化学习(IRL)旨在恢复一个奖励函数,使得基于该奖励函数得到的策略能够复现专家演示中观察到的行为。一种自然的方法是将IRL表述为双层优化问题,其中内层对应于在学习到的奖励下的策略优化,外层则衡量诱导策略与专家数据之间的差异。然而,这种表述在实际应用中计算难度较大,因为外层更新需要涉及内层目标的逆海森向量乘积的超梯度。我们通过证明在内层最优解处,内层目标的海森与策略的费舍尔信息矩阵成比例,从而得到了一种与自然超梯度下降密切相关的结构化基于费舍尔的超梯度,以此解决这一挑战。为解决与大型费舍尔矩阵相关的可扩展性瓶颈,我们使用流式频谱草图近似所需的逆费舍尔向量乘积,避免显式构建费舍尔矩阵。我们在离散和连续控制环境中,将我们的方法与一阶随机双层基线进行了评估。结果表明,该方法具有有竞争力的策略性能和出色的奖励排序质量,同时费舍尔草图降低了曲率存储复杂度,并且相较于显式费舍尔求解器可提高计算效率。

英文摘要

Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. A natural approach is to formulate IRL as a bilevel optimization problem, in which the inner level corresponds to policy optimization under the learned reward and the outer level measures the discrepancy between the induced policy and expert data. However, this formulation is computationally challenging in practice because the outer update requires a hypergradient involving an inverse-Hessian-vector product for the inner objective. We address this challenge by showing that, at the inner optimum, the Hessian of the inner objective is proportional to the Fisher information matrix of the policy, yielding a structured Fisher-based hypergradient closely related to Natural Hypergradient Descent. To address the resulting scalability bottleneck associated with large Fisher matrices, we approximate the required inverse-Fisher-vector product using a streaming spectral sketch, avoiding explicit construction of the Fisher matrix. We evaluate our approach against a first-order stochastic bilevel baseline across discrete- and continuous-control environments. The results demonstrate competitive policy performance and strong reward-ranking quality, while Fisher sketching reduces curvature-storage complexity and can improve computational efficiency relative to an explicit Fisher solver.

URL PDF HTML 收藏