arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
全部学科分类 1986
2607.21595 2026-07-24 cs.CV cs.AI cs.LG 新提交

3D-Aware VLMs with Implicit and Explicit Geometries

具有隐式和显式几何的3D感知视觉语言模型

Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao, Ran Xu, Shijian Lu, Gongjie Zhang

机构 * Nanyang Technological University(南洋理工大学) DAMO Academy, Alibaba Group(阿里巴巴达摩院) HuPan Lab(湖畔实验室)

AI总结 研究针对多数2D视觉输入的VLM处理3D任务困难的问题,提出VLM-IE3D框架,通过引入隐式和显式几何令牌及3D感知适配器,融合几何表示与视觉线索,在多种3D任务中表现优异。

Comments Accepted by ECCV 2026, Open Sourced

详情
AI中文摘要

尽管取得了快速进展,但大多数基于2D视觉输入构建的现有视觉语言模型(VLM)在处理需要细粒度空间理解和推理的各种3D任务时往往面临困难。为了弥合这一差距,我们提出了VLM-IE3D,这是一个统一框架,通过为VLM配备从RGB视频中学习的隐式和显式3D几何来增强其3D空间感知。我们的VLM-IE3D引入了隐式几何令牌(IGT)来从输入视频中捕获高级几何先验,以及补充的显式几何令牌(EGT)来编码来自重建3D属性的详细几何结构。在此基础上,VLM-IE3D还配备了一个3D感知适配器,可有效地将这两种几何表示与2D视觉线索融合。这种仅基于RGB的设计为细粒度空间理解和推理注入了强大的3D归纳偏差,而无需任何额外的3D输入。广泛的实验表明,VLM-IE3D在包括3D视频检测、3D视觉定位、3D密集字幕和空间推理在内的各种3D任务中始终取得优异性能。代码和模型可在该https URL获取。

英文摘要

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RGB-only design injects strong 3D inductive biases for fine-grained spatial understanding and reasoning without requiring any additional 3D inputs. Extensive experiments show that VLM-IE3D achieves superior performance consistently across various 3D tasks including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning. Code and models are available at https://github.com/Vegetebird/VLM-IE3D.

URL PDF HTML 收藏
2607.21594 2026-07-24 cs.CV 新提交

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

具有世界状态寄存器的流式多智能体自回归扩散模型

Sicheng Mo, Yuheng Li, Ziyang Leng, Krishna Kumar Singh, Bolei Zhou

机构 * University of California, Los Angeles(加利福尼亚大学洛杉矶分校) Adobe Research(Adobe研究院)

AI总结 研究多智能体交互世界模型共享状态维护难题,提出WorldWeaver模型,利用跨智能体世界状态寄存器及混合变压器设计,经双智能体我的世界视频生成实验验证,该模型能提升逻辑一致性与生成质量。

Comments Project page: https://vail-ucla.github.io/worldweaver/

详情
AI中文摘要

多智能体交互世界模型不仅要生成一致的观测结果,还要维护跨智能体持久且跨视图演变的世界状态。现有的自回归视频扩散管道将观测历史作为条件上下文传递,这使得在多智能体和多视图设置中难以维护共享状态。我们提出了WorldWeaver(W^2),一种流式多智能体视频扩散模型,它通过跨智能体世界状态寄存器增强了展开过程:可学习的令牌存储共享世界信息、跟踪单个智能体状态,并在每个生成块后动态更新。我们用跨越单个智能体状态、包括鸟瞰图的全局状态视图和场景文本的监督信号来为这些寄存器提供基础。我们进一步用混合变压器设计改进了架构,该设计对世界状态建模和视觉帧建模使用单独的权重。在双智能体我的世界视频生成中的广泛实验表明,显式的世界状态建模提高了逻辑一致性和生成质量。

英文摘要

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. We ground these registers with supervision signals spanning individual agent status, global state views including bird's-eye views, and scene text. We further improve the architecture with a Mixture-of-Transformers design that uses separate weights for world state modeling and visual frame modeling. Extensive experiments in two-agent Minecraft video generation show that explicit world-state modeling improves logical consistency and generation quality.

URL PDF HTML 收藏
2607.21592 2026-07-24 cs.CV 新提交

Unified Video Dense Prediction from Disjoint Data

从不相交数据进行统一视频密集预测

Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee

机构 * Adobe Research(Adobe 研究院) Cornell University(康奈尔大学)

AI总结 研究旨在解决现有任务特定注释分散问题,提出统一视频模型UniD,从不相交特定领域数据集联合预测八个密集场景属性。通过简单蒸馏步骤,利用预训练扩散模型视觉先验弥合领域差距,性能优于特定任务专家和多任务基线,泛化能力强。

Comments ECCV 2026

详情
AI中文摘要

场景理解需要同时预测几何、外观和语义。然而,现有的特定任务注释分散在不兼容的特定领域数据集中。当前的统一系统通过将训练限制在完全共同注释的数据上,或通过产生伪标签的巨大计算成本来规避这一问题。为了缓解这一问题,我们引入了UniD,一个统一的视频模型,它联合预测八个密集场景属性——深度、表面法线、语义分割、边界、人体部位、反照率、阴影和材质——所有这些都从不相交的特定领域数据集中学习。我们提出了一个简单而有效的蒸馏步骤,其中每个任务的专家通过轻量级任务投影仪监督统一的主干,消除了对注释重叠或伪标签的需求。我们的关键见解是,预训练扩散模型的强大视觉先验足以弥合不相交训练源引入的领域差距,从而能够对训练期间从未见过的场景任务组合进行稳健泛化。UniD在与特定任务专家和多任务基线的比较中取得了有竞争力的性能,对分布外场景具有强大的泛化能力,并增强了时间和跨任务的一致性。代码和视频结果可在这个https URL上获取。

英文摘要

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labeling. Our key insight is that the strong visual priors of a pretrained diffusion model are sufficient to bridge the domain gaps introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training. UniD achieves competitive performance against per-task specialists and multi-task baselines, with strong generalization to out-of-distribution scenarios and enhanced temporal and cross-task consistency. Code and video results are available at https://unid-video.github.io/.

URL PDF HTML 收藏
2607.21591 2026-07-24 cs.CV 新提交

Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning

通过渐进式种子剪枝实现扩散模型的推理时间缩放

Rogerio Guimaraes, Pietro Perona

机构 * California Institute of Technology(加州理工学院)

AI总结 研究扩散模型推理时间缩放问题,提出渐进式种子剪枝方法,通过早期评估多种子并剪枝,有效利用固定计算预算,在扩散和流匹配主干上比其他基线方法在奖励引导选择及提示对齐评估上表现更优。

Comments Project page: https://www.vision.caltech.edu/psp. Code: https://github.com/rogerioagjr/psp

详情
AI中文摘要

扩散模型和流匹配模型在条件图像生成中占主导地位,但其推理时间缩放远不如自回归语言模型成熟。由于最终质量对初始噪声种子高度敏感,许多方法在种子搜索或重采样上花费额外计算。我们表明,放宽这一约束能实现未充分探索的推理时间缩放轴:通过早期加载探索、评估多个种子并积极剪枝,可更有效地使用固定计算预算。渐进式种子剪枝(PSP)对中间去噪估计进行评分并逐步缩小候选集,在保持模型评估总数固定的情况下,仅对有希望的轨迹进行完全去噪。在扩散和流匹配主干上,PSP始终改进奖励引导选择,在匹配计算时比最佳N、重要性采样和树搜索基线获得更高的GenEval分数(自动化)和更好的人类提示对齐评估。

英文摘要

Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the initial noise seed, many approaches spend extra compute on seed search or resampling under a black-box reward, but typically maintaining a constant memory footprint throughout inference. We show that relaxing this constraint enables an underexplored inference-time scaling axis: by front-loading exploration, evaluating many seeds early, and pruning aggressively, we can use a fixed compute budget more effectively. \emph{Progressive Seed Pruning} (\PSP) scores intermediate denoised estimates and progressively narrows the candidate set so that only promising trajectories are fully denoised, while keeping the total number of model evaluations fixed. Across diffusion and flow-matching backbones, \PSP \ consistently improves reward-guided selection and achieves higher GenEval scores (automated) and better human evaluation on prompt-alignment than best-of-$N$, importance-sampling, and tree-search baselines at matched compute. Project page: https://www.vision.caltech.edu/psp. Code: https://github.com/rogerioagjr/psp.

URL PDF HTML 收藏
2607.21588 2026-07-24 cs.RO 新提交

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

AXIS:用于可扩展机器人操作的可增长社区驱动数据引擎

Mengfei Zhao, Dihong Huang, Yikai Tang, Peihao Li, Mingxuan Yan, Ruiqi Zhuang, Yanjia Huang, Jie Wang, Hai Zhai, Tony Zhou, Rui Zhang, Zhexi Luo, Yuchen Huang, Jianfei Yang, Jiachen Li

机构 * Axis Robotics(轴机器人公司) University of California, Berkeley(加州大学伯克利分校) Georgia Institute of Technology(佐治亚理工学院) Texas A&M University(德州农工大学) Johns Hopkins University(约翰·霍普金斯大学) University of Pennsylvania(宾夕法尼亚大学) University of Michigan(密歇根大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

AI总结 研究针对机器人操作策略学习中数据管道难扩展的问题,提出AXIS这一可增长社区驱动数据引擎,它能收集、处理数据并组织成任务快照,通过实验表明其能提升模型性能且随数据量增加有一致扩展性。

Comments Project Website: https://axisaiorg.github.io/AXIS-V1/

详情
AI中文摘要

学习有效的机器人操作策略需要多样、高质量的演示,但现有数据管道因依赖专业硬件、集中式操作员或固定任务套件而难以扩展。我们提出了AXIS,一个用于可扩展机器人学习的可增长社区驱动数据引擎和基准。它支持基于浏览器的远程操作以收集大规模演示,自动生成并验证新操作任务,通过自动成功检查、质量过滤等将社区收集的演示转化为可训练数据。AXIS数据集目前包含207个不同任务和50K+轨迹,还将数据组织成任务快照并通过系统的留出协议评估策略。我们在统一的AXIS评估套件下比较视觉-语言-动作(VLA)策略并分析不同数据量下的扩展行为。在AXIS上持续预训练大幅提高了π0.5的总体成功率,比在RoboCasa365上预训练的模型性能优37.3%,且随着数据量增加呈现一致的扩展性。

英文摘要

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalable robot learning, which enables browser-based teleoperation for large-scale demonstration collection, automatically generates and validates new manipulation tasks, and transforms community-collected demonstrations into training-ready data through automated success checking, quality filtering, trajectory smoothing, and visual and physics-based augmentation. The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories. Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-out protocol. We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyze scaling behavior across different data volumes. Continual pretraining on AXIS substantially improves the overall success rate of $π_{0.5}$ by 5.8%, outperforms the model pretrained on RoboCasa365 by 37.3%, and exhibits consistent scaling with increasing data volume, with the largest gains observed under layout, sensor-noise, and camera perturbations.

URL PDF HTML 收藏
2607.21582 2026-07-24 cs.RO cs.CV 新提交

Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

战略扩展:通过偏差感知评估和数据收集学习机器人操作中的组合泛化

Yu Qi, Zhang Ye, Xinyi Xu, Yuxuan Lu, Amitoj Sandhu, Boce Hu, Haojie Huang, Jonathan Tremblay, Lawson L. S. Wong

机构 * Northeastern University(东北大学) NVIDIA(英伟达)

AI总结 研究机器人操作中组合泛化问题,通过引入诊断框架量化指令因素偏差,提出偏差感知数据收集策略,能定位预训练策略捷径问题,实现更高效可泛化的策略学习。

详情
AI中文摘要

组合泛化对机器人遵循各种指令至关重要。然而,预训练策略会走捷径,依赖显著线索而非基于语言。我们引入一个诊断框架,将这种失败定位到各个“指令因素”,如颜色、动词、物体、大小和空间属性等可重复使用的语义组件。该框架形式化了指令因素偏差,即微调策略过度依赖主导因素走捷径的倾向,并通过两个指标量化:因素主导率(FDR)捕捉因素间的成对偏差,因素主导层次结构(FDH)将其汇总为全局排名。对六个基础策略的评估揭示了大致一致的排序,即颜色≥物体≥空间≥动词≥大小,颜色占主导,动词和大小最缺乏基础。我们进一步表明这种诊断是可行的:一种偏差感知数据收集策略,将固定预算重新分配到缺乏基础的因素上,在模拟和真实机器人上使用一半演示时优于基线,从而实现更具样本效率和可泛化的策略学习。

英文摘要

Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual \textit{instruction factors}, \textit{e.g.,} reusable semantic components such as color, verb, object, size, and spatial attribute. Our framework formalizes instruction factor bias, the tendency of fine-tuned policies to over-rely on dominant factors as shortcuts, and quantifies it through two metrics: Factor Dominance Rate (FDR), capturing pairwise bias between factors, and Factor Dominance Hierarchy (FDH), aggregating these into a global ranking. Evaluation on six foundation policies reveals broadly consistent ordering, \textit{i.e.}, color $\geq$ object $\geq$ spatial $\geq$ verb $\geq$ size, with color dominant, and verb and size most under-grounded. We further show the diagnosis is actionable: a bias-aware data collection strategy that reallocates a fixed budget toward under-grounded factors outperforms baselines in simulation and on a real robot using half the demonstrations, thereby enabling more sample-efficient and generalizable policy learning.

URL PDF HTML 收藏
2607.21577 2026-07-24 cs.CV cs.AI cs.LG eess.IV 新提交

Synthetic data generation framework for quality control automation in gravure printing

凹版印刷质量控制自动化的合成数据生成框架

Korota Arsène Coulibaly, Mohamed Hamlich, Khalid Hmali, Andrea Trombin

机构 * univh2c(滨海大学)

AI总结 针对凹版印刷质量控制中人工检测的不足,提出合成数据生成框架,自动生成特定印刷缺陷图像及标注,用其训练模型在实际工业测试样本上达80.9%的mAP,提供零成本、快速部署的缺陷检测自动化方案。

Comments 27 pages, 15 figures. To be submitted to Journal of Engineering Research (Elsevier). Certain TeX commands are supported

详情
AI中文摘要

印刷中的质量控制,尤其是轮转凹版印刷,仍依赖缓慢、昂贵且主观的人工检查。自动表面缺陷检测对维持凹版印刷的高质量标准至关重要。深度学习模型为自动化带来希望,但训练如YOLO或视觉Transformer等强大的深度学习模型因现实工业缺陷图像极度稀缺而受阻。本文引入专为凹版印刷质量控制定制的新型合成数据生成框架。该框架自动生成特定印刷缺陷(褶皱、条纹、套准误差等)的高保真图像,并输出相应边界框和注释。为验证框架,生成7533张图像的合成数据集并用于训练最先进的目标检测模型RFDETR。实验结果表明,在我们的合成数据上训练的模型在实际工业测试样本上实现了80.9%的平均精度均值(mAP)。该框架为印刷生产线中的缺陷检测自动化提供了零成本、快速部署的解决方案,无需大量人工数据收集。

英文摘要

Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface defect detection is critical for maintaining high-quality standards in rotogravure printing. Deep learning models give prospects for automation. However, training robust deep learning models, such as YOLO or Vision Transformers, is heavily hindered by the extreme scarcity of real-world industrial defects images. To overcome this limitation, this paper introduces a novel synthetic data generation framework tailored for rotogravure printing quality control. The proposed pipeline automatically generates high-fidelity images of specific printing defects (creases, streaks, misregistration, etc.) and outputs corresponding bounding boxes and annotations. To validate the framework, a synthetic dataset of 7533 images was generated and used to train the state-of-the-art object-detection model RFDETR. Experimental results demonstrate that the model trained on our synthetic data achieves a Mean Average Precision (mAP) of 80.9\% on real industrial testing samples. This framework provides a zero-cost, rapid-deployment solution for automating defect inspection in printing lines without requiring massive manual data collection.

URL PDF HTML 收藏
2607.21576 2026-07-24 cs.CV 新提交

Self-Supervised Learning of Structured Dynamics from Videos

从视频中进行结构化动力学的自监督学习

Lukas Knobel, Andrew Zisserman, Yuki M. Asano

机构 * Fundamental AI Lab, UTN(基础人工智能实验室,UTN) VGG, University of Oxford(VGG,牛津大学)

AI总结 研究能否从预训练图像视觉Transformer的冻结特征中恢复结构化运动表示,提出结构化动力学模型(SDM),通过未来特征预测分离动力学来源,结合自监督与弱监督训练,在新评估套件上表现出色,证明预训练图像模型可用于结构化视频动力学表示。

Comments preprint, Project page: https://lukasknobel.github.io/projects/StructuredDynamics

详情
AI中文摘要

理解视频中的运动是视觉学习的一项基本挑战,因为帧间变化包含相机运动和物体运动这两种动力学来源。在表示学习中,这种分解尚未得到充分探索,部分原因是这些因素在自然视频中紧密耦合且难以单独监督。然而,恢复这种分解对于学习将有意义的物体动力学与相机引起的变化分离的鲁棒运动表示很重要。我们研究是否可以从预训练图像视觉Transformer的冻结特征中恢复这种结构化运动表示。我们提出了结构化动力学模型(SDM),它通过未来特征预测明确地将时间变化的主要来源与残余动力学分开,而不是用单个纠缠的潜在特征或无结构的空间密集过渡令牌来表示视频变化。训练将对真实视频的自监督学习与对合成Kubric数据的场景动力学弱监督相结合。我们在ProbeMotion上评估SDM,这是一个新的评估套件,涵盖具有相机运动、物体运动和组合动力学的合成和真实视频。SDM在使用全局CLS或平均池化特征方面优于主干基线,并且在几个探针上与诸如VGGT等强监督表示相比具有优势,尽管使用的监督要弱得多。这些结果表明,预训练的图像模型可以很容易地重新用于结构化视频动力学表示,为学习和分析潜在视频动力学提供有用的归纳偏差。

英文摘要

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.

URL PDF HTML 收藏
2607.21574 2026-07-24 cs.CL 新提交

Surprisal Theory is Tautological (without Rational Grounding)

惊奇理论是同义反复的(缺乏理性基础)

Ryan Cotterell

机构 * ETH Zürich(苏黎世联邦理工学院)

AI总结 研究惊奇理论,指出其在无额外约束时是同义反复,任何难度模式都与某语言模型一致,无法证伪。长期被隐含假设掩盖,近期实证削弱该假设。结论是打破同义反复需理性主义干预,相关语言模型应源自非经验驱动模型。

Comments Under "Review" at ARR

详情
AI中文摘要

惊奇理论认为,语境中语言单元的人类处理难度是其在某种语言模型下惊奇度的仿射函数。本文认为该主张在没有进一步约束时是同义反复:对于语境中单元的任何非负难度度量,在温和技术条件下,存在一个语言模型,其惊奇度是该难度度量的仿射函数。所以,由于任何难度模式都与某个语言模型一致,若无对语言模型的额外约束,惊奇理论无法做出可证伪的预测。这一同义反复长期被心理语言学工作中隐含的假设掩盖,即相关语言模型是生成训练语料库的分布,所以改善语料库拟合能改进对人类行为的预测。近期实证工作削弱了这一假设,表明更好的语料库模型对处理难度的预测可能更差。本文结论是,打破同义反复需要理性主义干预,即相关语言模型必须源自基于理解者的非经验驱动模型,比如基于记忆约束或处理目标,且不依赖于惊奇理论旨在解释的行为数据。

英文摘要

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore, because any pattern of difficulty is consistent with some language model, without an additional constraint on the language model, surprisal theory makes no falsifiable predictions. The tautology was long obscured by an assumption implicit in two decades of psycholinguistic work---that the relevant language model is the distribution that generated the training corpus, so that improving corpus fit improves predictions of human behavior. Recent empirical work has undermined this assumption, demonstrating that better corpus models can be worse predictors of processing difficulty. I conclude that breaking the tautology requires a rationalist intervention, i.e., the relevant language model must be derived from a non-empirically motivated model of the comprehender, which could be based on, for instance, memory constraints or processing goals, and that, thus, does not depend on the behavioral data surprisal theory is meant to explain.

URL PDF HTML 收藏
2607.21573 2026-07-24 cs.LG cs.AI 新提交

Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity

超越充分性:具有反事实必要性的时间序列解释

Hongnan Ma, Yiwei Shi, Mengyue Yang, Weiru Liu

机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院) School of Engineering Mathematics and Technology, University of Bristol(布里斯托大学工程数学与技术学院)

AI总结 研究时间序列分类器解释问题,提出必要性感知框架TimePNS,受Pearl反事实必要性概念启发,采用两阶段设计,通过实验证明其能更准确识别决策关键子序列,改进充分性-必要性权衡。

详情
AI中文摘要

时间序列分类器的可靠解释应识别出不仅足以维持黑箱模型预测,而且对维持预测必不可少的子序列。然而,现有的面向充分性的方法可能会将高重要性赋予支持预测但对模型决策并非必不可少的虚假子序列。我们引入了TimePNS,这是一个用于时间序列解释的必要性感知框架。受Pearl反事实必要性概念的启发,TimePNS通过对时间因素进行干预并测量原始预测是否被破坏来评估其是否必要。该框架采用两阶段设计。第一阶段学习一个可识别的因果生成过程以及一个面向充分性的解释掩码。第二阶段对时间因素进行反事实干预以得出必要性信号,该信号监督一个时间门,通过抑制非必要成分并强调反事实必要成分来完善初始解释。在合成和真实世界时间序列基准上的实验表明,TimePNS能更准确地识别决策关键子序列,并始终优于强大基线改进充分性-必要性权衡。

英文摘要

Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficiency-oriented methods can assign high importance to spurious subsequences that support the prediction without being essential to the model's decision. We introduce \textbf{TimePNS}, a necessity-aware framework for time-series explanation. Inspired by Pearl's counterfactual notion of necessity, TimePNS assesses whether a temporal factor is necessary by intervening on it and measuring whether the original prediction is disrupted. The framework adopts a two-stage design. Stage I learns an identifiable causal generative process together with a sufficiency-oriented explanation mask. Stage II performs counterfactual interventions on temporal factors to derive necessity signals, which supervise a temporal gate that refines the initial explanation by suppressing non-essential components and emphasizing counterfactually necessary ones. Experiments on synthetic and real-world time-series benchmarks show that TimePNS more accurately identifies decision-critical subsequences and consistently improves sufficiency-necessity trade-offs over strong baselines.

URL PDF HTML 收藏
2607.21571 2026-07-24 cs.RO 新提交

Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

超越情节性评估:序列式具身问答中的内存架构瓶颈

Zikui Cai, Kaushal Janga, Tan Dat Dao, Seungjae Lee, Shivin Dass, Mingyo Seo, Kaiyu Yue, Mintong Kang, Nandhu Pillai, Monte Hoover, Aadi Palnitkar, Ruchit Rawal, Ruijie Zheng, Bo Li, Yuke Zhu, Roberto Martín-Martín, Tom Goldstein, Furong Huang

机构 * University of Maryland, College Park(马里兰大学帕克分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 研究序列式具身问答中不同内存架构表现,发现仅保留现有记忆不足,短视情节数据训练的智能体有时间不匹配问题。强调结构化、基于空间记忆的必要性,实验表明其能打破准确性-效率权衡,对真实机器人连续智能操作至关重要。

Comments Accepted to IROS 2026

详情
AI中文摘要

具身问答(EQA)传统上是在情节性框架下进行评估的,即智能体独立解决每个任务并在情节之间重置内部状态。然而,现实世界中的机器人是持续运行的,必须积累、保留并选择性地重用从先前交互中获取的信息。尽管有此实际需求,但在EQA中支持序列记忆所需的架构机制仍未得到充分探索。在这项工作中,我们研究了在对EQA智能体进行序列评估时,即在同一场景中回答多个问题且记忆在查询之间传递时,不同内存架构的表现。我们发现仅仅保留现有记忆往往是不够的。仅保留可遍历性信息(如二维占用地图)的智能体,记住了机器人探索过的位置,但没有记住后续问题所需的视觉语义证据。在短视情节数据上训练的智能体面临不同的挑战:当面对连续的多查询历史时,它们继承的上下文存在严重的时间不匹配,而不是形成可重用的场景表示。为了克服这一架构瓶颈,我们强调了结构化的、基于空间的记忆的必要性:将持久视觉观察映射到度量三维几何上的架构,在连贯的场景表示中保留视觉语义证据。在模拟环境中的大量实验表明,这种形式的记忆打破了序列设置中的准确性-效率权衡,同时实现了更高的答案准确性和更低的导航成本。我们还在真实世界的移动机器人上验证了这些发现,证明基于空间的视觉记忆对于在物理环境中实现连续、智能操作至关重要。

英文摘要

Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continuously and must accumulate, retain, and selectively reuse information acquired from prior interactions. Despite this practical requirement, the architectural mechanisms needed to support sequential memory in EQA remain underexplored. In this work, we investigate how different memory architectures behave when EQA agents are evaluated sequentially, with multiple questions answered in the same scene while memory is carried forward across queries. We find that simply preserving existing memory is often insufficient. Agents that retain only traversability information, such as 2D occupancy maps, remember where the robot has explored but not the visual-semantic evidence needed for later questions. Agents trained on short-horizon episodic data face a different challenge: when exposed to continuous, multi-query histories, their inherited context suffers from severe temporal mismatch, rather than forming a reusable scene representation. To overcome this architectural bottleneck, we highlight the necessity of structured, spatially grounded memory: architectures that map persistent visual observations onto metric 3D geometry preserve visual-semantic evidence in a coherent scene representation. Extensive experiments in simulated environments reveal that this form of memory breaks the accuracy-efficiency tradeoff in sequential settings, simultaneously achieving higher answer accuracy and lower navigation costs. We further validate these findings on a real-world mobile robot, demonstrating that spatially grounded visual memory is critical for enabling continuous, intelligent operation in physical environments.

URL PDF HTML 收藏
2607.21570 2026-07-24 cs.CL cs.HC 新提交

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

MedGame:由大语言模型赋能的用于医学教育的故事化游戏化

Qian Wu, Xinrong Zhou, Zizhan Ma, Kai Chen, Zheyao Gao, Xun Lin, Hongqiu Wu, Longfei Gou, Yixiao Liu, Ann Sin Nga Lau, Qi Dou

机构 * CUHK(香港中文大学) Southern Medical University(南方医科大学) Peking University(北京大学) Tencent(腾讯)

AI总结 研究旨在利用大语言模型助力医学教育,提出MedGame框架,通过双引擎设计将临床病例转化为故事化游戏,构建MedGame Bench进行评估,实验显示特定任务微调提升开源模型表现,学生研究表明其更具吸引力和实用性。

Comments Work in Progress; an explorational design and study on AI+Education+Game

详情
AI中文摘要

大语言模型在医学教育中展现出潜力,但现有多数系统专注于局部交互,而非将整个临床病例组织成以决策为中心的学习轨迹。我们引入MedGame,一个将静态临床病例转化为结构化、可执行的故事化游戏的框架。MedGame采用双引擎设计,医学叙事设计师合成基于病例的临床故事情节,故事导演将其转化为依赖感知的多模态编排计划。我们构建了MedGame Bench作为医学叙事生成和故事指导的基准和评估协议。实验表明特定任务微调显著提升了开源大语言模型在MedGame Bench上的表现并缩小与商业模型的差距。一项学生试点研究进一步表明学习者认为MedGame比纯文本替代方案更具吸引力和实用性。

英文摘要

Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or single-turn feedback, rather than organizing an entire clinical case into a decision-centered learning trajectory. We introduce \textit{MedGame}, a framework that transforms static clinical cases into structured, executable storytelling games. MedGame uses a dual-engine design: a Medical Narrative Designer synthesizes case-grounded clinical storylines with states and decision nodes, while a Story Director converts them into dependency-aware multimodal orchestration plans rendered by our released interactive platform. We construct MedGame Bench, a 5,000-case benchmark and evaluation protocol for Medical Narrative Generation and Story Direction. Experiments show that task-specific fine-tuning substantially improves open-source LLMs on MedGame Bench and narrows the gap with commercial models. A pilot student study further shows that learners perceive MedGame as more engaging and useful than text-only alternatives.

URL PDF HTML 收藏
2607.21562 2026-07-24 cs.CV cs.GR 新提交

Scene Parameter Saliency via Differentiable Light Transport

通过可微光线传输实现场景参数显著性

Linas Beresna, Eugene Fiume

机构 * Simon Fraser University(西蒙弗雷泽大学)

AI总结 研究通过可微光线传输实现场景参数显著性,计算针对不同目标的度量显著性图,发现同一场景不同度量下显著性排名差异大,显著性图因度量而异,结果表明可微渲染器的导数图像对场景理解有重要价值。

Comments 13 pages, 5 figures

详情
AI中文摘要

基于梯度的显著性方法揭示哪些输入特征对神经网络输出影响最大,是模型可解释性的标准工具。我们发现,常用于参数优化的可微渲染器会产生类似的显著性形式:给定在渲染图像上评估的任何标量度量,单次反向模式微分传递会产生每个参数的梯度,以识别哪些场景元素对该度量影响最大。我们将这些梯度场称为度量显著性图。与通过学习权重传播归因的神经显著性不同,度量显著性通过图像形成过程本身传播,包括多次反射光线传输,捕捉到难以通过人工检查发现的参数依赖性。我们针对不同性质的目标计算度量显著性图:心理视觉眩光指数、平均场景亮度和神经感知分数。对于同一场景,不同度量的显著性排名差异很大,对一个目标起主导作用的参数对另一个目标可忽略不计。显著性图特定于度量,而非场景的固有属性。我们的结果表明,可微渲染器产生的导数图像对于场景理解与它们设计生成的原始图像一样具有信息价值。

英文摘要

Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability. We observe that differentiable renderers, which are conventionally used for parameter optimisation, produce an analogous form of saliency: given any scalar metric evaluated on a rendered image, a single reverse-mode differentiation pass yields per-parameter gradients that identify which scene elements most influence the metric. We call these gradient fields metric saliency maps. Unlike neural saliency, which propagates attribution through learned weights, metric saliency propagates through the image formation process itself, including multi-bounce light transport, capturing parameter dependencies that are semi-opaque to manual inspection. We compute metric saliency maps for qualitatively different objectives: psychovisual glare indices, mean scene luminance, and neural perceptual scores. The saliency rankings differ substantially across metrics for the same scene, with parameters that dominate one objective being negligible for another. The saliency map is specific to the metric, not an intrinsic property of the scene. Our results suggest that differentiable renderers produce derivative images that are as informative for scene understanding as the primal images they were designed to generate.

URL PDF HTML 收藏
2607.21559 2026-07-24 cs.AI cs.CE cs.ET stat.AP stat.ML 新提交

Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana

基于无监督共识的加纳时空疟疾发病率异常检测

T. Ansah-Narh, Y. Asare Afrane

机构 * Ghana Space Science and Technology Institute, Ghana Atomic Energy Commission(加纳空间科学与技术研究所,加纳原子能委员会) Department of Medical Microbiology, University of Ghana Medical School(加纳大学医学院医学微生物学系)

AI总结 利用基于共识的异常检测框架分析加纳2014 - 2023年月度疟疾监测数据,发现时空异常特征及异常负担与频率的空间差异,能区分高发区与异常传播区,助力加强疟疾监测与防控。

Comments 32, 15 figures, under review at spatial and spatio-temporal epidemiology

详情
AI中文摘要

一个基于共识的异常检测框架应用于加纳2014 - 2023年的月度疟疾监测数据,以识别非典型传播模式。异常在时空上具有高度结构化特征。阿散蒂和北部地区异常情况最为频繁,塔马利、库马西和阿克拉存在持续热点。关键发现是异常负担(异常时期累计病例数)和异常频率(异常行为持续性)的空间差异。该框架通过区分疟疾高发区和传播异常区,可加强监测、优化调查并支持针对性控制策略。

英文摘要

A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transmission patterns. Anomalies were highly structured in space and time. Ashanti and Northern Regions accounted for most recurrent anomalies, with persistent hotspots at Tamale, Kumasi, and Accra. A key finding was the spatial distinction between anomaly burden (cumulative cases during anomalous periods) and anomaly frequency (persistence of unusual behaviour). Tamale had the highest burden during anomalies, whereas the highest anomaly rates clustered in Ashanti districts, showing that high-burden areas are not necessarily those with the most frequent anomalous transmission. Anomalous months formed a statistically distinct group, with much higher case counts (Cohen's $d = 3.252$) and large seasonal deviations ($d > 1.2$) compared with normal months. Malaria burden alone provides an incomplete picture of transmission dynamics. By distinguishing where malaria is most prevalent from where transmission behaves most unusually, this framework can strengthen surveillance, prioritise investigations, and support targeted control strategies.

URL PDF HTML 收藏
2607.21557 2026-07-24 cs.AI cs.CL 新提交

OpenForgeRL: Train Harness-native Agents in Any Environment

OpenForgeRL:在任何环境中训练原生利用工具的智能体

Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao

机构 * Columbia University(哥伦比亚大学) Dartmouth College(达特茅斯学院) Microsoft Research(微软研究院)

AI总结 研究针对现代AI智能体依赖复杂推理工具难端到端训练的问题,提出OpenForgeRL框架,通过轻量级代理和Kubernetes编排器,在多环境下对基于工具的智能体端到端训练,验证了框架效果并分析了工具选择和RL对智能体行为的影响。

详情
AI中文摘要

现代人工智能智能体依赖复杂的推理工具(如Claude Code、Codex和OpenClaw)来驱动多轮推理、工具使用和访问外部系统。这些强大但复杂的工具使得智能体难以通过开放基础设施进行端到端训练,因为其SFT/RL堆栈无法原生表达有状态的多进程工具推理。为此,我们提出了OpenForgeRL,一个用于在各种环境中对基于工具的智能体进行端到端训练的开源框架。OpenForgeRL通过一个轻量级代理实现这一目标,该代理在记录工具模型调用作为标准RL代码库(如veRL)的训练数据时为其提供服务,以及一个Kubernetes编排器,它在自己的远程容器中运行每次展开,共同实现了在任何环境中对任何工具进行大规模训练。通过解耦训练和推理,OpenForgeRL允许研究人员在智能体所部署的实际工具和环境中轻松地训练、研究和改进智能体。我们在各种复杂的工具和环境中验证了我们的框架,涵盖基于工具/爪子的智能体以及多模态GUI浏览器和计算机使用智能体。仅使用数百到数千个任务,OpenForgeClaw在ClawEval上达到31.7 pass^3和55.9 pass@3,在QwenClawBench上达到33.7。OpenForgeGUI在OSWorld-Verified上达到37.7,在Online-Mind2Web上达到63.0,在WebVoyager上达到72.3。两者在几乎所有基准测试中都优于类似规模的开放基线,并且在GUI设置中与几倍大的模型相匹配或超越。除了基准测试,我们分析了工具选择(如ZeroClaw、OpenClaw、Codex)和强化学习如何塑造智能体行为。我们发现一些工具比其他工具更难学习,并且强化学习提高了智能体的可靠性,如自我验证、工具覆盖和完成多步计划,尽管诸如错误恢复等关键能力仍然较弱。

英文摘要

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Both outperform open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger. Beyond benchmarks, we analyze how harness choice (e.g., ZeroClaw, OpenClaw, Codex) and RL shape agent behavior. We find that some harnesses are substantially harder to learn than others, and that RL improves agentic reliability, such as self-verification, tool coverage, and completing multi-step plans, though critical abilities such as error recovery remain weak.

URL PDF HTML 收藏
2607.21556 2026-07-24 cs.CV cs.AI 新提交

Visual Contrastive Self-Distillation

视觉对比自蒸馏

Yijun Liang, Yunjie Tian, Yijiang Li, Yuqi Jia, Furong Huang, Tianyi Zhou, Di Fu

机构 * University of Maryland, College Park(马里兰大学帕克分校) University of California, San Diego(加利福尼亚大学圣地亚哥分校) Duke University(杜克大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

AI总结 研究提出视觉对比自蒸馏(VCSD),将图像内容去除转换为策略内自蒸馏信号,通过对比突出相关候选者,锐化教师原始图像分布并蒸馏到学生中,在多个模型上优于匹配的OPSD,且无需外部教师等额外条件。

Comments 15 pages

详情
AI中文摘要

策略内自蒸馏(OPSD)很有前景,它去除了策略内蒸馏(OPD)所需的外部教师,但仍需要教师和学生之间的不对称信息。现有方法通过特权答案或视觉证据来创造这种不对称。本文提出是否可以消除这两者,产生一种仅由输入条件驱动的更简单的OPSD形式。为此提出视觉对比自蒸馏(VCSD),将图像内容去除转换为策略内自蒸馏信号。在每个学生生成的响应前缀处,指数移动平均(EMA)教师在相同提示和前缀下产生两个下一个token分布,其token-wise对数概率差异突出由实例级视觉内容特别增加可能性的候选者。利用此对比锐化教师在合理支持范围内的原始图像分布,并将结果全分布目标蒸馏到学生中。使用ViRL39K数据集,VCSD在Qwen3-VL和Qwen3.5模型上始终优于匹配的OPSD。例如,在Qwen3-VL上,2B时七基准聚合从62.27%提高到67.04%,4B时从71.30%提高到73.16%,8B时从72.51%提高到76.26%。此外,VCSD不需要外部教师、特权答案、视觉证据信号、推理痕迹或额外推理时间成本。

英文摘要

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from $62.27\% \rightarrow 67.04\%$ at 2B, $71.30\% \rightarrow 73.16\%$ at 4B, and $72.51\% \rightarrow 76.26\%$ at 8B. Furthermore, VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.

URL PDF HTML 收藏
2607.21553 2026-07-24 cs.CV 新提交

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

SANA-Video 2.0:具有注意力残差的混合线性注意力用于高效视频生成

Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie

机构 * NVIDIA(英伟达)

AI总结 SANA-Video 2.0是一种混合视频扩散Transformer,采用混合线性-softmax注意力和块注意力残差,在单个GPU上生成高质量视频,成本大幅降低,实现可扩展长时、高分辨率视频生成,性能与大模型竞争且速度更快。

Comments 13 pages, 9 figures, 5 tables

详情
AI中文摘要

我们介绍了SANA-Video 2.0,这是一种在统一架构下实例化的5B和14B规模的混合视频扩散Transformer。旨在在单个GPU上生成高达720p的高质量视频,SANA-Video 2.0在质量上与全softmax视频DiTs相匹配,同时保留线性注意力良好的长序列缩放特性。混合线性-softmax注意力以3:1的比例将门控线性注意力与周期性门控-softmax锚点相结合,避免了二次注意力,恢复了纯线性注意力所缺乏的满秩令牌交互。通过块注意力残差将完整的块摘要路由到后续线性层,实现锚点特征重用并将深层有效秩提高约12%。通过从头开始训练,SANA-Video 2.0直接学习完整的混合模型而非线性化预训练模型,通过降低分辨率的代理研究确定25%的softmax为最佳质量-效率权衡。在40步采样下,SANA-Video 2.0在单个H100上以480p在13.2秒内实现VBench分数84.30,在延迟仅为一小部分的情况下与大得多的softmax视频DiTs竞争。其编译的DiT前向传递在720p/60s时比匹配的全softmax基线快3.2倍,随着视频持续时间差距扩大。全栈Sol-Engine优化进一步加速了这个硬件友好的主干,使5B管道在720p/5s时达到13.06秒,比Wan 2.2-A14B在一个H100上快120倍。总体而言,我们的混合设计以大幅降低的成本恢复了softmax级别的表现力,实现了可扩展的长时、高分辨率视频生成。

英文摘要

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.

URL PDF HTML 收藏
2607.21552 2026-07-24 cs.AI cs.LG 新提交

MIRROR: Learning from the Other View for Multi-Modal Reasoning

MIRROR:从其他视角学习以进行多模态推理

Wen Ye, Yuxiao Qu, Aviral Kumar, Xuezhe Ma

机构 * University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究视觉语言模型在几何问题多模态推理上的不足,构建ODA-Data数据集,提出模态感知互惠推理优化(MIRROR)方法,通过强化学习自我监督,提升多模态推理能力,在相关基准测试中表现更优。

详情
AI中文摘要

与具有强大推理能力的大语言模型不同,视觉语言模型在视觉推理方面存在困难,即使是在有等效文本、图表和图表+文本视图的几何问题上。不同视图会引发不同行为,标准多模态后训练未充分利用这些互补推理路径和失败模式。为此构建了ODA-Data数据集,开发了模态感知互惠推理优化(MIRROR)方法,通过自我监督改进多模态推理。在几何问题推理基准测试中,MIRROR优于标准强化学习,跨模态行为更准确、一致。

英文摘要

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To study and exploit this phenomenon, we construct ODA-Data, a high-quality paired multimodal geometry dataset with text-dominant, image-dominant, and combined image+text views of the same problems, together with splits for training and evaluating modality-dependent reasoning behaviors. We then develop Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision. For each problem, MIRROR evaluates the model under all views, selects the best-performing view as a teacher, and trains other views with a reverse-KL objective towards the teacher. Across reasoning benchmarks that evaluate on geometry problems, MIRROR improves over standard RL and yields more accurate and consistent behavior across modalities

URL PDF HTML 收藏
2607.21550 2026-07-24 cs.LG 新提交

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

X³-OPD:通过策略对齐将推理能力提炼到大型音频-语言模型中

Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin

机构 * Tencent Hunyuan(腾讯混元) Zhejiang University(浙江大学)

AI总结 针对大型音频-语言模型逻辑推理能力不足,提出X³-OPD跨模态策略蒸馏框架,借助教师模型指导与构建的三层对称语料库训练学生模型,实验证明该方法能提升音频推理及思维链质量,还能保留模型域转移下的能力。

详情
AI中文摘要

虽然大型音频-语言模型在听觉感知方面取得了显著进展,但在深度逻辑推理方面仍落后于基于文本的大型语言模型,主要原因是高质量音频推理数据稀缺。为弥合这一差距,我们提出了X³-OPD,这是一个跨模态策略蒸馏框架,将推理能力从强大的文本教师模型转移到音频-语言学生模型。训练期间,学生模型根据自身声学感知生成推理轨迹,教师模型使用匹配的文本输入和验证答案提供令牌级指导。我们还构建了一个三层对称语料库。实验表明,X³-OPD显著提高了基于音频的推理和思维链质量,同时在很大程度上保留了模型在域转移下的现有能力。

英文摘要

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using matched textual inputs and verified answers. We further construct a three-tier symmetric corpus covering textual reasoning rendered into speech, audio-event reasoning grounded in complex acoustic scenes, and spoken-dialogue reasoning involving paralinguistic cues. This design extends cross-modal distillation beyond textually recoverable content to reasoning grounded in non-linguistic events, prosody, and conversational context. Experiments on MMSU, MMAU, BIG Bench Audio, and MMAR demonstrate that X$^3$-OPD substantially improves audio-grounded reasoning and chain-of-thought quality while largely preserving the model's existing capabilities under domain shift.

URL PDF HTML 收藏
2607.21547 2026-07-24 cs.AI cs.CL cs.ET cs.LG cs.MA 新提交

The Boundaries of Automation: A Theory of Persistent Human Participation

自动化的边界:持续人类参与理论

Fares Fourati, Hinrich Schütze, Eyke Hüllermeier, Iryna Gurevych

机构 * TU Darmstadt(达姆施塔特工业大学) LMU Munich(慕尼黑大学)

AI总结 探讨自动化概念极限,指出即便人工智能能力强,人类参与仍可能持续,原因包括技术互补、规范发展、目标涌现。认为人类与人工智能共同构建是目标通过参与涌现的活动的持久特征,对自动化极限及人工智能系统相关方面有重要意义。

详情
AI中文摘要

人工智能的快速发展加剧了对自动化的长期追求,即尽可能用算法取代人类参与,这种追求隐含着人类仅因当前人工智能系统能力不足而留在循环中的假设。本文挑战这一假设,探讨自动化的概念极限,认为即使在人工智能系统能力很强时,人类参与仍可能持续,原因有技术或互补性、规范性或发展性、目标涌现性。人类与人工智能共同构建不仅是对不完善人工智能的临时应对,而是目标通过参与涌现的活动的持久特征,这对自动化极限及未来人工智能系统设计、评估和伦理有重要意义。

英文摘要

The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the loop only because current AI systems are not yet sufficiently capable. This paper challenges that assumption. Rather than asking how far automation can extend, we ask where its conceptual limits lie and argue that human participation may persist even with highly capable AI systems for three distinct reasons. Technical or complementarity grounds arise when humans contribute capabilities or perspectives unavailable to AI. Normative or developmental grounds arise when participation itself is valuable for human agency or learning. Most importantly, emergence grounds arise from target emergence: in some activities, the target is not fully specified in advance but instead emerges through the interaction itself. In these cases, human participation is not merely a means of improving execution but is constitutive of the target being produced. Human--AI co-construction, understood as the joint production of outcomes by humans and AI systems, is therefore not simply a temporary response to imperfect AI, but a persistent feature of activities whose objectives emerge through participation. This perspective has important implications for the limits of automation and for the design, evaluation, and ethics of future AI systems.

URL PDF HTML 收藏
2607.21546 2026-07-24 cs.CV 新提交

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

UnDA:用于医学成像中跨模态知识转移的无配对域对齐

Rafsan Jany, Shadab Tanjeed Ahmad, Ahsan Bulbul, Tahsinul Islam, Md Azam Hossain, Abu Raihan Mostofa Kamal

机构 * Korea Institute of Oriental Medicine(韩国韩医学研究院) Islamic University of Technology(伊斯兰科技大学)

AI总结 针对医学成像中跨模态知识转移时获取配对数据难、处理模态差距及噪声传播问题,提出UnDA框架,通过与主干无关的对齐模块、不确定性加权最优传输及每个类别的ProtoNCE目标,实现无配对跨模态蒸馏,提升目标模态分割精度。

详情
AI中文摘要

基于多模态的方法在下游任务中通常优于单模态方法,因为不同模态提供互补信息,但获取配对临床数据在现实场景中仍是重大挑战。虽然跨模态知识蒸馏解决了这一问题,但现有方法在处理大模态差距和不确定源域预测的噪声传播方面存在困难。为克服这些挑战,我们提出了UnDA,这是一个用于无配对跨模态蒸馏的锚点引导框架。我们的方法引入了一个与主干无关的对齐模块,通过基于注意力的池化机制提取语义结构化的类令牌。为确保稳健的知识转移,我们提出了不确定性加权最优传输(UCT-OT),它基于预测置信度动态加权特征级对齐,有效抑制噪声监督。此外,每个类别的ProtoNCE目标维持稳定的原型记忆,以在无配对批次中强制全局可辨别性。在严格无配对设置下对代表性分割任务的评估表明,目标模态的准确性和边界精度持续提高,证明了无需配对数据集即可在异构数据源之间转移有意义的结构知识。

英文摘要

Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions. To overcome these challenges, we propose UnDA, an anchor-guided framework for unpaired cross-modal distillation. Our approach introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism. To ensure robust knowledge transfer, we propose Uncertainty-Weighted Optimal Transport (UCT-OT), which dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision. Furthermore, a per-class ProtoNCE objective maintains stable prototype memories to enforce global discriminability across unpaired batches. Evaluations on representative segmentation tasks under strictly unpaired settings show consistent improvements in accuracy and boundary precision in the target modality, demonstrating that meaningful structural knowledge can be transferred across heterogeneous data sources without paired datasets.

URL PDF HTML 收藏
2607.21545 2026-07-24 cs.CV 新提交

Towards Robust Iris Recognition Through Occlusion Identification and Conditional Diffusion-Based Reconstruction

通过遮挡识别和基于条件扩散的重建实现鲁棒虹膜识别

Kamrul Hasan, Mylene C. Q. Farias, Oleg V. Komogortsev

机构 * Texas State University(德克萨斯州立大学)

AI总结 针对虹膜纹理被遮挡时识别性能下降的问题,提出含遮挡类型识别、基于扩散重建及深度学习识别三个模块的框架,通过确定遮挡类别、重建受损区域并提取特征,提升了在CASIA-Iris-Thousand数据集上的虹膜识别性能。

Comments Accepted by IEEE International Joint Conference on Biometrics (IJCB) 2026

详情
AI中文摘要

虹膜识别是一种可靠的生物识别方法,利用虹膜独特且稳定的纹理识别个体。然而,当有判别力的虹膜纹理被眼睑、睫毛、镜面反射或其他采集伪像部分遮挡时,识别性能会下降。现有方法常直接对退化样本进行识别或仅依赖剩余可见区域,在大量纹理受损时可能不足。我们提出一个具有三个连续模块的遮挡感知虹膜识别框架:遮挡类型识别、基于扩散的重建和基于深度学习的识别。首先,基于残差二维卷积神经网络的网络确定虹膜图像是否未被遮挡或属于受控遮挡类别之一。其次,被遮挡图像、二进制掩码和预测的遮挡类型为去噪扩散概率模型提供条件以重建受损区域。最后,VGG19-HPMNet(一种具有水平金字塔映射的改进VGG19模型)提取有判别力的全局和局部虹膜特征进行识别。在受控合成遮挡协议下对CASIA-Iris-Thousand数据集的实验表明,该框架通过识别遮挡类型、重建掩码区域和重新评估恢复的虹膜样本提高了虹膜识别性能。

英文摘要

Iris recognition is a reliable biometric approach that identifies individuals using the distinctive and stable texture of the iris. However, recognition performance can degrade when discriminative iris texture is partially occluded by eyelids, eyelashes, specular reflections, or other acquisition artifacts. Existing approaches often perform recognition directly on degraded samples or rely only on the remaining visible iris region, which may be inadequate when substantial texture is corrupted. To address this limitation, we propose an occlusion-aware iris recognition framework with three sequential modules: occlusion-type identification, diffusion-based reconstruction, and deep-learning-based recognition. First, a residual 2D CNN-based network determines whether an iris image is non-occluded or belongs to one of the controlled occlusion categories. Second, the occluded image, binary mask, and predicted occlusion type condition a denoising diffusion probabilistic model to reconstruct the corrupted region. Finally, VGG19-HPMNet, a modified VGG19 model with horizontal pyramid mapping, extracts discriminative global and part-wise local iris features for recognition. Experiments on the CASIA-Iris-Thousand dataset under a controlled synthetic-occlusion protocol show that the proposed framework improves iris recognition performance by identifying the occlusion type, reconstructing masked regions, and re-evaluating the restored iris samples.

URL PDF HTML 收藏
2607.21542 2026-07-24 cs.LG stat.ML 新提交

Zero-Flow Two-Sample Tests

零流双样本检验

Yakun Wang, Leyang Wang, Song Liu, Taiji Suzuki

机构 * University of Bristol(布里斯托大学) RIKEN AIP(理化学研究所先进智能项目中心) University of Tokyo(东京大学)

AI总结 该论文提出零流双样本检验新方法,基于零流准则构建ZFD,通过分离见证学习与假设评估,利用灵活神经网络并保持统计校准,开发两种学习见证的方法,实验证明其在结构化分布变化检验中能力强且I型错误校准良好。

详情
AI中文摘要

我们提出一种用于判定两组样本是否来自同一分布的双样本检验新方法。该检验基于零流准则构建统计差异,即零流差异(ZFD)。我们证明了ZFD的有效性并提出实用检验程序,即零流双样本检验(ZF2ST)。关键是了解两个分布的样本如何局部不对齐,并将所得方向模式用作分布差异的证据。通过分离见证学习和假设评估,ZF2ST能使用灵活神经网络并保持有效统计校准。我们开发了基于回归和功率最大化的方法来学习见证。在合成和图像数据集上的实验表明,ZF2ST能在保持良好校准的I型错误的同时,对结构化分布变化实现强大的检验能力。

英文摘要

We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the same distribution. The test is built on a statistical discrepancy based on the zero-flow criterion, termed zero-flow discrepancy (ZFD). We prove the validity of ZFD and propose a practical testing procedure, termed the zero-flow two-sample test (ZF2ST). The key idea is to learn how samples from the two distributions are locally misaligned and use the resulting directional pattern as evidence of distributional difference. By separating witness learning from hypothesis evaluation, ZF2ST can use flexible neural networks while maintaining valid statistical calibration. We develop both regression-based and power-maximized approaches for learning the witness. Experiments on synthetic and image datasets demonstrate that ZF2ST can achieve strong testing power for structured distributional changes while maintaining well-calibrated type-I error.

URL PDF HTML 收藏
2607.21540 2026-07-24 cs.CL 新提交

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

DONDO:用于非洲语言的开放w2v-BERT语音识别基础模型

Paul Azunre

机构 * Khaya AI(Khaya人工智能公司)

AI总结 该研究提出DONDO一族非洲语言语音识别基础模型,基于w2v-BERT 2.0构建,含单语和多语模型。介绍了两步微调及语言调节机制,多语模型平均字错误率10 - 13%,缩小与单语差距,模型开源可自由微调,覆盖众多非洲语言使用者。

详情
AI中文摘要

我们展示了DONDO,这是一族基于w2v-BERT 2.0自监督语音编码器构建的、开放且遵循宽松许可的非洲语言自动语音识别基础模型。DONDO包含21个单语模型和5个多语模型,涵盖加纳、塞拉利昂、尼日利亚、塞内加尔、肯尼亚和津巴布韦的27种语言变体。模型主要在宗教文本的朗读语音上进行微调,这些文本为缺乏转录音频的语言提供了广泛、许可清晰且拼写一致的覆盖。我们描述了一种两步(对于一个族为三步)学习率退火微调过程,首先以高学习率调整共享多语模型,然后退火以恢复并在某些情况下超过强大的单语基线。我们还描述了一种轻量级语言调节机制,在推理时将独热语言标识作为前缀帧序列注入声学特征,使单个多语检查点能转向目标语言。在五个多语族中,退火模型的平均字错误率达到10 - 13%,缩小了与单语模型的大部分差距,同时在单个检查点中覆盖多种语言。所有模型在Hugging Face KhayaAI组织下以Apache - 2.0许可(仅需署名)发布,以便他人可自由微调,包括商业用途。我们保守估计所涵盖语言的母语使用者约有一亿,若算上第二语言使用者则更多。

英文摘要

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily on read speech drawn from religious texts, which offer broad, license-clear and orthographically consistent coverage for languages that otherwise lack transcribed audio. We describe a two-step (and, for one family, three-step) learning-rate-annealed fine-tuning procedure that first adapts a shared multilingual model at a high learning rate and then anneals it to recover, and in several cases surpass, strong monolingual baselines. We further describe a lightweight language-conditioning mechanism that injects a one-hot language identity as a sequence of prefix frames prepended to the acoustic features, allowing a single multilingual checkpoint to be steered to a target language at inference. Across the five multilingual families the annealed models reach average word error rates (WER) of 10-13%, closing most of the gap to monolingual models while covering many languages in a single checkpoint. All models are released on the Hugging Face KhayaAI organisation under the Apache-2.0 license (attribution only) so that others may fine-tune them freely, including for commercial use. We provide a conservative estimate that the languages covered are spoken by on the order of one hundred million first-language speakers, and by substantially more when second-language use is included.

URL PDF HTML 收藏
2607.21535 2026-07-24 cs.LG cs.CL cs.PF 新提交

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

窗口化多令牌预测:消除百万令牌上下文下的全上下文草稿键值开销

Alagappan Valliappan

机构 * NVIDIA(英伟达)

AI总结 研究百万令牌上下文下MTP草稿头成本过高问题,提出窗口化MTP方法,仅对草稿注意力应用滑动窗口等,可降低解码成本、改善延迟,在多种架构上效果显著,还能回收未读取草稿键值。

Comments 25 pages, 2 figures, 11 tables

详情
AI中文摘要

推测性解码通过让低成本的草稿提议令牌并由目标并行验证来加速自回归生成。前沿模型越来越多地内置多令牌预测(MTP/NEXTN)草稿头,假设草稿成本可忽略不计。但在百万令牌上下文时此假设不成立,MTP草稿头在每个草稿步骤通常对整个键值缓存进行全注意力计算,其读取成本随上下文线性增长并主导草稿成本。我们仅对草稿的注意力应用流式语言模型风格的滑动窗口加注意力汇聚(窗口化MTP),保持全注意力验证不变。它无需训练、即插即用且无损,将草稿的键值工作集限制为常数,在100万个令牌时减少约99%的键值条目。在单GPU上的SGLang中,针对三种架构系列在100万个令牌上下文下,窗口化使每个解码步骤的成本比原生MTP草稿降低28%至44%,端到端解码延迟也相应改善,同时保留目标的验证输出分布,且未读取的草稿键值可通过紧凑环形缓冲区回收而不影响接受率或质量。

英文摘要

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap. At million-token context this breaks: an MTP draft head typically runs full attention over the entire KV cache at every draft step, so its read grows linearly with context and comes to dominate the draft cost -- precisely where speculation is most valuable. The effect compounds with draft length (a deep native draft can turn net-negative, slower than no speculation) and sharpens under hybrid/linear-attention targets, where cheaper verification leaves the draft's full-attention read exposed. We apply a StreamingLLM-style sliding window plus attention sink to the draft's attention only (Windowed-MTP), leaving full-attention verification intact. It is training-free, drop-in, and lossless by construction: the full-attention target still decides every accepted token, so windowing changes only which tokens are proposed, never which are accepted. It bounds the draft's KV working set to a constant, dropping ~99% of KV entries at 1M. Across three architecture families (Qwen GDN-MoE 35B/122B and a Mamba2-hybrid NoPE 120B) at 1M context on a single GPU in SGLang, windowing cuts the per-decode-step cost over the shipping native MTP draft by +28% to +44%, an input-invariant margin that widens with context. Since per-token latency is this cost divided by acceptance length, at matched acceptance end-to-end decode latency improves by the same amount, and more where windowing also lifts acceptance, while preserving the target's verified output distribution. Finally, the unread draft KV -- 7.7-11% of total KV at 1M -- is reclaimed via a compact ring buffer at no acceptance or quality cost.

URL PDF HTML 收藏
2607.21529 2026-07-24 cs.CV cs.AI 新提交

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

ElasticTTT:用于视频编辑的保留先验的测试时调优

Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu

机构 * College of AI, Tsinghua University(清华大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

AI总结 研究针对预训练扩散模型的测试时调优中存在的问题,提出ElasticTTT框架,通过目标分布正则化、对比条件采样和异步噪声调度等方法,成功保留基础模型生成先验,在一次性视频编辑中达领先性能。

详情
AI中文摘要

预训练扩散模型上的测试时调优(TTT)已成为视频编辑的强大范式。然而,生成模型的分布映射性质与标准TTT的单点优化之间存在根本不匹配。本文证明这种不匹配会引发“先验崩溃”,即模型丢弃文本条件和空间潜在信息,使生成退化为源视频,或混淆不同区域的特征。为解决此问题,我们提出了ElasticTTT,这是一个保留先验生成分布并恢复生成弹性的新框架。具体而言,我们提出了目标分布正则化以防止尖锐的记忆最小值,对比条件采样以引导推理远离源偏差,以及异步噪声调度以保留未编辑区域。广泛的评估表明,ElasticTTT成功保留了基础模型的生成先验,在一次性视频编辑中实现了领先性能。

英文摘要

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.

URL PDF HTML 收藏
2607.21526 2026-07-24 cs.CV 新提交

Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving

增强自动驾驶中全天候自监督深度估计的鲁棒性

Mengshi Qi, Xiaoyang Bi, Xianlin Zhang, Huadong Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 研究自动驾驶中全天候自监督深度估计问题,提出用多教师蒸馏和鲁棒雷达融合的自训练管道,包括不确定性感知多教师蒸馏及POV-BEV雷达融合方法,经实验验证该方法具鲁棒性且性能达最优。

详情
AI中文摘要

自监督深度估计在各种恶劣天气条件下对安全自动驾驶具有挑战性,因为传感器感知能力下降。挑战主要来自两方面:恶劣条件会扭曲像素对应关系并违反自监督损失函数中的假设,导致深度预测错误;雷达虽广泛用于恶劣天气,但稀疏分布的雷达点给自监督融合带来挑战。为解决这些问题,我们引入一种新颖的自训练管道,通过多教师蒸馏和鲁棒雷达融合使用未配对的真实全天候数据。我们提出不确定性感知多教师蒸馏方法生成不同的教师模型,并用不确定性建模权衡知识蒸馏损失。此外,设计了POV-BEV雷达融合方法,利用相机像素射线约束在相机视角和雷达鸟瞰图之间建立连接,有效利用更密集的雷达点,捕捉互补视角。大量定量和定性实验证明了该方法在全天候数据集上的鲁棒性,实现了当前最优性能。代码和模型可获取。

英文摘要

Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradation. These challenges arise from two main aspects. Firstly, adverse conditions can distort pixel correspondences and violate the assumptions embedded in the self-supervised loss function, leading to erroneous depth predictions. Secondly, while radar is a widely adopted sensor in adverse weather conditions, the sparse distribution of radar points in the Point of View (POV) poses challenges for self-supervised fusion. To address these issues, we introduce a novel self-training pipeline using unpaired real all-weather data through multi-teacher distillation and robust radar fusion. We propose the Uncertainty-Aware Multi-Teacher Distillation method to generate diverse teacher models with different adverse condition inputs, and then employ uncertainty modeling to weigh the knowledge distillation loss. Additionally, we design the POV-BEV Radar Fusion approach, which leverages camera-pixel ray constraints to establish connections between the camera's Point of View (POV) and the radar's Bird's-Eye View (BEV). This approach enables the utilization of denser radar points, effectively capturing the complementary perspectives of both POV and BEV. Extensive quantitative and qualitative experiments demonstrate the robustness of our proposed method on all-weather datasets, achieving state-of-the-art performance. Our code and models are available at https://github.com/MICLAB-BUPT/RobustDepth.

URL PDF HTML 收藏
2607.21522 2026-07-24 cs.RO cs.AI cs.CL cs.CV 新提交

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

GS-Agent:通过生成式模拟创建4D物理世界

Hongxin Zhang, Chunru Lin, Junyan Li, Zhou Xian, Tsun-Hsuan Wang, Chuang Gan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Genesis AI(创世纪人工智能公司)

AI总结 研究旨在从自然语言描述创建4D物理世界。核心方法是提出GS-Agent这一端到端多智能体框架,分解任务并让多智能体与物理引擎协作。主要贡献是能有效转换自然语言为物理合理的4D世界,实现相机和灯光控制,为新范式奠基。

详情
AI中文摘要

从自然语言描述创建动态且物理逼真的4D世界既迷人又具有挑战性。传统计算机图形方法依赖手动创建,需大量人力微调材料、运动和视觉保真度。生成基础模型的进展引发了从大规模数据生成此类4D世界的兴趣,但现有方法仍难以确保物理合理性和可控性。本文利用基础模型构建代理系统,提出端到端多智能体框架GS-Agent,将任务分解为实体管理和渲染配置,多智能体协作通过代码与物理引擎交互,迭代构建符合描述的4D世界。实验结果表明,GS-Agent能有效将自然语言转换为多样且物理合理的4D世界,实现电影级相机和灯光控制,为4D世界生成新范式奠定基础。

英文摘要

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an agentic system that emulates how humans traditionally create 4D worlds, yet automates the entire process. We present GS-Agent, an end-to-end multi-agent framework that integrates physics engines in the loop to generate realistic, dynamic, and controllable 4D physical worlds from natural language. Inspired by how humans build 4D worlds, GS-Agent decomposes the task into entity management, covering 3D asset curation, material tuning, placement, and motion control, and rendering configuration, including camera and lighting manipulation. Multiple agents with distinct expertise interact with the physics engine via code, seek multimodal feedback, and collaborate to iteratively construct 4D worlds that align with the given descriptions. Experimental results show that GS-Agent effectively converts natural language into diverse and physically plausible 4D worlds exhibiting rich interactions among liquids, deformable objects, and rigid bodies, while achieving cinematic camera and lighting control. We envision GS-Agent as a foundation for a new paradigm in 4D world generation, empowering creative content creation and physical AI. Project page at https://umass-embodied-agi.github.io/gs-agent/

URL PDF HTML 收藏
2607.21518 2026-07-24 cs.AI 新提交

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

相同危险目标,相反建议:直接暴露与多智能体调解

Linjun Li

机构 * University of Pennsylvania(宾夕法尼亚大学)

AI总结 研究通过测试发现,语言模型直接面对危险目标和经智能体转换传递指令时表现不同,存在行为反向转变及组合安全漏洞,高性能模型可用于特定自动化工作流程的用户端组件,却未明确其内部机制。

Comments 21 pages; welcome comments

详情
AI中文摘要

即使是当前高性能的语言模型,直接展示危险目标时比其他智能体转换并传递其指令方向时看起来更安全。使用OpenAI的gpt-5.6-sol模型别名,测试了25个预先指定的镜像权衡配置文件。直接暴露授权隐瞒、伪造和施压的目标产生的建议与目标相反。经过Id和Censor转换后,面向用户的Superego产生的建议与目标一致。这种行为反向转变与模型识别或不信任操纵动机一致。还揭示了一个组合安全漏洞:高性能模型可用于服务明确操纵目标的自动化多阶段工作流程的用户端组件。

英文摘要

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective. The workflow can keep the raw instruction, its manipulation-authorizing clauses, and its provenance outside the downstream model's context while preserving the objective's target direction. A user with endpoint-only access likewise cannot directly inspect those upstream messages including the objective.

URL PDF HTML 收藏
2607.21504 2026-07-24 cs.CV 新提交

Texture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion Model

Texture++:使用区域感知扩散模型提升3D资产纹理分辨率

Shuaiwei Wang, Shi Li, Jieting Xu, Yuchi Huo, Qi Wang, Wenting Zheng, Rengan Xie

机构 * State Key Laboratory of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室) North China Electric Power University(华北电力大学)

AI总结 针对3D资产纹理分辨率低问题,提出Texture++框架,通过在UV空间重新表述超分辨率任务,采用自适应视图选择、四叉树纹理区域组织及基于扩散的超分辨率模型,提升纹理分辨率,相比现有方法显著改进了纹理细节与连贯性。

详情
AI中文摘要

由于纹理分辨率低,大量3D资产被废弃,而当前超分辨率模型忽略纹理图并专注于自然图像。一个高效且通用的纹理超分辨率模型可使电影和视频游戏等行业中大量老旧但有价值的资产重焕生机。我们提出Texture++,一种新颖的纹理超分辨率框架,可增强资产的低分辨率纹理以产生高分辨率、高质量的结果。具体而言,我们将超分辨率任务在UV空间中重新表述为跨多个渲染视图执行并合并输出。首先,为在视图空间中实现更完整和连续的纹理,我们提出一种自适应视图选择策略来整合分散在UV纹理补丁中的纹理。此外,我们引入一种基于四叉树的纹理区域组织方法来组合来自不同视点的超分辨率纹理,提供掩码以区分需要改进的区域。最后,我们设计一种基于扩散的超分辨率模型,为指定的掩码区域增强纹理分辨率,并与周围区域无缝集成。通过全面评估,我们证明我们的方法比现有方法产生的纹理在细节和连贯性上有显著改进。

英文摘要

Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore texture maps and focus on natural images. An efficient and generalizable texture super-resolution model can revitalize a large corpus of aging yet valuable assets across industries such as film and video games. We present Texture++, a novel framework for texture super-resolution, which enhances the low-resolution textures of assets to produce high-resolution, high-quality results. Specifically, we reformulate the task of super-resolution in UV space into performing it across multiple rendered views and merging the outputs. Firstly, to achieve more complete and continuous textures in the view space, we propose an adaptive view selection strategy to integrate textures dispersed across UV texture patches. Furthermore, we introduce a quadtree-based texture region organization method for combining super-resolved textures from different viewpoints, providing masks to distinguish regions that require improvement. Finally, we design a diffusion-based super-resolution model that enhances the texture resolution for specified masked regions, seamlessly integrating with surrounding regions. Through comprehensive evaluations, we demonstrate that our approach yields textures with substantially improved detail and coherence over existing methods.

URL PDF HTML 收藏