arXivDaily arXiv每日学术速递 周一至周五更新
全部学科分类 1774
2607.20417 2026-07-23 cs.CV 新提交

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

ATSplat:具有自适应令牌扩展的紧凑型前馈3D高斯点渲染

Cho In, Jeonghwan Cho, Mijin Yoo, Gim Hee Lee, Seon Joo Kim

机构 * Yonsei University(延世大学) National University of Singapore(新加坡国立大学)

AI总结 研究针对现有前馈3DGS方法不足,提出ATSplat框架。通过自适应3D令牌恢复自适应分配能力,先提升深度等形成场景支架,再回归高斯并解耦放置,还用模块预测不确定性分数并扩展令牌。实验表明其实现高质量渲染且减少高斯数量,提升效率。

详情
AI中文摘要

3D高斯点渲染(3DGS)通过在3D中优化自由放置的原语并在重建不足的区域自适应地使其密集化来实现高质量的新视图合成。然而,现有的前馈3DGS方法在很大程度上失去了这种场景自适应能力分配,这些方法通常在输入像素处回归高斯并沿相机光线提升它们。这种像素对齐的公式使得原语的数量和位置取决于图像分辨率和输入视点,而不是场景复杂性,导致密集且通常冗余的高斯集。我们提出了ATSplat,一个前馈3DGS框架,通过自适应3D令牌恢复3DGS优化的自适应分配能力。ATSplat首先将粗略的补丁级深度和相机线索提升到稀疏的3D锚定令牌中,形成场景的紧凑支架。然后,每个令牌通过可学习的3D偏移量回归到局部高斯,将原语放置与输入图像网格解耦。一个自适应令牌扩展模块预测令牌级不确定性分数,由渲染误差图监督,并通过可学习的扩展层选择性地扩展高不确定性令牌。这种从稀疏到自适应的公式使ATSplat能够在具有挑战性的区域集中原语,同时保持紧凑的表示。在两个代表性数据集RealEstate10K和DL3DV上的实验表明,ATSplat实现了最先进的渲染质量,同时与密集的前馈3DGS方法相比,高斯数量减少了5.7倍以上。从12张分辨率为512×960的输入图像中,ATSplat使用单个商用GPU在不到一秒的时间内完成重建,并以1136 FPS(512×960)的速度渲染高质量的新视图,仅使用311K个高斯。

英文摘要

3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than $5.7\times$ compared with dense feed-forward 3DGS methods. From 12 input images at $512 \times 960$ resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS ($512 \times 960$) with only 311K Gaussians.

URL PDF HTML 收藏
2607.20410 2026-07-23 cs.CL 新提交

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

LKValues:使大语言模型与斯里兰卡社会价值观保持一致

Nethmi Muthugala, Supryadi, Surangika Ranathunga, Nisansa de Silva, Ruijie Tao, Ovindu Gunatunga, Pengyun Zhu, Shaowei Zhang, Jingting Zheng, Deyi Xiong

机构 * TJUNLP Lab, School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院TJUNLP实验室) School of Mathematical and Computational Sciences, Massey University(梅西大学数学与计算科学学院) Department of Computer Science & Engineering, University of Moratuwa(莫拉图瓦大学计算机科学与工程系) Johns Hopkins University(约翰霍普金斯大学) School of Computing, University of Colombo(科伦坡大学计算机学院)

AI总结 研究针对大语言模型价值对齐存在西方文化偏见问题,以斯里兰卡为例,提出LKValues资源套件,通过调查得出社会价值观,构建语料库和评估基准,经实验发现其能改善模型表现,为低资源国家价值对齐提供可复制流程。

Comments 37 pages, 10 figures, and 15 tables. Includes appendices. Datasets are available at the project repository

详情
AI中文摘要

大语言模型(LLMs)的价值对齐在文化上偏向西方规范,导致像斯里兰卡这样具有独特文化动态的多语言社会中当地价值观被错误处理。现有基准忽略了斯里兰卡官方语言僧伽罗语中的情境化价值观。为弥补这一差距,我们提出LKValues,首个基于调查的斯里兰卡价值对齐资源套件。通过对205名受访者的三语调查,融合全球框架和大语言模型引出的本地结构,得出40个多数认可的社会价值观。利用这些价值观构建了包含15万个基于场景实例的僧伽罗语 - 英语新闻衍生指令语料库LKvaluesIT和1000个实例的价值敏感评估基准LKvaluesBench。我们用LKvaluesBench评估了一系列专有和开源权重的大语言模型,并对三个开源权重基础模型进行微调。实验表明,更新更大的大语言模型仍存在资源和文化价值对齐差距。LKValues微调改善了Qwen系列模型在英语和僧伽罗语方面的表现,减少了无效输出和跨语言差异,不过收益仍依赖于模型家族。这些突出了LKValues在嵌入斯里兰卡价值观方面的功效,为低资源、特定国家的多元价值对齐提供了可复制的流程。数据集可在指定网址公开获取。

英文摘要

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning. To bridge this gap, we propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment. From a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs, we derive 40 majority-endorsed societal values. Using these values, we construct LKvaluesIT, a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances, and LKvaluesBench, a value-sensitive evaluation benchmark of 1,000 instances. We evaluate a set of proprietary and open-weight LLMs with LKvaluesBench. We fine-tune three open-weight base models (Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base). Our experiments show that newer and larger LLMs still exhibit low-resource and cultural value-alignment gaps. LKValues fine-tuning improves Qwen-family models in English and Sinhala, reducing invalid outputs and cross-lingual disparities, though gains remain model-family dependent. These highlight LKValues efficacy in embedding Sri Lankan values, offering a replicable pipeline for low-resource, country-specific pluralist value alignment. The dataset is publicly available at this https URL.

URL PDF HTML 收藏
2607.20402 2026-07-23 cs.AI 新提交

SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data

SoftReason:一种用于高维感知数据的完全可微神经软符号演绎推理架构

Wael AbdAlmageed

机构 * Clemson University(克莱姆森大学)

AI总结 研究针对前提需从高维输入推断且由知识图谱提供相关信息的推理问题,提出神经软符号架构SoftReason,核心是对直接后果算子可微提升,能在知识感知视觉问答中支持多种功能。

详情
AI中文摘要

在许多推理问题中,前提并非作为离散符号被观察到,而是必须从高维输入中推断出来。此外,谓词词汇、论证结构和可信证据由知识图谱(KG)或规则定义提供。经典的神经符号管道在感知和演绎之间有离散接口。我们提出了一种神经软符号架构,用于对潜在感知事实和知识提供的谓词进行可微演绎推理。SoftReason通过将演绎状态表示为候选常量和谓词上的局部软解释张量来消除梯度差距。感知提出概率性基本事实,KG三元组作为高置信度软证据进入,每个查询锚点、谓词选择和闭包更新都是可微的。我们的核心创新是对直接后果算子的可微提升。它使用谓词定义嵌入和潜在组合通道来形成软主体谓词混合,在所有可能的见证上聚合,提出查询条件头部事实,并通过单调概率或运算更新解释。我们在知识感知视觉问答(KVQA)上实例化了该框架,并展示了SoftReason如何在一个可训练架构中支持端到端感知基础、KG证据注入和可微演绎闭包。

英文摘要

In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or rule definitions. Classical neuro-symbolic pipelines have a discrete interface between perception and deduction. We present a neuro-soft-symbolic architecture for differentiable deductive reasoning over latent perceptual facts and knowledge-provided predicates. SoftReason removes the gradient gap by representing the deductive state as a local soft interpretation tensor over candidate constants and predicates. Perception proposes probabilistic base facts, KG triples enter as high-confidence soft evidence, and every query anchor, predicate choice, and closure update remains differentiable. Our core innovation is a learned differentiable lift of the immediate-consequence operator. It uses predicate-definition embeddings and latent composition channels to form soft body-predicate mixtures, aggregate over all possible witnesses, propose query-conditioned head facts, and update the interpretation through a monotone probabilistic OR. We instantiate the framework on Knowledge-aware Visual Question Answering (KVQA), and demonstrates how SoftReason supports end-to-end perceptual grounding, KG evidence injection, and differentiable deductive closure in one trainable architecture.

URL PDF HTML 收藏
2607.20399 2026-07-23 cs.RO cs.HC cs.LG 新提交

Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning

利用虚拟现实和强化学习实现微型仿人机器人的远程定位操作

Nicolas Kosanovic, Jordan Dowdy, Jean Chagas Vaz

机构 * University of Louisville(路易斯维尔大学)

AI总结 研究针对微型仿人机器人缺乏类似全尺寸机器人控制堆栈的问题,利用虚拟现实和强化学习开发了控制堆栈,经实验验证能实现一定速度行走及远程定位操作,展现了微型仿人机器人远程定位操作的潜力。

Comments 8 pages, 6 figures. Accepted manuscript. Published in the 2025 IEEE-RAS 24th International Conference on Humanoid Robots (Humanoids), pp. 1233-1240

详情
AI中文摘要

近年来,全尺寸仿人机器人能力呈指数级增长,旨在在人类环境中进行通用部署。制造商常用的一种控制方法是利用虚拟现实进行上身遥操作,利用强化学习进行下身平衡和运动控制。然而,这种强大的控制堆栈通常只用于昂贵的全尺寸机器人,许多研究团队无法使用。微型仿人机器人更为普遍,但设计中仿生学应用较少,且缺乏类似的发展。本文描述了一种专门为微型仿人机器人从头开发的柔顺全身临场感控制堆栈。在ROBOTIS OP3硬件上进行的框架实验展示了最高可达0.45 m/s的独立于手臂运动的行走速度。通过与专业人类操作员进行的立方体重新定位实验演示了远程定位操作。平均而言,遥控操作系统在10分钟内移动了2个不同的40 g立方体,总共行走了5 m。总体而言,所开发的系统显示出微型仿人机器人远程定位操作的潜力。

英文摘要

Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilizes Virtual Reality for upper-body teleoperation and Reinforcement Learning for lower-body balance and locomotion control. As a result, a single remote operator can see, manipulate, and navigate about a real, distant physical environment. This powerful control stack is often relegated to expensive full-sized robots, many of which are inaccessible to the research community. Miniature humanoids are more prevalent, but employ less biomimicry in their design (e.g. fewer sensors, Degrees of Freedom, etc) and lack similar developments. This paper describes a compliant full-body telepresence control stack developed from the ground up for miniature humanoids. Framework experimentation on ROBOTIS OP3 hardware showcases walking at speeds up to 0.45 m/s independent of arm motions. Tele-loco-manipulation is demonstrated via a cube relocation experiment with an expert human operator. On average, the teleoperated system moved 2 different 40 g cubes within 10 mins, walking a total distance of 5 m. Overall, the developed system shows potential for miniature humanoid tele-loco-manipulation.

URL PDF HTML 收藏
2607.20392 2026-07-23 cs.RO 新提交

Distributed Acoustic Localization Array Deployed Using a Soft Everting Vine Robot

使用软外翻藤蔓机器人部署的分布式声学定位阵列

Sebastian Lorca Godoy, Ciera McFarland, Michael Val, Antonio Alvarez Valdivia, Nathaniel Hanson, Margaret McGuinness

机构 * University of Notre Dame(圣母大学) Pontifical Catholic University of Chile(智利天主教大学) Lincoln Laboratory, Massachusetts Institute of Technology(麻省理工学院林肯实验室)

AI总结 研究在受限非结构化环境中用软外翻藤蔓机器人定位灾难受害者,提出动态转向响应功率与相位变换框架,通过实验测量不同放置和配置下的定位准确性,展示藤蔓机器人生长定位声源能力,凸显分布式声学传感潜力。

Comments Sebastian Lorca Godoy, Ciera McFarland, Michael Val, Antonio Alvarez Valdivia, Nathaniel Hanson, and Margaret McGuinness, "Distributed Acoustic Localization Array Deployed Using a Soft Everting Vine Robot", in IEEE International Conference on Intelligent Robots and Systems, 2026

详情
AI中文摘要

软机器人外部感知在各种现场应用中得到越来越多的探索。在这项工作中,我们提出了一种基于声音的系统,用于在受限和非结构化环境中定位灾难受害者,该系统基于沿软外翻藤蔓机器人身体嵌入的分布式声学传感架构。我们提出了一种动态转向响应功率与相位变换框架,该框架支持远场到达方向估计和机器人接近声源时的近场三维源定位。为了更好地理解与使用柔软、可变形机器人身体定位声音相关的设计和控制空间,我们进行了实验,测量了附着在机器人身体上的五麦克风阵列在相对于机器人外膜的三种放置方式(加压体内、内尾内和外壁外)以及四种机器人配置(线性、双线性、圆形和正弦形)下这些方法的准确性。我们测量了随着信噪比、接近方向和声源与阵列中心的距离变化时准确性的变化。最后,我们展示了一个藤蔓机器人在其外壁携带麦克风时生长成任意形状,并表明在只有三个麦克风从机器人身体外翻后,位于阵列近场的声源可以高精度定位。这些结果突出了分布式声学传感在使用软生长机器人进行可靠受害者定位方面的潜力。

英文摘要

Soft robot exteroception is increasingly being explored for a variety of field applications. In this work, we present a sound-based system for localizing disaster victims in confined and unstructured environments, based on a distributed acoustic sensing architecture embedded along the body of a soft everting vine robot. We propose a dynamic Steered Response Power with Phase Transform framework that supports both far-field direction-of-arrival estimation and near-field three-dimensional source localization as the robot approaches the sound source. To better understand the design and control space related to localizing sound using a soft, shape-morphing robot body, we conduct experiments measuring the accuracy of these methods for a five-microphone array attached to the robot body using three placements relative to the outer membrane of the robot (inside the pressurized body, inside the inner tail, and outside the outer wall) and in four robot configurations (linear, double linear, circular, and sinusoidal). We measure the change in accuracy as the signal-to-noise ratio, the direction of approach, and the distance of the sound source from the center of the array change. Finally, we demonstrate a vine robot growing into an arbitrary shape while carrying microphones along its outer wall, and show that a sound source located with the array's near field can be localized with high accuracy after only three microphones have everted from the robot body. These results highlight the potential of distributed acoustic sensing for reliable victim localization using soft growing robots.

URL PDF HTML 收藏
2607.20389 2026-07-23 cs.CV 新提交

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

PercepCap:具有结构化时空感知的视频字幕生成器

Yifan Xu, Zihao Wang, Zhixiao Wang, Jiaming Zhang, Yichun Yang, Desen Meng, Yuanxing Zhang, Pengfei Wan, Limin Wang

机构 * Nanjing Univerisity(南京大学) Kuaishou Technology(快手科技) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 研究视频字幕生成问题,提出PercepCap框架,遵循感知-描述生成链,设计两阶段训练策略及相关数据构建方法,在评估中优于基线,提升视频字幕生成质量。

详情
AI中文摘要

视频字幕需要对视频进行细粒度的时空理解,包括物体位置的空间感知和事件发生时间的时间感知。现有多模态语言模型通常直接从视频输入生成字幕,而不展示描述背后的感知证据。因此,时空感知错误只能在最终字幕中观察到,难以直接识别潜在的感知错误。为了解决这些问题,我们提出了PercepCap,这是一个感知感知视频字幕框架,在生成最终字幕之前使感知证据明确。具体来说,PercepCap遵循感知-描述生成链,模型首先生成包含物体轨迹和时间事件的时空感知轨迹,然后根据感知到的证据生成最终字幕。为了支持这一点,我们设计了一个两阶段训练策略。感知-然后-描述监督微调使模型从仅字幕生成适应到提出的感知-描述链,而感知基础强化学习通过对感知链和最终字幕的联合奖励来优化感知轨迹和字幕质量。为了支持我们的两阶段训练,我们引入了字幕锚定感知数据构建。该管道通过首先生成仅字幕描述,提取其中提到的物体和事件,并将它们用框和时间戳重新定位到视频中,来构建SFT和RL训练数据。这产生了字幕对齐的感知数据,提供了可靠的训练地面真值,确保明确的感知轨迹和最终字幕指的是相同的物体和事件。在直接字幕和字幕到问答评估中,PercepCap始终优于Qwen3-VL基线,并展示了领先的字幕质量。

英文摘要

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without exposing the perceptual evidence behind descriptions. As a result, mistakes in spatiotemporal perception are only observed in the final caption, making it difficult to identify the underlying perceptual errors directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evidence explicit before producing the final caption. Specifically, PercepCap follows a perceive-describe generation chain, where the model first produces a spatiotemporal perception trace comprising object trajectories and temporal events, and then generates the final caption conditioned on the perceived evidence. To support this, we design a two-stage training strategy. Perceive-then-Describe Supervised Fine-tuning adapts the model from caption-only generation to the proposed perceive-describe chain, while Perception-Grounded Reinforcement Learning optimizes perception trace and caption quality with joint rewards over perception chain and the final caption. To support our two-stage training, we introduce Caption-Anchored Perception Data Construction. This pipeline builds the SFT and RL training data by first generating a caption-only description, extracting the objects and events it mentions, and grounding them back in the video with boxes and timestamps. This yields caption-aligned perception data that provides solid training ground truth, ensuring that the explicit perception trace and final caption refer to the same objects and events. Across direct caption and caption-to-QA evaluation, PercepCap consistently improves upon the Qwen3-VL baseline and demonstrates leading caption quality.

URL PDF HTML 收藏
2607.20374 2026-07-23 cs.LG 新提交

Online Variance Reduction for Domain Adaptation on Streaming Data

流数据域适应的在线方差缩减

Andrea Napoli

机构 * University of Southampton(南安普顿大学)

AI总结 研究流数据域适应的随机方差缩减问题,提出在线SVR算法ARROW,通过维护移动平均参考和自适应重新加权小批量数据,使小批量与参考统计量对齐,实验表明该算法在运行时间、方差缩减及目标域准确性方面与离线算法相当。

详情
AI中文摘要

本文研究了最大均值差异(MMD)和相关对齐(CORAL)损失函数的随机方差缩减(SVR)问题。尽管已提出多种针对这些损失的离线SVR算法,但它们与在线、分布式或增量学习设置不兼容。本文提出了通过在线重新加权的自适应方差缩减(ARROW),这是首个用于流数据的MMD和CORAL在线SVR算法。该方法维护对齐统计量的移动平均参考,并自适应地重新加权输入的小批量数据,以使小批量和参考统计量对齐。此外,还提出了一种宽松的重新加权方案,以使后续的权重优化问题易于处理。在实验和模拟中,ARROW在运行时间、方差缩减程度和目标域准确性方面与离线算法具有竞争力。

英文摘要

This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. Although various offline SVR algorithms for these losses have been proposed, these are incompatible with online, distributed, or incremental learning settings. This paper presents Adaptive vaRiance Reduction via Online reWeighting (ARROW), the first online SVR algorithm for the MMD and CORAL for streamed data. The method maintains moving average references of the alignment statistics, and adaptively reweights incoming minibatches so that the minibatch and reference statistics are aligned. Further, we propose a relaxed reweighting scheme so that the ensuing weight-optimisation problem is tractable. In experiments and simulations, we show that ARROW performs competitively with offline algorithms in terms of runtime, degree of variance reduction achieved, and target domain accuracy.

URL PDF HTML 收藏
2607.20372 2026-07-23 cs.CL 新提交

Notes to Self: Can LLMs Benefit from Experiential Abstractions?

给自己的笔记:大语言模型能从经验抽象中受益吗?

Chang Liu, Xinyu Li, Artur Dubrawski

机构 * Auton Lab, Carnegie Mellon University(卡内基梅隆大学自动实验室)

AI总结 研究大语言模型能否从经验抽象中受益,通过从其在MATH训练集的痕迹提取抽象存入可检索库,探索推理时检索和强化学习两种使用模式,发现能提高模型在数学和逻辑推理基准上的性能,且框架可转移。

详情
AI中文摘要

人类将经验提炼为可复用的抽象,如策略和警示提醒,并应用它们逐步更有效地解决问题。我们研究大语言模型(LLMs)是否能同样从这种经验抽象中受益。从LLMs在MATH训练集上的解决方案痕迹中,一个更强的教师模型或LLMs自身将自然语言抽象提取到一个可检索库中。我们探索两种使用模式:(1)推理时检索和(2)使用抽象增强训练提示的强化学习(RL)。经验抽象提高了LLMs在数学和逻辑推理基准上的性能。自我提取的抽象与教师提取的抽象相匹配,且我们的抽象使用框架可转移到其他数据集和模型。这些发现表明LLMs能像人类利用提炼的经验一样提取和应用经验抽象。

英文摘要

Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply them to gradually solve problems more effectively. We study whether Large Language Models (LLMs) can similarly benefit from such experiential abstractions. From LLMs' solution traces on the MATH training set, a stronger teacher or the LLMs themselves extract natural-language abstractions into a retrievable library. We explore two usage modes: (1) inference-time retrieval and (2) reinforcement learning (RL) with abstraction-augmented training prompts. Experiential abstractions improve LLM performance on mathematical and logical reasoning benchmarks. Self-extracted abstractions match teacher-extracted ones, and our abstraction usage framework can transfer to other datasets and models. These findings suggest LLMs can extract and apply experiential abstractions much as humans leverage distilled experience.

URL PDF HTML 收藏
2607.20368 2026-07-23 cs.CV 新提交

Self Gradient Forcing: Native Long Video Extrapolation

自梯度强制:原生长视频外推

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian, Yaowei Li, Yawen Luo, Yijun Liu, Weiyang Jin, Songchun Zhang, Xianglong He, Xuying Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan

机构 * Joy Future Academy(快乐未来学院)

AI总结 研究针对自回归视频扩散方法的历史上下文梯度差距问题,提出自梯度强制(SGF)两阶段训练策略,利用未来视频潜变量损失监督模型将上下文编码为有效因果记忆,在长视频外推实验中效果优于自强制。

Comments Project page: this https URL (https://zhuang2002.github.io/SelfGradientForcing/)

详情
AI中文摘要

近期自回归视频扩散方法多基于自强制构建,学生模型在自身展开生成的历史上训练,减少了曝光偏差,但存在历史上下文梯度差距问题。本文提出自梯度强制(SGF),一种两阶段训练策略。第一阶段进行无梯度自回归展开匹配推理并记录相关信息,第二阶段对记录步骤进行并行上下文梯度重建。通过该方法,SGF在原生自回归训练目标中提供缺失的内存写入监督。实验表明,SGF在不同初始化下的长视频外推效果优于自强制,仅用5秒训练窗口就能外推到几分钟的视频。代码和模型将发布以推动自回归视频生成研究。

英文摘要

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.

URL PDF HTML 收藏
2607.20367 2026-07-23 cs.LG 新提交

Variance-reduced Domain Adaptation using Paired Sampling

使用配对采样的方差减少域适应

Andrea Napoli

机构 * University of Southampton(南安普顿大学)

AI总结 针对无监督域适应中分布匹配框架损失方差高且缺乏有限和结构的问题,提出PSDA技术,通过在域内和域间配对观测值形成四元组,最小化预期梯度方差,经实验验证该方法能降低方差并提高目标域准确性。

详情
AI中文摘要

相关对齐和最大均值差异是无监督域适应(UDA)中广泛使用的两种分布匹配框架。然而,这些损失中的高方差已被证明会削弱它们在小批量优化设置中的有效性。此外,这些损失缺乏有限和结构,这使得它们与经典随机方差减少(SVR)方法不兼容。本文提出了用于域适应的配对采样(PSDA),这是一种针对此类目标量身定制的新型SVR技术。PSDA在域内和域间对观测值进行配对,形成四元组,在训练期间总是一起采样。配对设计用于最小化预期梯度方差,并简化为解决一组线性分配问题。我们的模拟表明与相关方法相比方差降低,并且在三个域转移数据集上的实验显示目标域准确性提高。

英文摘要

Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, high variance in these losses has been shown to undermine their effectiveness in minibatch optimisation settings. Furthermore, the losses lack finite-sum structure, which renders them incompatible with classical stochastic variance reduction (SVR) methods. This paper proposes Paired Sampling for Domain Adaptation (PSDA), a novel SVR technique tailored to such objectives. PSDA pairs observations both within and across domains, to form quadruplets that are always sampled together during training. The pairings are designed to minimise expected gradient variance, and reduce to solving a set of linear assignment problems. Our simulations demonstrate reduced variance compared to related methods, and experiments on three domain shift datasets show improved target domain accuracy.

URL PDF HTML 收藏
2607.20357 2026-07-23 cs.CV 新提交

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

少看,快思考:多模态大语言模型的联合令牌-计算适配

Pengcheng Wang, Zhiquan Wang, Jayoung Lee, Zhuoyan Xu, Ran Xu, Saurabh Bagchi, Yin Li, Somali Chaterji

机构 * Purdue University(普渡大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) NVIDIA(英伟达)

AI总结 针对多模态大语言模型推理成本高的问题,提出SmartVL框架,通过视觉侧令牌控制器和LLM侧计算控制器联合控制视觉令牌数量和模型计算能力,实验证明该框架优于先前方法,实现更好的精度-效率平衡。

Comments Accepted at ECCV 2026

详情
AI中文摘要

多模态大语言模型(MLLMs)在视觉-语言任务中表现出色,但推理成本高阻碍实际部署。近期工作尝试单独优化各维度来降低成本,却忽视了计算资源需根据输入内容动态分配这一耦合关系。为此提出SmartVL统一自适应推理框架,它通过视觉侧令牌控制器和LLM侧计算控制器联合控制视觉令牌数量和模型计算能力,并使其相互协调以满足目标预算。实验表明,SmartVL优于先前自适应方法,实现了更好的精度-效率帕累托前沿。

英文摘要

Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large number of input visual tokens and the heavy computation of the large language model (LLM), remains a key barrier to practical deployment. Recent work attempts to reduce the cost by adaptively optimizing individual dimensions, e.g., pruning redundant visual tokens or skipping LLM layers and heads. Nonetheless, prior approaches typically treat these dimensions independently and overlook a fundamental coupling: the available compute resources must be dynamically allocated across all dimensions based on the input content. To bridge the gap, we propose SmartVL, a unified adaptive inference framework that jointly controls vision token number and model compute capability in response to varying input contents and compute budgets. SmartVL introduces a vision-side token controller that dynamically selects informative visual tokens and an LLM-side compute controller that adaptively adjusts LLM computation. Importantly, these controllers are trained to coordinate with each other so that the overall inference cost satisfies a target budget. To allow this joint scheduling, we connect the controllers using a shared budget encoding and leverage a differentiable latency estimator for end-to-end training. This design enables SmartVL to learn cross-stage allocation strategies that adapt to both input complexity and runtime compute constraints. Experiments across multiple MLLM benchmarks demonstrate that, with joint scheduling, SmartVL consistently outperforms prior adaptive methods and achieves superior accuracy-efficiency Pareto frontiers. Project page: this https URL.

URL PDF HTML 收藏
2607.20352 2026-07-23 cs.RO 新提交

Distributed Motion Planning with Safety Guarantees for Self-Reconfiguring Robotic Boats

具有安全保障的自重构机器人船分布式运动规划

Alejandro Gonzalez-Garcia, Wei Wang, Wei Xiao, Wilm Decre, Jan Swevers, Carlo Ratti, Daniela Rus

机构 * KU Leuven(鲁汶大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Nanyang Technological University(南洋理工大学) SMART(新加坡制造技术研究院) SENSEable City Laboratory, Massachusetts Institute of Technology(麻省理工学院可感知城市实验室) Computer Science and Artificial Intelligence Lab (CSAIL), Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

AI总结 研究水生自重构机器人多智能体形状形成与重构问题,提出将分布式MPC与CBF相结合的混合框架,利用ADMM求解MPC方案,通过局部优化和信息交换计算轨迹,并应用基于CBF的滤波器保障安全,经仿真和实验验证了框架的有效性与可扩展性。

Comments Submitted to IEEE

详情
AI中文摘要

水生自重构机器人必须在确保多个智能体之间安全交互的同时组装成所需形状。本文提出了一种混合框架,将分布式模型预测控制(MPC)与控制障碍函数(CBF)相结合,用于多智能体形状形成和重构。给定所需形状和目标分配,通过交替方向乘子法(ADMM)求解的分布式MPC方案通过局部优化和信息交换计算协调轨迹。为实时确保安全,应用基于分布式CBF的滤波器来强制避免智能体间碰撞。该方法利用MPC的预测能力减轻局部极小值,而CBF尽管基础优化问题非凸仍提供形式上的安全保障。多达25个智能体的仿真结果和四个物理机器人的实验验证证明了该框架的有效性和可扩展性。

英文摘要

Aquatic self-reconfigurable robots must assemble into desired shapes while ensuring safe interactions among multiple agents. This paper proposes a hybrid framework that combines distributed Model Predictive Control (MPC) with Control Barrier Functions (CBFs) for multi-agent shape formation and reconfiguration. Given a desired shape and target assignment, a distributed MPC scheme, solved via the Alternating Direction Method of Multipliers (ADMM), computes coordinated trajectories through local optimization and information exchange. To ensure safety in real time, distributed CBF-based filters are applied to enforce inter-agent collision avoidance. The proposed approach leverages the predictive capabilities of MPC to mitigate local minima, while CBFs provide formal safety guarantees despite the nonconvexity of the underlying optimization problem. Simulation results with up to 25 agents and experimental validation with four physical robots demonstrate the effectiveness and scalability of the framework.

URL PDF HTML 收藏
2607.20351 2026-07-23 cs.CV cs.CL 新提交

Test-Time Training for Modality Order Consistency in Vision-Language Models

视觉语言模型中模态顺序一致性的测试时训练

Aditi Gupta, Yossi Gandelsman

机构 * University of Chicago(芝加哥大学)

AI总结 研究发现视觉语言模型对图像和问题呈现顺序敏感,利用此设计测试时训练方法,缩小模态顺序差距,使两种顺序相互一致,定位顺序失败区域,证明该方法可缓解故障并提升性能。

Comments 16 pages, 7 figures, preprint

详情
AI中文摘要

我们发现视觉语言模型对特定语义无关变化敏感:图像和问题呈现的顺序。在三个模型和三个基准测试中,图像优先提示始终优于问题优先提示,揭示了可重复的模态顺序失败。我们利用这一差距设计了一种顺序一致的测试时训练方法。该方法在所有评估设置中大幅缩小了模态顺序差距。令人惊讶的是,它还在更强的图像优先分支上相对于基线产生了一致的改进,使两种顺序相互一致。激活修补将顺序失败定位到网络中间一个狭窄区域,测试时训练方法修复了各层的这种错位。我们的结果表明模态顺序敏感性是视觉语言模型中的电路级故障,并证明简单的非对称测试时适应可以有效缓解它甚至提高性能。

英文摘要

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. Surprisingly, it also yields consistent improvements in the stronger image-first branch over the baseline, hence bootstrapping both orderings toward mutual consistency. Activation patching localizes the ordering failure to a narrow mid-network region where representations diverge sharply between prompt orders. We find that the test-time training method repairs this misalignment across layers. Together, our results identify modality-order sensitivity as a circuit-level failure in VLMs and demonstrate that simple, asymmetric test-time adaptation can effectively mitigate it and even improve performance over the baseline.

URL PDF HTML 收藏
2607.20349 2026-07-23 cs.CL cs.AI cs.CY 新提交

Generative AI floods and dilutes the market for books

生成式人工智能充斥并稀释了图书市场

Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon

机构 * Stony Brook University(纽约州立大学石溪分校) Columbia Law School(哥伦比亚法学院) University of Michigan(密歇根大学) MIT Initiative on the Digital Economy(麻省理工学院数字经济倡议)

AI总结 研究通过对亚马逊上自出版小说检测,发现含大量AI文本书籍占书目比重大但销售占比小,其达商业规模且重塑市场,单本销售收入下降,还影响不同类型书籍及畅销书语言借鉴情况,结果关乎版权侵权合理使用抗辩的市场效应问题。

Comments Working Paper Under Review

详情
AI中文摘要

生成式人工智能能够以近乎零成本创作出篇幅达书籍长度的虚构作品。这些书籍常被视为低质量的“垃圾”而遭买家忽视,且被认为商业价值不大。我们通过对2023年至2026年在亚马逊上销售的14419本自出版类型小说进行全文人工智能检测,并匹配到2026年6月的每日销售记录来验证这一假设。这些书均未披露是否包含人工智能生成的内容。我们发现,检测出大量人工智能文本(>25%)的书籍在书目总量中占比大,但销售占比小。即便如此,它们达到了商业规模,随着时间推移销售份额不断增加,占据了更多原本由未检测出人工智能文本的书籍所占据的稀缺高位。在此期间,一个季度内有销售记录的书籍数量增长了19.2倍,而季度收入仅增长了8.9倍。因此,市场上书籍销售数量的增长速度超过了收入增长速度,大多数类型书籍的单本销售收入下降。在人工智能传播率高的类型中,尤其是Kindle Unlimited可用性高的地方,无人工智能文本的书籍损失最大。在畅销书当中,有大量人工智能文本的书籍比无人工智能文本的书籍借鉴了更多现有书籍中的独特语言;对于这些书来说,重叠度随收入上升,而无人工智能文本的书籍则未检测到这种梯度变化。因此,生成式人工智能可通过规模而非质量重塑创意市场。我们的结果直接关乎版权侵权合理使用抗辩核心的市场效应问题。

英文摘要

Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buyers will ignore, and are assumed to carry little commercial weight. We test that assumption with full-text AI detection across 14,419 self-published genre-fiction books sold on Amazon from 2023 to 2026, matched to daily sales records through June 2026. None of these books disclose whether or not they contain AI-produced content. We find that books for which we detected substantial AI text ($>$ 25\%) make up a large share of the catalog but a smaller share of sales. Even so, they reach commercial scale, winning a growing share of sales over time and taking more of the scarce top-rank positions once held by books with no detected AI text. Over this period, the number of books with observed sales in a quarter grew 19.2-fold, while quarterly revenue grew only 8.9-fold. The market therefore added selling books faster than it added revenue, and revenue per selling book fell across most genres. Books with no AI text lose the most ground in genres with high AI diffusion, and most of all where Kindle Unlimited availability is high. Among top-selling books, those with substantial AI text draw on more distinctive language from existing books than do books with no AI text; for these books overlap rises with revenue, a gradient we do not detect for books with no AI text. Generative AI can thus reshape a creative market through scale rather than quality. Our results bear directly on the market-effect question at the center of the fair use defense to copyright infringement.

URL PDF HTML 收藏
2607.20345 2026-07-23 cs.RO cs.AI 新提交

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

弥合实验室与商店之间的差距:一种用于零售类人机器人的数据高效训练后及经验驱动学习的VLA框架

Roger Sala Sisó, Tiago Silvério, Jakob Sand, Tran Nguyen Le

机构 * HIVE Robots(蜂巢机器人公司) Technical University of Denmark(丹麦技术大学)

AI总结 研究针对VLA类人机器人弥合实验室与实际应用差距的问题,提出DEED系统级方法,含数据高效训练后管道、经验驱动细化研究及潜在空间分析工具,经实验证明精心设计数据和训练后处理可解决该问题。

Comments 8 pages. This work has been submitted to the IEEE for possible publication

详情
AI中文摘要

弥合基准性能与可靠的实际操作之间的差距仍然是视觉语言动作(VLA)类人机器人面临的核心挑战,这类机器人必须应对执行错误、分布变化和环境变异性。本文提出了DEED(数据高效训练后及经验驱动学习),这是一种在超市芯片补货任务中使用宇树G1-Edu类人机器人和GR00T N1.6基础模型进行评估的系统级方法。DEED包括三个关键组件:(1)一个具有控制频率对齐、数据管理、任务相关视觉突出显示和降低VLA依赖性的数据高效训练后管道;(2)通过基于文本的优势前缀和视觉语言价值函数从RECAP改编而来的经验驱动细化的实际研究;(3)用于研究分布内和分布外行为的潜在空间分析工具。我们的结果表明,弥合实验室与商店之间的差距主要是系统集成挑战而非架构挑战:精心的数据设计和有针对性的训练后处理可以仅使用单个GPU将在简单微调下失败的策略转变为一个能胜任实际操作的系统。

英文摘要

Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, data curation, task-relevant visual highlighting, and reduced VLA dependence; (2) a real-world study of experience-driven refinement, adapted from RECAP via a text-based advantage prefix and a vision-language value function; and (3) a latent-space analysis tool for studying in- and out-of-distribution behavior. Our results suggest that bridging the lab-to-store gap is primarily a systems integration challenge rather than an architectural one: careful data design and targeted post-training can transform a policy that fails under naive fine-tuning into a competent real-world system using only a single GPU.

URL PDF HTML 收藏
2607.20327 2026-07-23 cs.CL 新提交

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference

PyroDash:具有成本效益的令牌级小语言模型与大语言模型协作推理

Niqi Lyu, Pengtao Shi, Wei Qiu, Jianlin Zhong, Sicong Xia, Jianyao Ma, Yicheng Ding

机构 * Pyromind Dynamics Inc.(Pyromind动力学公司)

AI总结 研究针对大语言模型推理成本高、小语言模型可靠性低的问题,提出PyroDash框架,通过令牌级协作推理,分三阶段训练小语言模型,在数学推理基准测试中能支持不同操作点,可减少大语言模型使用并保持推理性能。

Comments 19 pages, 3 figures

详情
AI中文摘要

大语言模型(LLMs)推理能力强但大规模服务成本高,小语言模型(SLMs)成本低但处理难题可靠性差。我们引入了PyroDash,一个用于令牌级SLM-LLM协作推理的成本感知框架。生成过程中,SLM通过发出控制令牌决定是否请求协助,协作引擎将查询和部分推理轨迹发送给冻结的LLM完成。PyroDash分三个阶段训练SLM,其奖励平衡答案准确性和仅使用LLM推理归一化后的推理成本。在五个数学推理基准测试中,PyroDash支持不同的准确性-成本操作点,结果表明学习到的令牌级交接可减少LLM使用并保持强大推理性能。

英文摘要

Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework for token-level SLM-LLM collaborative inference. During generation, the SLM decides whether to request assistance by emitting a control token. A Collaborate Engine then sends the query and partial reasoning trace to a frozen LLM for completion through a single handoff. The policy is internalized in the SLM, requiring neither a separate router, LLM retraining, nor access to LLM logits. PyroDash trains the SLM in three stages: control-token embedding learning, offloading-oriented supervised fine-tuning, and cost-aware alignment with Group Relative Policy Optimization. Its reward balances answer accuracy against inference cost normalized by LLM-only inference. Across five mathematical reasoning benchmarks, PyroDash supports different accuracy-cost operating points. With $\lambda=0.05$, it achieves 64.04 percent average accuracy, 6.36 percentage points above the LLM-only baseline, while reducing cost by 20.4 percent. With $\lambda=0.6$, it achieves 54.55 percent accuracy with a 1.90 percent LLM token ratio and 0.012 LLM calls per example, reducing total cost from USD 49.36 to USD 1.78. These results show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance.

URL PDF HTML 收藏
2607.20301 2026-07-23 cs.LG cs.CL 新提交

The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

维度的祝福:高维空间中的近正交性如何解释时间可移植性

Abigail Woodring, Adrian Chan, Rana Muhammad Shahroz Khan, Sukwon Yun, Chau-Wai Wong, Tianlong Chen

机构 * NC State University(北卡罗来纳州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

AI总结 研究探讨PortLLM在持续预训练中LoRA补丁的长期时间可移植性及有效性。通过对Mistral、Gemma和Qwen基础模型进行实证研究,并提供理论分析,发现其可移植性持久,高维向量近正交性是关键,还展示了损失景观几何视角。

详情
AI中文摘要

微调已广泛用于使大语言模型适应特定领域任务。参数高效微调(PEFT)方法如低秩适应(LoRA)常被用于降低计算成本。PortLLM是一种在持续预训练后用于适应大语言模型的无需训练和数据的方案。虽然初始结果显示LoRA补丁有短期时间可移植性,但PortLLM在多次持续预训练更新中的长期性能仍未充分探索,其有效性也缺乏理论理解。我们通过对PortLLM补丁在10个持续预训练步骤中的长期时间可移植性进行广泛实证研究,并提供两种理论分析来解决这两个问题。实证发现可移植性在更长时间持续存在,理论发现高维向量的近正交性是时间可移植性的关键依据,分析还展示了损失景观的几何视角以促进不同适应选项的理论比较。

英文摘要

Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computational costs. PortLLM is a training-free and data-free scheme used to adapt LLMs after continual pretraining. Although the initial PortLLM results show that LoRA patches exhibit short-term temporal portability, the long-term performance of PortLLM across several updates of continual pretraining remains underexplored. Furthermore, the intriguing effectiveness of PortLLM is not well understood from a theoretical standpoint. We address these two open questions by (1) performing an extensive empirical study of the long-term temporal portability of PortLLM patches across 10 continual pretraining steps using base models Mistral, Gemma, and Qwen; and (2) offering two theoretical analyses to explain our observation that the simple PortLLM method achieves competitive performance. We find empirically that the portability persists across longer time duration, indicating that repeated fine-tuning is not required when the base model is periodically updated. We find theoretically that near-orthogonality of high-dimensional vectors is a key justification for temporal portability. Our analyses also demonstrate a geometric perspective of the loss landscape in facilitating the theoretical comparison of different adaptation options.

URL PDF HTML 收藏
2607.20293 2026-07-23 cs.CV 新提交

Evolving Cache Schedules for Fast Diffusion Policy Inference

用于快速扩散策略推理的演进缓存调度

Siying Wang, Kangye Ji, Di Wang, Fei Cheng

机构 * School of Telecommunications Engineering, Xidian University(西安电子科技大学通信工程学院) School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

AI总结 研究针对扩散策略实时部署计算需求大的问题,提出演进缓存调度(EVO)框架,通过进化搜索全局调度缓存刷新,引入冗余感知初始化和目标条件早期停止,大幅减少计算量并保留性能,实现动作生成加速和FLOPs降低。

Comments 15 pages, 3 figures, supplementary material included. Accepted by PRCV 2026

详情
AI中文摘要

扩散策略通过迭代去噪动作块实现强大的视觉运动控制,但重复去噪使实时部署计算需求大。基于缓存的方法通过重用中间激活来降低推理成本,但现有无训练调度通常均匀分配计算,忽略块间异构冗余。我们引入演进缓存调度(EVO),通过进化搜索全局调度缓存刷新。它将候选者表示为块时间步晶格上的完整调度,可跳过冗余计算并保留闭环展开性能。还引入冗余感知初始化和目标条件早期停止。离线优化调度可直接插入预训练扩散策略。大量操纵基准表明,EVO在大幅减少计算的同时保留近全性能,动作生成加速高达8.05倍,FLOPs从15.77G降至1.96G。

英文摘要

Diffusion policies achieve strong visuomotor control by iteratively denoising action chunks, but repeated denoising makes real-time deployment computationally demanding. Cache-based methods reduce inference cost by reusing intermediate activations, but existing training-free schedules typically allocate computation uniformly across blocks, ignoring heterogeneous redundancy across blocks and leading to a suboptimal performance-efficiency trade-off. To bridge this gap, we introduce Evolving Cache Schedules (EVO), a training-free acceleration framework that globally schedules cache refreshes via evolutionary search. EVO represents each candidate as a complete schedule over the block-timestep lattice. Thus, redundant transformer computations during iterative denoising can be skipped through cache reuse while preserving closed-loop rollout performance. To make the search practical, EVO introduces redundancy-aware initialization, which seeds the population with promising schedules, and target-conditioned early stopping, which verifies and terminates once a desired performance target is reached. The offline-optimized schedule can be directly plugged into pretrained diffusion policies without retraining. Extensive manipulation benchmarks show that EVO preserves near-full performance while substantially reducing computation, achieving up to 8.05x action-generation speedup and reducing FLOPs from 15.77G to as low as 1.96G. Source code is available at this https URL.

URL PDF HTML 收藏
2607.20289 2026-07-23 cs.RO cs.AI 新提交

Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments

礼貌预期:改善持久共享环境中的长期任务规划

Md Ridwan Hossain Talukder, Roshan Dhakal, Elizabeth Phillips, Gregory J. Stein

机构 * George Mason University(乔治梅森大学)

AI总结 研究共享持久环境中机器人任务规划问题,提出礼貌预期规划方法,基于模型规划器联合最小化即时成本与预期未来成本,经独立估计器估计,在家庭和餐厅环境实验中相比其他规划降低了任务序列总成本。

Comments 9 Pages

详情
AI中文摘要

我们考虑一种任务规划场景,即共享持久环境的机器人从一个保留序列中依次被分配任务。标准任务规划器缺乏对未来任务的预见且不考虑其他机器人的约束,孤立地解决每个任务,留下会增加所有机器人未来成本的终端状态以及在长任务序列中复合的副作用。为降低序列成本,机器人必须预期其当前行动如何影响共享环境中所有机器人未来任务的性能。因此,我们提出礼貌预期规划,其中基于模型的规划器提出候选计划并选择能使所有机器人的即时成本和聚合预期未来成本联合最小化的计划,通过独立的每个机器人的学习估计器进行估计。这种分解式公式避免了组合式联合展开并支持模块化部署。我们在两个持久的PDDL领域进行评估,一个是具有相似能力但不同职责的机器人的家庭环境,另一个是机器人的不同能力会产生其他机器人无法解决的状态的餐厅环境。在长任务序列中,在双机器人家庭环境中,我们的规划器与近视规划相比总成本降低了10.43%,与自私预期规划相比降低了4.03%;在三机器人餐厅中,分别降低了17.41%和13.24%。

英文摘要

We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a time from a held-out sequence. Standard task planners, lacking foresight of future tasks and inconsiderate of others' constraints, solve each task in isolation, leaving terminal states that increase future cost for all, side effects that compound over lengthy task sequences. To reduce cost over the sequence, a robot must anticipate how its actions now may impact performance on future tasks for all robots sharing the environment. Therefore, we present courteous anticipatory planning, wherein a model-based planner proposes candidate plans and selects the one that jointly minimizes immediate cost and aggregated expected future cost across all robots, estimated via independent per-robot learned estimators. This factored formulation avoids combinatorial joint rollouts and supports modular deployment: adding a robot requires only training its own estimator. We evaluate in two persistent PDDL domains, a home environment with robots that have similar capabilities but different responsibilities, and a restaurant environment where robots' distinct capabilities create states that other robots lack the capability to resolve. During lengthy task sequences, our planner reduces total cost by 10.43% versus myopic and 4.03% versus selfish anticipatory planning in a two-robot home environment and by 17.41% and 13.24%, respectively, in a three-robot restaurant.

URL PDF HTML 收藏
2607.20286 2026-07-23 cs.CL cs.AI 新提交

Sound Probabilistic Safety Bounds for Large Language Models

大语言模型的可靠概率安全边界

Mahdi Nazeri, Anne-Kathrin Schmuck, Sadegh Soudjani, Alessandro Abate

机构 * University of Oxford(牛津大学) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所) University of Birmingham(伯明翰大学)

AI总结 研究提出框架计算大语言模型生成有害输出概率的严格边界,利用克洛普 - 皮尔逊置信区间,通过潜在空间特征优先探索有害分支,能高效计算可靠下界,实验验证方法有效性,为模型评估和认证提供新途径。

Comments The Initial version of this manuscript has been available on OpenReview, see this https URL (https://openreview.net/forum?id=papImkPLf5)

详情
AI中文摘要

我们提出了一个新颖的框架,用于计算大语言模型(LLM)针对给定提示生成有害输出的概率的严格边界。我们研究了克洛普 - 皮尔逊置信区间的新应用,以获得该问题的概率近似正确(PAC)边界。作为主要技术贡献,我们提出一种算法,利用潜在空间中的特征,优先探索自回归生成树中更可能产生有害输出的分支。我们的方法尤其能高效计算有用的下界,即便真实危害概率极小,且所获下界是可靠的,即经形式证明小于实际危害概率。实验结果通过计算现有先进LLMs的非平凡下界证明了方法的有效性。本研究为LLMs的评估和统计认证提供了新途径。

英文摘要

We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of useful lower bounds, even in scenarios where the true harm probability is extremely small, and crucially, the obtained lower bounds are sound, i.e., formally proven to be less than the actual harmfulness probability: our experimental results demonstrate the effectiveness of our method by computing non-trivial lower bounds on state-of-the-art LLMs. This study newly enables the evaluation and statistical certification of LLMs.

URL PDF HTML 收藏
2607.20284 2026-07-23 cs.CV 新提交

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

用于遥感图像理解的多模态大语言模型:领域特定还是通用?

Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室) Yuelushan Center for Industrial Innovation(岳麓山工业创新中心)

AI总结 本文针对遥感图像理解的多模态大语言模型展开研究,通过系统调查和评估,比较其与通用模型在不同任务上的表现,发现当前模型存在局限,进而给出未来方向,为开发相关模型提供系统参考。

Comments 27 pages, 11 figures

详情
AI中文摘要

多模态大语言模型(MLLMs)的快速发展为遥感图像场景理解(RSISU)带来了灵活范式,实现与遥感图像的自然语言交互。然而,对现有遥感MLLMs(RS - MLLMs)的能力边界、跨任务泛化和特定任务局限性仍缺乏系统理解。本文对用于RSISU的MLLMs进行系统调查和诊断评估。回顾RS - MLLMs技术演变,关注模型设计等。比较RS - MLLMs与通用计算机视觉MLLMs(CV - MLLMs)在不同RSISU任务和基准上的表现。发现RS - MLLMs在特定领域有竞争力,通用CV - MLLMs在一些任务上可匹配甚至超越。当前MLLMs在空间和关系推理等方面也有局限。基于此给出未来方向,为开发用于RSISU的强大、通用且实用的MLLMs提供系统参考。

英文摘要

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.

URL PDF HTML 收藏
2607.20274 2026-07-23 cs.CV cs.AI cs.CL cs.LG 新提交

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

自我监督比临床监督更能推动医学基础模型中的表征趋同

Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn

机构 * RWTH Aachen University(亚琛工业大学) University Hospital RWTH Aachen(亚琛工业大学附属医院) Technical University of Munich(慕尼黑工业大学) Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡弗里德里希-亚历山大大学) Technical University Dresden(德累斯顿工业大学) University Hospital Dresden(德累斯顿大学附属医院) University Hospital Heidelberg(海德堡大学附属医院)

AI总结 研究探讨医学基础模型中表征趋同问题,通过对多个编码器剖析发现自我监督比临床监督更能驱动收敛,虽收敛有限但线性分类器可跨编码器转移,表明医学成像收敛由预训练目标决定,为互操作性设计与验证提供依据。

详情
AI中文摘要

不同团队的医学图像编码器越来越被视为可互换的,人们认为规模和临床监督会将其表征集中到一个共享结构上。但这种趋同是否真实、其产生原因以及是否具有临床可用性都未经检验,且相关相似性度量很脆弱。我们对18个图像和7个文本编码器进行了可控剖析,涵盖多种参数规模和成像模态。结果表明,收敛适度但高于随机水平,是由自我监督目标驱动,而非临床监督。线性分类器可跨编码器转移并应用于五家医院,保留约85%的编码器内性能。因此,医学成像中的收敛由预训练目标决定,而非规模或临床监督。应通过该目标设计互操作性,并在共享几何结构最薄弱的地方进行验证。

英文摘要

Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure. Whether this convergence is real, what produces it, and whether it is clinically usable are untested, and the similarity measures behind such claims are fragile. We present a controlled dissection across 18 image and 7 text encoders, all open-weight and run locally, spanning 7M to 27B parameters and five imaging modalities, including 650,982 chest radiographs from six datasets. To isolate cause, we train encoders that vary only the objective under fixed data, architecture, and scale, and reproduce the effect in a synthetic model. Convergence is modest but above a random floor, driven by the self-supervised objective, not clinical supervision: matched self-supervised encoders aligned most (40.4% on chest radiography), with label-supervised (21.1%) and image-text (3.3%) far lower, and did not grow with size (Spearman 0.302, p=0.223) or capability. It is within-modality, does not reach clinical language, and does not reproduce how radiologists judge case similarity. Yet a linear classifier transfers across encoders and to five held-out hospitals, retaining about 85% of within-encoder performance. Convergence in medical imaging is therefore set by the pretraining objective, not inherited from scale or clinical supervision. Interoperability is accordingly something to design for through that objective, and to validate where the shared geometry is weakest, across patient subgroups and against clinical judgment.

URL PDF HTML 收藏
2607.20270 2026-07-23 cs.CL 新提交

Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study

大语言模型混淆了哪些价值观?基于施瓦茨的识别研究

Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Mikhail Solovev, Valeriia Kuschenko, Maria Chistyakova, Sergey Bolovtsov

机构 * Ivannikov Institute for System Programming of the Russian Academy of Sciences(俄罗斯科学院伊万尼科夫系统编程研究所) Russian Presidential Academy of National Economy and Public Administration(俄罗斯总统国民经济与公共管理学院)

AI总结 研究大语言模型对施瓦茨十个基本价值观的识别,通过俄语情境文本评估 21 个指令微调的模型运行,分析准确率、混淆情况等,推动结合精确准确率、排序恢复和有向错误分析的价值识别评估。

Comments 14 pages, 7 figures, 3 tables

详情
AI中文摘要

大语言模型越来越多地通过它们认可的价值观来评估,但这种评估预先假定模型能够识别具体情境中表达的价值观。我们将此前提研究为对施瓦茨的十个基本价值观的受控 top-1 识别。我们的评估集包含 1000 篇俄语情境文本,在十个价值观上保持平衡,每个项目由两名人类注释者独立标注。我们在固定的排序响应协议下评估 21 个指令微调的大语言模型运行;20 个具有可靠输出的运行形成语义面板。汇总的 Acc@1 为 0.683,Acc@3 为 0.892,表明模型在定位正确动机区域时,对相近替代选项的排序不稳定。相邻价值观占语义错误的 50.9%,而特定检查点的空值下为 24.4%。八个有向混淆在检查点和人类确认的子集中反复出现。结果推动了结合精确准确率、排序恢复和有向错误分析的价值识别评估。

英文摘要

Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values. Our evaluation set contains 1,000 Russian situational texts, balanced across the ten values and independently labeled by two human annotators per item. We evaluate 21 instruction-tuned LLM runs under a fixed ranked-response protocol; 20 runs with reliable outputs form the semantic panel. Pooled Acc@1 is 0.683 and Acc@3 is 0.892, showing that models often locate the correct motivational region while ranking close alternatives unstably. Adjacent values account for 50.9% of semantic errors, compared with 24.4% under a checkpoint-specific null. Eight directed confusions recur across checkpoints and human-confirmed subsets. Several are strongly asymmetric, including Universalism to Benevolence, Tradition to Conformity, and Security to Power, whereas Stimulation-Hedonism forms a bidirectional boundary. Their severity is checkpoint-specific and can bias higher-order value profiles. The results motivate value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.

URL PDF HTML 收藏
2607.20265 2026-07-23 cs.CL cs.AI 新提交

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

可掩码性指数:预测预训练语言模型中的任务目标对齐

Ahmad Pouramini, Mahsa Afsharzadeh

机构 * Sirjan University of Technology(锡尔詹理工大学) Department of Computer Engineering(计算机工程系)

AI总结 研究提出可掩码性指数(MI),用于评估知识关系适合的提示方式,通过掩码与未掩码模板的DepthRank分数差异计算。在ATOMIC2020基准上评估发现MI与下游生成性能正相关,有助于选择提示模板和策略提取预训练语言模型的关系知识。

详情
AI中文摘要

大规模预训练语言模型如T5和BERT在生成结构化知识方面展现出强大能力,但其性能取决于提示策略与预训练目标的匹配程度。我们引入可掩码性指数(MI),这是一种定量指标,用于估计在少样本生成中知识关系更适合掩码式提示还是前缀式提示。MI通过掩码和未掩码模板之间的DepthRank分数差异计算得出,可衡量目标与模板的对齐程度。我们在ATOMIC2020知识库完成基准的各种关系上评估MI,结果表明它与下游生成性能呈正相关。这些结果表明MI可帮助选择合适的提示模板和适应策略,从预训练语言模型中提取关系知识,特别是在低资源环境中。

英文摘要

Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used during pretraining. We introduce the Maskability Index (MI), a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot generation. MI is computed from differences in DepthRank scores between masked and unmasked templates, providing a principled measure of objective-template alignment. We evaluate MI on a diverse set of relations from the ATOMIC2020 knowledge base completion benchmark and show that it is positively correlated with downstream generation performance. These results indicate that MI can help select appropriate prompting templates and adaptation strategies for extracting relational knowledge from pretrained language models, especially in low-resource settings.

URL PDF HTML 收藏
2607.20263 2026-07-23 cs.CV 新提交

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

城市环境与住宅建筑健康有何关联?一种用于建筑层面房屋检查的视觉与兴趣点融合框架

Kun Zhao, Helei Ren, Guilin Tang, Tianyi Chen, Zhehui Song, Xing Liu, Lijian Zhou, Yuhong Zhao, Xiang Gao, Jinming Jiang, Qichao Ban

机构 * School of Information Management, Qingdao University of Technology(青岛理工大学信息管理学院) Embodied AI & Robot Research Institute, Qingdao University of Technology(青岛理工大学具身人工智能与机器人研究所) College of Architecture and Urban Planning, Qingdao University of Technology(青岛理工大学建筑与城乡规划学院) Innovation Institute for Sustainable Maritime Architecture Research and Technology (iSMART), Qingdao University of Technology(青岛理工大学可持续海洋建筑研究与技术创新研究所) Department of Urban Planning and Design, Xi’an Jiaotong-Liverpool University(西交利物浦大学城市规划与设计系)

AI总结 研究探讨城市环境与住宅建筑健康的关联,提出视觉与兴趣点融合框架,通过多视图视觉检查与兴趣点邻里环境结合评估建筑健康。经实验,多视图聚合提升性能,兴趣点上下文补充信息,提高了建筑层面宏F1分数。

详情
AI中文摘要

房屋层面的城市体检对于识别住宅建筑问题和支持有针对性的城市更新至关重要。现有自动化检查研究主要依赖单个图像,很少研究周围城市功能环境能否为建筑层面评估提供补充信息。本研究提出了一种视觉与兴趣点融合框架,将多视图视觉检查与从兴趣点得出的邻里环境相结合,用于住宅建筑健康评估。实证数据集涵盖中国青岛的92个老旧住宅小区、3237栋住宅建筑和25608张实地采集的检查图像,包含七类与住房相关的问题。首先,评估多个目标检测模型以从单个图像中提取问题位置、类别和置信度分数,然后将图像层面的输出跨多个视图聚合以构建可解释的建筑层面表示。其次,在500米、1000米和1500米的邻里缓冲区中提取兴趣点特征以表征周围功能环境,使用皮尔逊和斯皮尔曼相关性分析并结合错误发现率校正来识别候选上下文特征。最后,在社区隔离空间交叉验证下,使用成本敏感随机森林分类器整合视觉和兴趣点特征。结果表明,多视图聚合带来了主要性能提升,将建筑层面的宏F1从直接检测下的60.84%提高到74.95%。纳入兴趣点上下文进一步将宏F1提高到76.79%,尽管额外增益不大且与类别相关。因此,兴趣点信息起到补充上下文先验的作用,而非直接视觉证据的替代品或建筑状况的因果决定因素。

英文摘要

Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment. This study proposes a vision-POI fusion framework that combines multi-view visual inspection with POI-derived neighborhood context for residential building health assessment. The empirical dataset covers 92 old residential communities, 3,237 residential buildings, and 25,608 field-acquired inspection images in Qingdao, China, encompassing seven categories of housing-related issues. First, multiple object detection models are evaluated to extract issue locations, categories, and confidence scores from individual images. The image-level outputs are subsequently aggregated across multiple views to construct interpretable building-level representations. Second, POI features are extracted within 500m, 1,000m, and 1,500m neighborhood buffers to characterize surrounding functional environments. Pearson and Spearman correlation analyses, combined with false discovery rate correction, are used to identify candidate contextual features. Finally, visual and POI features are integrated using a cost-sensitive Random Forest classifier under community-isolated spatial cross-validation. The results show that multi-view aggregation provides the main performance improvement, increasing the building-level Macro-F1 from 60.84% under Direct Detection to 74.95%. Incorporating POI context further increases Macro-F1 to 76.79%, although the additional gain is modest and category-dependent. POI information therefore functions as a supplementary contextual prior rather than a substitute for direct visual evidence or a causal determinant of building condition.

URL PDF HTML 收藏
2607.20253 2026-07-23 cs.SD cs.AI eess.AS 新提交

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

推动全曲生成的前沿:分层自回归规划与流匹配渲染

Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu

机构 * Alibaba(阿里巴巴)

AI总结 该研究提出统一歌曲生成框架,支持多项任务,由四个组件构成。通过特定编码、建模、匹配及模块实现歌曲生成,还研究多种后训练策略,实验表明该框架在评估中性能具有竞争力。

详情
AI中文摘要

在本报告中,我们提出了一个统一的歌曲生成框架,能够从歌词、文本描述和音乐属性中生成高质量的全长音乐。该框架支持三项任务:歌词到歌曲生成、器乐音乐生成和翻唱歌曲生成。系统架构上由语义感知分词器、hybird-LM、FullDiT和两级旋律模块四个主要组件组成。分词器将音频编码为8码本RVQ令牌,hybird-LM进行分层自回归音频令牌建模,FullDiT在连续VAE潜在空间中进行全曲流匹配,旋律模块用于翻唱歌曲生成。最后研究了DPO、GRPO和OPD作为基于奖励的后训练策略,并将基于流的GRPO应用于FullDiT。实验结果表明该框架在评估设置中取得了有竞争力的性能。

英文摘要

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module. The tokenizer encodes audio into 8-codebook RVQ tokens for efficient discrete music representation. Based on these tokens, hybird-LM performs hierarchical autoregressive audio-token modeling for full-song generation. To improve audio fidelity, FullDiT performs full-song flow matching in a continuous VAE latent space conditioned on codec tokens, lyrics, and text captions. For cover song generation, the melody module extracts and discretizes melody cues from reference audio to guide generation while preserving the original melodic content. Finally, we investigate DPO, GRPO, and OPD as reward-based post-training strategies for hybird-LM and apply flow-based GRPO to FullDiT to improve musicality and rendering quality. Experimental results on a multilingual automatic benchmark, complemented by the Artificial Analysis Music with Vocals leaderboard, show that the proposed framework achieves competitive performance in the evaluated settings.

URL PDF HTML 收藏
2607.20251 2026-07-23 cs.CL 新提交

Exposure is Optional: Learning Unlike Coordination in Language Models

曝光非必需:语言模型中学习不同类别的并列结构

Jiamu Luo, Shane Steinert-Threlkeld

机构 * University of Washington(华盛顿大学)

AI总结 研究语言模型中不同类并列结构习得是否需直接曝光,用过滤语料库训练GPT-2模型,发现无需直接曝光,模型能泛化处理,还揭示了模型处理方式及可从同类并列结构学习,助力理解语言模型结构表示。

Comments 13 pages, 6 tables, 2 figures, to submit to TACL

详情
AI中文摘要

并列结构作为一种基本的语言结构,仍然是激烈辩论的主题,其确切性质仍然困扰着理论语言学。一种常见观点认为只有同类成分才能并列,这一观点受到自然语言中许多不同类并列结构的挑战。我们将语言模型视为计算测试平台,研究不同类并列结构的习得是否需要训练数据中的直接曝光,或者它是否可以从一般的组合能力中有机出现。我们使用过滤语料库训练(FiCT),在去除所有不同类并列结构实例的语料库上训练GPT - 2模型。我们发现直接曝光并非必要:在过滤数据上训练的模型成功地将不同类并列结构进行了泛化,在困惑度和语法判断上与在未过滤文本上训练的模型相当。此外,我们对内部表示的分析表明,语言模型通过将并列元素视为属于相似结构类别或通过类似删除的机制来处理不同类并列结构,这两者似乎都可以仅从接触同类并列结构中学习。这项工作有助于加深对语言模型如何在内部表示语言结构的理解,同时也通过展示模型在没有直接曝光的情况下如何泛化和处理不同类并列结构,为关于并列结构的更广泛辩论增添了内容。

英文摘要

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language. Treating language models as a computational testbed, we investigate whether the acquisition of unlike coordination requires direct exposure in the training data, or whether it can emerge organically from general compositional abilities. Using Filtered-Corpus Training (FiCT), we train GPT-2 models on corpora from which all instances of unlike coordination have been removed. We find that direct exposure is not necessary: models trained on filtered data successfully generalize to unlike coordination, achieving perplexity and grammaticality judgments comparable to models trained on unfiltered text. Furthermore, our analyses of internal representations indicate that language models process unlike coordination by treating the conjoined elements as belonging to similar structural categories or through a mechanism akin to deletion, both of which appear learnable from exposure to alike coordination alone. This work contributes to the growing understanding of how language models internally represent linguistic structures, while also adding to the broader debate on coordination by showing how models generalize and process unlike coordination without direct exposure.

URL PDF HTML 收藏
2607.20247 2026-07-23 cs.CV 新提交

Vera: Identity-Faithful Human Subject-to-Video Generation

Vera:身份忠实的人类主体到视频生成

Yulong Xu, Xinyue Liu, Shujuan Li, huafeng shi, Yan Zhou, Jiwen Liu, Xintao Wang, Yu Shen Liu, Huaibo Huang

机构 * Kuaishou Technology(快手科技) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学)

AI总结 针对主体到视频生成中人类身份一致性不足问题,提出Vera框架,通过构建百万对身份对齐数据集,引入身份聚焦掩码监督和参考感知分层注意力两种设计,提升了身份一致性、主体绑定及运动自然性等。

详情
AI中文摘要

主体到视频(S2V)生成在跨不同类别保留参考主体方面取得了重大进展,但以人类为中心的生成中通用的主体一致性仍然不足。在多人场景中问题更严重,身份角色绑定错误会导致主体混淆等。我们提出了Vera,一个用于单人和多人生成的统一的以人类为中心的S2V框架。首先通过人物级跨剪辑检索构建了一个百万对身份对齐的人类图像 - 视频数据集。在此数据集基础上,Vera引入了两种互补设计。身份聚焦掩码监督(IFMS)加强身份感知学习,参考感知分层注意力(RALA)调节视频令牌与参考身份线索的交互。大量实验表明,Vera提高了人类身份一致性、多人主体绑定和运动自然性,同时减少了身份混淆和过度的参考图像复制。

英文摘要

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.

URL PDF HTML 收藏
2607.20241 2026-07-23 cs.CL cs.AI 新提交

On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens

论文化负载机器翻译的系统性挑战:以《红楼梦》为文化视角

Yiming Wang, Jiayuan Di

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系) School of Foreign Languages, East China University of Science and Technology(华东理工大学外国语学院) Mejiro University(目白大学)

AI总结 研究以《红楼梦》构建数据集,系统探究基于大语言模型的机器翻译系统中文化负载翻译的挑战,揭示任务、人工评估及自动评估方面的问题,为文化导向的翻译研究提供见解。

详情
AI中文摘要

文化负载翻译给机器翻译带来独特挑战,因为意义深深嵌入社会文化语境而非表面语言形式。虽大语言模型使机器翻译系统在许多场景达到类人质量,但其处理文化负载表达的能力仍未充分探索。本研究系统调查基于大语言模型的机器翻译系统中文化负载翻译带来的挑战。我们从具有文化代表性的《红楼梦》语料库构建中日双语数据集,含500个不同文化类别的片段。通过综合评估协议,揭示三个主要挑战:前沿大语言模型在文化负载内容上表现有差距;评估者背景导致翻译判断存在重大分歧;广泛使用的指标无法可靠评估此任务的翻译质量。这些发现可为计算科学和语言学中面向文化的翻译研究提供有价值的见解。

英文摘要

Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally loaded translation in LLM-based MT systems. We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of the Red Chamber, containing 500 segments across diverse cultural categories. Using a comprehensive evaluation protocol, we reveal three main challenges: (1) task challenges, where frontier LLMs exhibit notable performance gaps and struggle with culturally loaded content; (2) human evaluation challenges, where evaluator backgrounds lead to substantial disagreement in translation judgments; and (3) automatic evaluation challenges, where widely used metrics fail to reliably assess translation quality for this task. These findings may offer valuable insights for culture-oriented translation research in both computational science and linguistics.

URL PDF HTML 收藏
2607.20238 2026-07-23 cs.CV 新提交

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

并非所有补丁都一样:可见-红外预训练中的采样很重要

Qiwei Ma, Bin Deng, Junjie Zhu, Qiangjuan Huang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li

机构 * Yuelushan Center for Industrial Innovation(岳麓山工业创新中心) Intelligent Game and Decision Lab(智能游戏与决策实验室)

AI总结 研究针对可见-红外预训练,提出重要性感知采样(IAS)方法,通过从红外结构线索导出权重、学习重要性掩码及采用补丁课程学习策略调整训练重点,在多个VIS-IR基准实验中比基线有改进。

Comments 13 pages, 11 figures,

详情
AI中文摘要

可见-红外(VIS-IR)对齐是强大的多传感器感知的关键预训练任务。大多数现有方法使用均匀的逐补丁对比学习,但在VIS-IR数据中这可能不可靠,因为成像物理差异使一些空间配对区域本质上可比性较低,同等强度对齐会阻碍表示学习和下游迁移。本文从采样角度重新审视VIS-IR预训练,提出重要性感知采样(IAS),它基于补丁可靠性调整训练重点。具体包括从红外结构线索导出补丁权重来重新加权对比目标;用轻量级采样器学习软重要性掩码,可从手工先验热启动;采用从高可靠性区域到更难补丁逐步扩展的补丁课程学习策略。IAS是即插即用的,与逐补丁/相关性级对齐和图像级对比基线都兼容。在多个VIS-IR基准上的广泛实验表明,相对于强大基线有持续改进,包括红外语义分割、红外目标检测、VIS语义分割和跨模态检索任务。代码将在该https网址发布。

英文摘要

Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can be unreliable in VIS-IR data because imaging-physics differences make some spatially paired regions inherently less comparable, and aligning them with equal strength hinders representation learning and downstream transfer. In this paper, we revisit VIS-IR pre-training from a sampling perspective and propose Importance-Aware Sampling (IAS), which adjusts training emphasis based on patch reliability. Specifically, IAS (i) derives patch weights from infrared structural cues and uses them to reweight the contrastive objective; (ii) learns a soft importance mask with a lightweight sampler, optionally warm-started from the hand-crafted prior; and (iii) employs a patch curriculum learning strategy that gradually expands from high-reliability regions to harder patches. It is worth noting that IAS is plug-and-play and works with both patch-/correlation-level alignment (e.g., UNIV-style) and image-level contrastive baselines (e.g., ImageBind-style). Extensive experiments on multiple VIS-IR benchmarks demonstrate consistent improvements over strong baselines, including for IR semantic segmentation, IR object detection and VIS semantic segmentation and cross-modal retrieval task. Code will be released on this https URL.

URL PDF HTML 收藏