arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

2026-08-11 至 2026-08-11 共收录 11
2608.09836 2026-08-11 cs.AI cs.CL 新提交

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

失配很重要:超越token一致性的在线蒸馏

Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou

机构 * The University of Hong Kong(香港大学) University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文针对在线蒸馏存在的退化一致性失效问题,提出TIDE方法修正师生失配,在数学推理基准上显著提升性能、缩短响应长度并减少格式错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09745 2026-08-11 cs.LG cs.AI stat.ML 新提交

SR-OPSD: Self-Referenced On-Policy Self-Distillation

SR-OPSD:自引用同策略自蒸馏

Zhuo Sun, Entong Li, Yanlong Zhao, Xiaoyuan Cheng, Wenxuan Yuan, Kaiyu Li, Che Liu, Huihang Liu, Harrison Bo Hua Zhu, Li Zeng

机构 * Shanghai University of Finance and Economics(上海财经大学) Imperial College London(伦敦帝国学院) University of Science and Technology of China(中国科学技术大学) University College London(伦敦大学学院) Nanyang Technological University(南洋理工大学) Technical University of Denmark(丹麦技术大学) University of Copenhagen(哥本哈根大学) Peking University(北京大学)

AI总结 该研究针对同策略自蒸馏的优化不稳定问题,提出SR-OPSD方法,通过几何插值结合Rényi散度优化蒸馏目标,在多任务大语言模型实验中取得最优或竞争性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09550 2026-08-11 cs.CV 新提交

PressureMesh: 3D Human Mesh Estimation from Multi-Device Pressure Images

PressureMesh:基于多设备压力图像的3D人体网格估计

Changhai Ma, Ziyu Wu, Yunkang Zhang, Fangting Xie, Mengting Niu, Heyu Ding, Quan Wan, Jiayue Yuan, Boyan Liu, Yi Ke, Xiaohui Cai

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 针对单设备压力监测范围有限的问题,提出端到端网络MDP-Net,结合MoE框架的多模态融合机制,构建MDP数据集,实现跨设备压力数据的3D人体网格估计,关节位置误差为12.6cm。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09529 2026-08-11 cs.CV 新提交

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

面向表达性与忠实性的音频到图像生成:一个统一的多模态数据集与合成框架

Dongxu Ge, Shansong Liu, Cheng Gong, Xiao-Lei Zhang, Chi Zhang, Xuelong Li

机构 * University of Science and Technology of China(中国科学技术大学) China Telecom (TeleAI)(中国电信(电信人工智能研究院)) Northwest Polytechnical University(西北工业大学)

AI总结 针对音频到图像生成受限于传统数据集的问题,提出A2I-Set数据集与AudioCanvas模型,实现了更优的跨模态对齐与视觉表达性。

Comments 23 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09244 2026-08-11 cs.CV 新提交

In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

用于高保真主体驱动文本到图像生成的耦合潜在噪声引导的循环内模型适配

Yushun Tang, Weiming Chen, Siyi Liu, Yi Zhang, Feng Wu, Zhihai He

机构 * Southern University of Science and Technology(南方科技大学) University of Science and Technology of China(中国科学技术大学) Pengcheng Lab(鹏城实验室)

AI总结 本研究针对主体驱动文本到图像生成中参考图像变化时模型难适配且难保持主体身份的问题,提出IMA方法,通过耦合潜在噪声损失引导循环内模型适配,提升了生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09168 2026-08-11 cs.AI 新提交

From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents

从相关性到执行效用:面向基于技能的大语言模型智能体的奖励感知动态执行门控

Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang

机构 * Tongji University(同济大学) The University of Sydney(悉尼大学) University of Science and Technology of China(中国科学技术大学) Nankai University(南开大学) Johns Hopkins University(约翰斯·霍普金斯大学) City University of Hong Kong(香港城市大学) University of Surrey(萨里大学)

AI总结 针对基于技能的LLM智能体执行决策难题,提出RADEG轻量型决策层,通过学习代理模型预测执行效用,在288个rollout的评估中减少不必要执行并优于基线方法

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08491 2026-08-11 cs.AI 新提交

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

TrustRoboReward:面向多范式机器人奖励模型的偏好有序保序分数编辑方法

Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang

机构 * Peking University(北京大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) University of Science and Technology of China(中国科学技术大学) Southeast University(东南大学) Southern University of Science and Technology(南方科技大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Beijing Language and Culture University(北京语言大学) Sichuan University(四川大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 针对现有机器人奖励模型的跨范式偏好与分数不一致问题,本文提出 TrustRoboReward 框架,通过 POISE 方法解决反转冲突,训练的 Qwen3-VL-4B 性能接近 GPT-5-mini,优于 RoboReward 基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08471 2026-08-11 cs.AI 新提交

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

昨日之盾,今日之矛:生产环境中的自演进安全护栏

Cong Ming, Jingyi Chen, Bin Liu, Qi Chu, Tao Gong, Nenghai Yu, Yingfei Xiang

机构 * University of Science and Technology of China(中国科学技术大学) Shenzhen University(深圳大学) Sangfor Technologies(深信服科技)

AI总结 针对静态LLM安全护栏无法应对新威胁的问题,提出自演进安全护栏SESG多智能体系统,可快速适配新威胁,性能优于静态护栏及自适应基线,已应用于深信服护栏并取得良好效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08236 2026-08-11 cs.AI cs.CL 新提交

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems

LatticeMind:面向多智能体系统的冲突感知记忆原语

Heng Zhou, Lian Zhang, Yutao Fan, Tiancheng He, Siki Chen, Hejia Geng, Philip Torr, Zhenfei Yin

机构 * University of Science and Technology of China(中国科学技术大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Shanghai AI Laboratory(上海人工智能实验室) Beijing University of Posts and Telecommunications(北京邮电大学) University of Oxford(牛津大学)

AI总结 本研究提出冲突感知记忆原语LatticeMind,解决多智能体LLM系统的主张信任决策问题,在ConflictBank评估中准确率达0.97,显著优于基线, ablation验证了其核心组件的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08103 2026-08-11 cs.LG 新提交

Support Selection Beyond Smooth DAG Exactness: Completion Geometry,Score Margins, and Selective Certificates

超越平滑DAG精确性的支持选择:完备几何、得分边际与选择性证书

Rui Wu, Zongyuan Chen, Hong Xie

机构 * School of Computer Science and Engineering, University of Science and Technology of China(中国科学技术大学计算机科学与工程学院)

AI总结 该研究针对超越平滑DAG精确性的支持选择,推导了孤立环精确选择时间,验证了相关规律,提出无真值分离统计量预测选择时间,用父集置信族等证明标签,区分DAG可行性等内容。

Comments 49 pages, 17 figures, 20 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07987 2026-08-11 cs.CV 新提交

Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence

优势引导门:重塑基于视觉的空间智能的开放式推理

Ling Lin, Yang Bai, Congcong Zhu, Jiangming Shi, Meng Wang, Yang Long, Jingrun Chen, Ling Shao, Huazhu Fu

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute for Advanced Research, USTC(中国科学技术大学苏州高等研究院) Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高级智能与计算研究所) East China Normal University(华东师范大学) National University of Singapore(新加坡国立大学) Durham University(杜伦大学) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 本文针对MLLMs开放式推理易出错的问题,提出优势引导门框架,结合蒙特卡洛价值评估与两类优势门,构建Reasoning-Tree-160k数据集并经两阶段学习,有效提升了基准MLLMs的视觉空间推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏