arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Zhejiang University(浙江大学)

2026-08-14 至 2026-08-14 共收录 15
2608.13552 2026-08-14 cs.CV 新提交

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

PlayWorld:基于智能体玩家的长程目标世界模型基准测试

Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhejiang University(浙江大学) Kuaishou Technology(快手科技)

AI总结 该研究针对现有世界模型跨模型公平比较的难题,推出含171个场景的PlayWorld基准,通过多模态智能体玩家从多维度评估9种先进世界模型,发现其长程交互式目标表现仍不可靠。

Comments project page: https://kxding.github.io/project/PlayWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13284 2026-08-14 cs.RO 新提交

Predictive Relative-Velocity Steering for Safe Robotic Manipulator Teleoperation in Dynamic Environments

动态环境下安全机器人操纵器遥操作的预测相对速度转向方法

Changhao Hu, Zeyi Liu, Songqiao Hu, Shuang Liu, Zihan Meng, Xiao He

机构 * Zhejiang University(浙江大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Tsinghua University(清华大学) Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学具身智能与机器人研究所) TetraBOT

AI总结 针对动态环境下机器人遥操作的安全问题,提出轻量级模块化预测相对速度转向框架,可提升末端执行器避障率并缓解死锁,仿真与物理实验验证了其有效性。

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13120 2026-08-14 cs.AI 新提交

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

SkillEvo:基于多轮交互反馈的自更新进化梯度

Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu

机构 * Tencent Cloud Andon(腾讯云Andon) Zhejiang University(浙江大学)

AI总结 SkillEvo通过多轮用户模拟生成反馈、独立治理层修复退化,在6类云服务等数据集上,相比两种基线方法显著提升了智能体技能进化效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13028 2026-08-14 cs.CV cs.RO 新提交

RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction

用于改进人机物体交接预测的RGB-D视频生成

Tianyu Sun, Zhoujie Fu, Zihui Gao, Bang Zhang, Guosheng Lin

机构 * Nanyang Technological University(南洋理工大学) College of Computing and Data Science(计算与数据科学学院) Zhejiang University(浙江大学) College of Computer Science and Technology(计算机科学与技术学院) Alibaba Group(阿里巴巴集团)

AI总结 针对人机物体交接数据集稀缺及现实-模拟差距问题,提出Hand2Bot数据集与PassGen生成流水线,实现高意图识别准确率与低误触发率,支持机器人稳健零样本迁移与早期意图预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12951 2026-08-14 cs.SD 新提交

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

VoxAudio:基于多奖励自回归流匹配的发声音频合成

Wenxiang Guo, Changhao Pan, Ziyue Jiang, Fei Wu, Zhou Zhao

机构 * Zhejiang University(浙江大学)

AI总结 VoxAudio是一种多奖励自回归流匹配模型,通过架构、偏好、数据层面的优化,解决现有T2A系统发声控制不足的问题,在多基准上验证了有效性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12939 2026-08-14 cs.LG 新提交

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

用动作条件预测一致性诊断JEPA世界模型

Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian

机构 * Huawei(华为) University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学) Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))

AI总结 本研究针对JEPAs世界模型易受视觉扰动影响的问题,提出动作条件预测一致性(ACPC)诊断方法,定义IR与SR指标,经四视觉控制任务实验验证其可预测扰动带来的预测及代价变化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12904 2026-08-14 cs.CV 新提交

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation

HounsWorld:用于隐藏患者状态读取、重建与模拟的多模态世界模型

Yunhao Bai, Zhongwei Qiu, Guangyu Guo, Yiming Huang, Tony C. W. Mok, Qinji Yu, Ling Zhang, Yan Wang

机构 * East China Normal University(华东师范大学) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) Hupan Laboratory(湖畔实验室) Zhejiang University(浙江大学)

AI总结 该研究提出3B参数多模态世界模型HounsWorld,结合CT扫描与临床语言,通过共享潜在患者状态实现读取、重建、模拟三类任务,在HounsBench基准上表现优异,提升了CT理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12780 2026-08-14 cs.CV 新提交

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention

SCOPE:用于稀疏视频注意力的在线分头Top-K估计的子空间聚类

Qi Zhao, Qirui Li, Hanlin Tang, Yiduo Li, Zhen Guo, Cuifeng Shen, Chao Xu, Zhaosheng Chi, Xiaojin Lu, Kan Liu, Tao Lan, Lin Qu, Xi Li

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

AI总结 SCOPE是一种无训练稀疏注意力框架,通过子空间聚类与在线分头Top-k估计解决视频DiTs的高成本问题,在6种配置中优于基线,实现1.99倍加速且保持28.46 dB PSNR。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12743 2026-08-14 cs.AI 新提交

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

空间记忆智能体:用于空间智能的基于经验的过程记忆

Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen

机构 * Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

AI总结 该研究提出SMA框架,让冻结VLM智能体无需外部空间工具,通过无参数更新自进化提升空间推理,在多基准和模型上表现最优,提供了空间自进化的实用路径。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12590 2026-08-14 cs.AI cs.CV 新提交

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

用于循证甲状腺超声诊断与报告的可审计智能体AI

Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Software Engineering, Sun Yat-sen University(中山大学软件工程学院) School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院) Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩雷尔曼医学院) Zhujiang Hospital, Southern Medical University(南方医科大学珠江医院) College of Mathematical Medicine, Zhejiang Normal University(浙江师范大学数学医学院)

AI总结 本文提出可审计的智能体AI系统ThyroidXAgent,基于多中心数据集开发,可协同完成甲状腺超声诊断多任务,提升分类准确率与报告一致性,缩短耗时,支持临床医生修正。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10433 2026-08-14 cs.LG 版本更新

From Recoverability to Functional Use: Certifying Temporal Reports in Time-Series Forecasting

时间序列预测器是否使用了正确的历史信息:时间延迟的可恢复性、恢复及功能使用

Qipeng Qian, Yuntao Qian

机构 * Supcon Technology(中控技术) College of Artificial Intelligence, Zhejiang University(浙江大学人工智能学院)

AI总结 该研究针对时间序列模型,分析了延迟可恢复性、报告与功能使用的问题,发现多数模型存在延迟报告正确但实际未使用对应历史的情况,提出的路由方法可实现预测与报告的对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09233 2026-08-14 cs.LG cs.CV 版本更新

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

DreOPD:用于流匹配模型的退化参考外推式在线策略蒸馏

Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Zhejiang University(浙江大学) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 本研究提出用于流匹配模型的DreOPD方法,将隐式奖励外推转为闭式速度回归并引入退化参考强化对比,在单/多教师设置下,其平均性能优于OPD及多任务RL基线,多数指标超专用教师。

Comments project page: https://sleepy1231.github.io/DreOPD/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03089 2026-08-14 cs.LG cs.AI 版本更新

Constitutional On-Policy Safe Distillation

宪法性在策略安全蒸馏

Ming Wen, Yuxuan Liu, Kun Yang, Yunhao Feng, Zhuoer Xu, Yuhao Sun, Shiwen Cui, Xiang Zheng, Yi Liu, Xingjun Ma, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Ant Group(蚂蚁集团) Zhejiang University(浙江大学) City University of Hong Kong(香港城市大学)

AI总结 针对在策略自蒸馏在安全对齐中因宪法条件导致教师分布收缩、表达能力下降的问题,提出宪法性在策略安全蒸馏(COPSD),通过交叉SFT冷启动校准教师分布,再进行宪法条件在策略蒸馏,在12个基准上实现了更优的安全-有用性权衡并降低安全税。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15489 2026-08-14 cs.NI cs.AI 版本更新

A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network

基于Q学习的面向服务质量的多路径路由协议

Mehdi Hosseinzadeh, Roohallah Alizadehsani, Amin Beheshti, Hamid Alinejad-Roknyd, Lu Chen, Mohammad Sadegh Yousefpoor, Efat Yousefpoor, Muneera Altayeb, Thantrira Porntaveetus, Sadia Din

机构 * School of Engineering and Technology, Duy Tan University(杜益坦大学工程与技术学院) Institute for Intelligent Systems Research and Innovation, Deakin University(德金大学智能系统研究与创新研究所) School of Computing, Macquarie University(麦考瑞大学计算机学院) UNSW BioMedical Machine Learning Lab (BML), School of Biomedical Engineering, UNSW Sydney(新南威尔士大学生物医学机器学习实验室(BML)) Visiting Scholar (Collaborative Projects), Center of Excellence in Precision Medicine and Digital Health, Chulalongkorn University(朱拉隆梭大学精准医学与数字健康中心) Department of Computer Science, Zhejiang University(浙江大学计算机科学系) Center of Research and Strategic Studies, Lebanese French University(黎巴嫩法语大学研究与战略研究所) Faculty of Engineering, Hourani Center for Applied Scientific Research, Al-Ahliyya Amman University(阿卡巴大学工程学院,Hourani应用科学研究中心) Center of Excellence in Precision Medicine and Digital Health, Chulalongkorn University(朱拉隆梭大学精准医学与数字健康中心) Department of Computer Engineering, Gachon University(加恩大学计算机工程系)

AI总结 本文提出QQMR协议,通过Q学习优化WBANs中的多路径路由决策,提升数据包交付率并降低延迟和能耗。

Comments Due to substantial changes in the contributions and responsibilities of the researchers involved in the project, the authorship of the manuscript requires revision to accurately reflect the current contributions. We therefore request withdrawal of the present version to appropriately resolve the authorship and contribution record

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17492 2026-08-14 cs.CV 版本更新

EvDiff: Event-Based Video Reconstruction using One-Step Diffusion Models

EvDiff:高质视频与事件相机

Weilun Li, Lei Sun, Ruixi Gao, Qi Jiang, Yuqin Ma, Kaiwei Wang, Ming-Hsuan Yang, Luc Van Gool, Danda Pani Paudel

机构 * Zhejiang University(浙江大学) INSAIT UC Merced(加州大学默塞德分校) Google DeepMind(谷歌DeepMind)

AI总结 EvDiff通过基于事件的扩散模型和替代训练框架,从单色事件流生成高质量彩色视频,提升视频生成的保真度和现实感。

Comments Replacement note: This manuscript has been transferred from the CVPR format to the ECCV 2026 format, with the corresponding title and template updated accordingly. The technical content remains largely unchanged from the previous version. (Current version: 21 pages, 6 figures, and 3 tables.) Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏