arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The University of Hong Kong(香港大学)

2026-05-11 至 2026-05-11 共收录 9
2605.07872 2026-05-11 cs.CV cs.AI

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

视频理解奖励建模:一个稳健的基准和高效的奖励模型

Yuancheng Wei, Linli Yao, Lei Li, Haojie Zhang, Hao Zhou, Fandong Meng, Xu Sun

机构 * South China University of Technology(华南理工大学) Peking University(北京大学) The University of Hong Kong(香港大学) Tencent(腾讯)

AI总结 本文提出Video Understanding Reward Bench基准和VideoDRM/VideoGRM模型,通过大规模高质量数据提升视频理解奖励建模性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03473 2026-05-11 cs.LG cs.CV

Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts

将Bi-Level路由混合专家扩展到300+任务的持续学习

Meng Lou, Yunxiang Fu, Yizhou Yu

机构 * School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Center for Embodied Artificial Intelligence and Computer Vision, Shenzhen Loop Area Institute(具身人工智能与计算机视觉中心,深圳环宇区研究所)

AI总结 本文提出CaRE,通过双级路由混合专家机制实现长任务序列的持续学习,提出OmniBenchmark-1K数据集,展示在多种数据集和任务设置上优于基线的性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07442 2026-05-11 cs.LG

GameGen-Verifier: Parallel Keypoint-Based Verification for LLM-Generated Games via Runtime State Injection

GameGen-Verifier:通过运行时状态注入实现LLM生成游戏的并行关键点验证

Chaobo Jia, Ruipeng Wan, Ting Sun, Weihao Tan, Borui Wan, Yuxuan Tong, Guangming Sheng, Hong Xu

机构 * CUHK(香港中文大学) HUST(华中科技大学) Lionrock AI Lab(Lionrock人工智能实验室) NTU(国立新加坡大学) HKU(香港大学)

AI总结 本文提出GameGen-Verifier,通过分解规范为可验证的关键点,实现LLM生成游戏的自动化验证,提升验证准确率并减少运行时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07302 2026-05-11 cs.LG

Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation

预训练诱导了可重用的谱基底以用于下游任务适应

Junjie Yu, Yue Wang, Zihan Deng, Yan Zhu, Wenxiao Ma, Quanying Liu

机构 * Department of Biomedical Engineering, Southern University of Science and Technology(生物医学工程系,南方科技大学) Department of Psychology, The University of Hong Kong(心理学系,香港大学)

AI总结 本文通过系统谱分析揭示预训练模型的主奇异向量在微调中保持稳定且跨任务共享,表明预训练建立了可重用的谱坐标系,且更大规模预训练提升几何可迁移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07274 2026-05-11 cs.AI cs.LG

Structured Role-Aware Policy Optimization for Multimodal Reasoning

结构化角色感知策略优化用于多模态推理

Bingqing Jiang, Difan Zou

机构 * School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) School of Computing & Data Science and Institute of Data Science, The University of Hong Kong(计算与数据科学学院和数据科学研究所,香港大学)

AI总结 本文提出结构化角色感知策略优化(SRPO),通过角色感知的token级信用分配提升多模态推理中的证据基础推理能力,无需外部奖励模型。

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06841 2026-05-11 cs.AI cs.LG

AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites

AGWM:基于 affordance 的世界模型用于具有组合前提条件的环境

Qinshi Zhang, Weipeng Deng, Zhihan Jiang, Jiaming Qu, Qianren Li, Weitao Xu, Ray LC

机构 * University of California, San Diego(加州大学圣地亚哥分校) University of Hong Kong(香港大学) Columbia University(哥伦比亚大学) Amazon(亚马逊) City University of Hong Kong(香港城市大学)

AI总结 本文提出 AGWM 模型,通过学习抽象 affordance 结构来跟踪动作的动态可执行性,以解决传统世界模型在多步预测中的误差累积问题。

Comments 16 pages, 3 figures, 4 tables. Appendix on pages 11-16 (main text is self-contained)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24372 2026-05-11 cs.CL cs.AI cs.NE

SeaEvo: Advancing Algorithm Discovery with Strategy Space Evolution

SeaEvo:通过策略空间进化推进算法发现

Sichun Luo, Yi Huang, Haochen Luo, Fengyuan Liu, Guanzhi Deng, Lei Li, Qinghua Yao, Zefa Hu, Junlan Feng, Qi Liu

机构 * The University of Hong Kong(香港大学) JIUTIAN Research, China Mobile(钧天研究院,中国移动) City University of Hong Kong(城市大学)

AI总结 SeaEvo通过策略空间进化层,将语言级策略推理转化为群体级进化状态,提升LLM引导的程序搜索效果,实现算法发现、系统优化和智能体设计任务的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21466 2026-05-11 cs.LG cs.IT math.IT

Normalized Maximum Likelihood Code-Length on Riemannian Data Spaces

在黎曼数据空间上的归一化最大似然代码长度

Kota Fukuzawa, Atsushi Suzuki, Kenji Yamanishi

机构 * The University of Tokyo(东京大学) NTT, Inc.(NTT公司) The University of Hong Kong(香港大学)

AI总结 本文提出在黎曼流形上适应的归一化最大似然代码长度(Rm-NML),克服了传统方法对坐标系的依赖,适用于具有层次结构的图数据,如双曲空间。

Comments 19 pages. This is a preprint of an article accepted for publication in the IEEE Transactions on Information Theory

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09929 2026-05-11 cs.LG cs.CV

Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise

对比学习中的数据增强是估计正激励噪声

Hongyuan Zhang, Yanchen Xu, Sida Huang, Xuelong Li

机构 * The University of Hong Kong(香港大学) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院) Fudan University(复旦大学) School of Artificial Intelligence, OPtics(人工智能学院) ElectroNics (iOPEN), Northwestern Polytechnical University(西北工业大学电子学院)

AI总结 本文研究对比学习与正激励噪声的联系,定义了任务熵并提出正激励噪声生成器,通过可视化证明所提方法有效。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏