arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

2026-04-16 至 2026-04-16 共收录 4
2604.14142 2026-04-16 cs.LG cs.AI cs.CL

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

从P(y|x)到P(y):在预训练空间中研究强化学习

Yuqiao Tan, Minzheng Wang, Bo Liu, Zichen Liu, Tian Liang, Shizhu He, Jun Zhao, Kang Liu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学) Tencent AI Lab(腾讯AI实验室)

AI总结 本文提出PreRL和DSRL方法,通过优化预训练空间中的边际分布P(y),提升LLM推理能力,并通过NSR机制增强推理效果,实验表明DSRL在推理任务中表现优异。

Comments Preprint. Our code is available at https://github.com/Trae1ounG/Pretrain_Space_RLVR

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20913 2026-04-16 cs.CV

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

LongVideo-R1: 低成本长视频理解的智能导航

Jihao Qiu, Lingxi Xie, Xinyue Huo, Qi Tian, Qixiang Ye

机构 * University of Chinese Academy of Sciences(中国科学院大学) Huawei Consumer Business Group(华为消费者业务集团)

AI总结 本文提出LongVideo-R1,一种基于多模态大语言模型的智能导航系统,通过高效视频上下文导航减少冗余搜索,提升长视频理解的效率与准确性。

Comments 17 pages, 9 figures, 8 tables, accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08605 2026-04-16 cs.CL cs.AI

ExpSeek: Self-Triggered Experience Seeking for Web Agents

ExpSeek:面向Web代理的自触发经验寻求

Wenyuan Zhang, Xinghua Zhang, Haiyang Yu, Shuaiyi Nie, Bingli Wu, Juwei Yue, Tingwen Liu, Yongbin Li

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Tongyi Lab , Alibaba Group(阿里云实验室,阿里巴巴集团)

AI总结 ExpSeek通过自触发机制实现Web代理的经验主动获取,提升交互能力,实验表明其在不同规模模型上均取得显著性能提升。

Comments ACL 2026 Findings, the code is accessible at https://github.com/WYRipple/ExpSeek

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12460 2026-04-16 cs.CL

Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generation

超越黑盒干预:用于忠实检索增强生成的潜在探测

Linfeng Gao, Qinggang Zhang, Baolong Bi, Bo Zeng, Zheng Yuan, Zerui Chen, Zhimin Wei, Shenghua Liu, Linlong Xu, Longyue Wang, Weihua Luo, Jinsong Su

机构 * School of Informatics, Xiamen University(厦门大学信息学院) The Hong Kong Polytechnic University(香港理工大学) University of Chinese Academy of Sciences(中国科学院大学) Alibaba Group(阿里巴巴集团) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),中华人民共和国文化和旅游部,中国)

AI总结 本文提出ProbeRAG框架,通过潜在冲突探测和注意力调节提升检索增强生成的忠实度,解决传统方法在评估知识冲突时的不足。

Comments ACL 2026 Findings; Code is available at https://github.com/LinfengGao/ProbeRAG

详情

展开后加载摘要…

URL PDF HTML 收藏