arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

The University of Hong Kong(香港大学)

2026-04-06 至 2026-04-06 共收录 6
2604.03098 2026-04-06 cs.LG cs.AI cs.CL

Co-Evolution of Policy and Internal Reward for Language Agents

语言代理中策略与内部奖励的共演化

Xinyu Wang, Hanwei Wu, Jingwei Song, Shuyuan Zhang, Jiayi Zhang, Fanqi Kong, Tung Sum Thomas Kwok, Xiao-Wen Chang, Yuyu Luo, Chenglin Wu, Bang Liu

机构 * McGill University(麦吉尔大学) McMaster University(麦克马斯特大学) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Peking University(北京大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校) DeepWisdom(深度智慧) Université de Montréal(蒙特利尔大学) Mila(米拉研究所)

AI总结 本文提出Self-Guide方法,通过自动生成内部奖励实现推理和训练时的协同优化,提升语言代理性能。

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03022 2026-04-06 cs.SI cs.AI cs.HC

Comparing the Impact of Pedagogy-Informed Custom and General-Purpose GAI Chatbots on Students' Science Problem-Solving Processes and Performance Using Heterogeneous Interaction Network Analysis

比较基于教学导向的定制与通用目的GAI聊天机器人对学生产生的科学问题解决过程和表现的影响:使用异质交互网络分析

Hanyu Su, Huilin Zhang, Shihui Feng

机构 * Faculty of Education, The University of Hong Kong(香港大学教育学院)

AI总结 本研究通过异质交互网络分析比较了基于教学导向的定制与通用目的GAI聊天机器人对学生产生的科学问题解决过程和表现的影响,发现定制聊天机器人能提高学生的认知互动强度和多样性。

Comments Full paper accepted to the 27th International Conference on AI in Education (AIED 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02934 2026-04-06 cs.CV

PolyReal: A Benchmark for Real-World Polymer Science Workflows

PolyReal:一个用于现实世界聚合物科学工作流程的基准

Wanhao Liu, Weida Wang, Jiaqing Xie, Suorong Yang, Jue Wang, Benteng Chen, Guangtao Mei, Zonglin Yang, Shufei Zhang, Yuchun Mo, Lang Cheng, Jin Zeng, Houqiang Li, Wanli Ouyang, Yuqiang Li

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Northwestern Polytechnical University(西北工业大学) Tongji University(同济大学) The University of Hong Kong(香港大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

AI总结 PolyReal是一个新的多模态基准,用于评估大规模语言模型在聚合物实验全流程中的能力,揭示了模型在实践任务上的不足,填补了评估空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02648 2026-04-06 cs.SE cs.AI

GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers

GBQA:用于评估LLM作为质量保证工程师的博弈基准

Shufan Jiang, Chios Chen, Zhiyang Chen

机构 * The University of Hong Kong(香港大学) Independent Researcher(独立研究者) Westlake University(西湖大学) Datawhale Org(Datawhale组织)

AI总结 本文提出GBQA基准,通过30款游戏和124个经人类验证的bug测试LLM自主发现软件缺陷的能力,实验表明最佳模型仅能识别48.39%的bug。

Comments Accepted as a workshop paper at the Fourteenth International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19090 2026-04-06 cs.CL

Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents

辩论真相:基于多个大语言模型代理的辩论驱动声明验证

Haorui He, Yupeng Li, Dacheng Wen, Yang Chen, Reynold Cheng, Donglong Chen, Francis C. M. Lau

机构 * Department of Interactive Media, Hong Kong Baptist University(香港浸会大学互动媒体系) School of Computing and Data Science, The University of Hong Kong(香港大学计算与数据科学学院) College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院) Shenzhen Institute of Advanced Technology(深圳先进技术研究院)

AI总结 本文提出DebateCV框架,通过多代理辩论和调解提升复杂声明验证的准确性与说服力。

Comments Accepted by the ACM Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07874 2026-04-06 cs.CL cs.AI

Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling

语言框架在神经符号语言建模中展开对决

Jakob Prange, Nathan Schneider, Lingpeng Kong

机构 * Georgetown University(乔治城大学) The University of Hong Kong(香港大学)

AI总结 研究探讨了语言图表示如何提升神经语言模型性能,发现语义短语结构最有效,且词性类别影响效果差异。

Comments Accepted to NAACL 2022 (slight typesetting divergences to NAACL camera-ready due to TexLive 2020/2021 mismatches)

详情

展开后加载摘要…

URL PDF HTML 收藏