R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning
R$^2$PO:用于大语言模型推理的解耦展开与推理策略
Jingchu Wang, Bingbing Xu, Yige Yuan, Dan Zhang, Bin Xie, Xiaoqian Sun, Huawei Shen
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
National University of Singapore(新加坡国家大学)
机构
*
Meta AI
;
Department of Computer Science, University of California, Santa Barbara(加州大学圣芭芭拉分校计算机科学系)
;
Google DeepMind(谷歌DeepMind)
;
Independent Researcher(独立研究者)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Sun Yat-sen University(中山大学)
;
Southern University of Science and Technology(南方科技大学)
;
National University of Singapore(新加坡国立大学)
Chuanhao Yan, Fengdi Che, Xuhan Huang, Xu Xu, Xin Li, Yizhi Li, Xingwei Qu, Jingzhe Shi, Chenghua Lin, Yaodong Yang, Binhang Yuan, Hang Zhao, Yu Qiao, Bowen Zhou, Jie Fu
机构
*
Shanghai AI Lab(上海人工智能实验室)
;
University of Alberta(阿尔伯塔大学)
;
Tsinghua University(清华大学)
;
Chinese University of Hong Kong, Shenzhen(香港大学(深圳))
;
Hong Kong University of Science and Technology(香港科技大学)
;
Nanyang Technological University(南洋理工大学)
;
University of Manchester(曼彻斯特大学)
;
Peking University(北京大学)
To Use AI as Dice of Possibilities with Timing Computation
将AI用作带有时序计算的可能性骰子
Jia Li, Vipin Kumar, Rui Zhang
机构
*
Department of Surgery, University of Minnesota(明尼苏达大学外科系)
;
Department of Computer Science & Engineering, University of Minnesota(明尼苏达大学计算机科学与工程系)
AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection
AULLM++:用于微表情动作单元检测的结构化令牌条件大语言模型
Zhishu Liu, Kaishen Yuan, Bo Zhao, Hui Ma, Zitong Yu
机构
*
Great Bay University(广东东莞大学)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构
*
Department of Computer Science and Engineering, University of Minnesota(计算机科学与工程系,明尼苏达大学)
;
Computational Biology Branch, National Library of Medicine(国家医学图书馆计算生物学分支)
;
Khoury College of Computer Sciences, Northeastern University(东北大学计算机科学学院)
;
State Key Laboratory for Novel Software Technology at Nanjing University, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室,南京大学计算机科学学院)
;
Developmental Therapeutics Branch, National Cancer Institute(国家癌症研究所发育治疗分支)
Toward Efficient Agents: Memory, Tool learning, and Planning
迈向高效智能体:记忆、工具学习与规划
Xiaofang Yang, Lijun Li, Heng Zhou, Tong Zhu, Xiaoye Qu, Yuchen Fan, Qianshan Wei, Rui Ye, Li Kang, Yiran Qin, Daizong Liu, Qi Li, Ning Ding, Siheng Chen, Jing Shao
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiaotong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))
;
Hong Kong Polytechnic University(香港理工大学)
;
Wuhan University(武汉大学)
;
Tsinghua University(清华大学)
CommentsPublished at SURGeLLM 2026, ACL 2026. Camera-ready version
Journal refProceedings of the First Workshop on Structured Understanding, Retrieval, and Generation in the LLM Era (SURGeLLM 2026), pp. 132-151, Association for Computational Linguistics, 2026
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
CoopEval:社会困境中合作维持机制与LLM代理的基准测试
Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin
机构
*
Carnegie Mellon University
;
Foundations of Cooperative AI Lab (FOCAL)
;
Jinesis Lab, University of Toronto \& Vector Institute
;
ETH Z\" u rich
;
Max Planck Institute for Intelligent Systems, T\" u bingen, Germany