CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
CSRP:基于效率感知奖励的强化学习链式推理中文文本纠错
Wei Tian, Yuhao Zhou, Man Lan
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Shanghai Institute of Artificial Intelligence for Education, East China Normal University(东华大学教育人工智能研究所)
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
人工推理之谜:探究大型推理模型中的生成-评估差距
Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan
机构
*
NUS Department of Computer Science(国立新加坡大学计算机科学系)
;
MIT EECS(麻省理工学院电子工程与计算机科学系)
;
A*STAR(新加坡科技研究局)
;
Singapore-MIT Alliance for Research and Technology (SMART)(新加坡-麻省理工联合研究技术机构(SMART))
CommentsThis manuscript is withdrawn to allow careful review and correction of bibliographic issues identified after submission, including references that could not be adequately verified. These matters should be resolved before further circulation
机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Institute of Artificial Intelligence and Future Networks, Beijing Normal University(北京师范大学人工智能与未来网络研究院)
;
Faculty of Arts and Sciences, Beijing Normal University(北京师范大学文理学院)
;
Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学)
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
TimeSage-MT:用于评估智能时间序列推理的多轮基准测试
Yaxuan Kong, Qingren Yao, Yuqi Nie, Yichen Li, Yilei Shao, Stefan Zohren, Anna Vettoruzzo, Joaquin Vanschoren, Ming Jin, Qingsong Wen
机构
*
University of Oxford(牛津大学)
;
VulpiVox Intelligence
;
Eindhoven University of Technology(埃因霍温理工大学)
;
Griffith University(格里菲斯大学)
;
Squirrel Ai Learning
;
East China Normal University(华东师范大学)
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
PolySpeech-100:面向100多种语言和方言的大规模语音理解基准
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu, Lu Fan, Zhi Li, You He
机构
*
Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学)
;
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
;
JD AI Research(京东人工智能研究院)
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
ADRA-Bank:面向学术深度研究智能体的模块化基准
Zhihan Guo, Feiyang Xu, Yifan Li, Muzhi Li, Shuai Zou, Jiele Wu, Han Shi, Haoli Bai, Ho-fung Leung, Irwin King
机构
*
The Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学)
;
Hong Kong Polytechnic University, Hong Kong SAR, China(香港理工大学)
;
National University of Singapore, Singapore, Singapore(新加坡国立大学)
;
Huawei Technologies, Hong Kong SAR, China(华为技术有限公司)
;
Independent(独立)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
SCOPE: 信号校准的在线策略蒸馏增强与双路径自适应加权
Binbin Zheng, Xing Ma, Yiheng Liang, Jingqing Ruan, Xiaoliang Fu, Kepeng Lin, Benchang Zhu, Ke Zeng, Xunliang Cai
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Meituan LongCat Interaction Team(美团 LongCat 交互团队)
;
Nanjing University(南京大学)
;
Fudan University(复旦大学)
;
Huazhong University of Science and Technology(华中科技大学)
When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs
当单一答案不够时:重新思考面向大语言模型的单步逆合成基准
Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Mathieu Reymond, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftakhutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov