GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
GR-Ben:一种通用推理基准,用于评估过程奖励模型
Zhouhao Sun, Xuan Zhang, Xiao Ding, Bibo Cai, Li Du, Kai Xiong, Xinran Dai, Fei Zhang, weidi tang, Zhiyuan Kan, Yang Zhao, Bing Qin, Ting Liu
机构
*
Research Center for Social Computing(社会计算研究中心)
;
Interactive Robotics, Harbin Institute of Technology, China(交互机器人学,哈尔滨工业大学,中国)
;
Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院,北京,中国)
AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification
AstroAlertBench: 评估多模态大语言模型在天文学分类中的准确性、推理和诚实性
Claire Chen, Jiabao Sean Xiao, Shuze Daniel Liu, Facundo Perez Paolino, Luke Handley, Theophile Jegou du Laz, Ricky Nilsson, Alice Zou, Matthew Graham, Ashish Mahabal
机构
*
California Institute of Technology(加州理工学院)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Purdue University(普渡大学)
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
无人机视角下视觉语言模型能否思考?统一无人机推理与生成
Jintao Sun, Gangyi Ding, Donglin Di, Hu Zhang, Zhedong Zheng
机构
*
Beijing Institute of Technology(北京理工大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
CSIRO Data61(澳大利亚联邦科学与工业研究组织 Data61 分部)
;
University of Macau(澳门大学)
CommentsThe authors have identified an issue in the evaluation protocol in Section 5.1.3. Feature extraction and semantic matching used to compute P-EHR require correction and re-validation, as they may not have been applied consistently across all generated explanations and baselines. This may affect part of the reported quantitative results and analysis, so the authors withdraw this version