SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
SATQuest:用于大语言模型逻辑推理评估和强化微调的验证器
Yanxiao Zhao, Yaqian Li, Zihao Bo, Rinyoichi Takezoe, Haojia Hui, Mo Guang, Lei Ren, Xiaolin Qin, Kaiwen Long
机构
*
Chengdu Institute of Computer Applications, Chinese Academy of Sciences(中国科学院成都计算机应用研究所)
;
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
专题命中
推理与问题求解
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Comments23 pages, 8 figures. ACL 2026 Main Conference long paper (oral presentation)
Journal refProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2109-2131. Association for Computational Linguistics, 2026
Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
机构
*
The Chinese University of Hong Kong, Sha Tin, NT, Hong Kong(香港中文大学)
;
Université de Montréal, Montréal, Quebéc, Canada(蒙特利尔大学)
;
McGill University, Montréal, Quebéc, Canada(麦吉尔大学)
;
Mila - Quebéc AI Institute, Montréal, Quebéc, Canada(魁北克AI研究院)
;
Huawei Noah’s Ark Lab, Montréal, Quebéc, Canada(华为诺亚实验室)
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Great Bay University(大湾区大学)
;
Dongguan Key Laboratory for Intelligence and Information Technology(东莞市智能信息技术重点实验室)
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);分类 cs.AI
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
TSRouter:用于时间序列推理的动态模态-模型选择
Fangxu Yu, Tao Feng, Dehai Min, Lu Cheng, Ge Liu, Tianyi Zhou
机构
*
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
MBZUAI(Mohamed Bin Zayed University of Artificial Intelligence)
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);分类 cs.LG
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);分类 cs.CL
Comments9 pages, 5 figures. This version substantially revises the previous preprint with a new method, updated experiments, and rewritten analysis. Code available at the GitHub project repository https://anonymous.4open.science/r/sca-B666
RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
RT-SHCUA:用于无人机控制的实时自托管计算机使用代理
Di Lu, Bo Zhang, Xiyuan Li, Yongzhi Liao, Xuewen Dong, Yulong Shen, Zhiquan Liu, Jianfeng Ma
机构
*
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
;
Shaanxi Key Laboratory of Network and System Security(陕西省网络与系统安全重点实验室)
;
College of Cyber Security, Jinan University(广州大学网络安全学院)
;
School of Cyber Engineering, Shaanxi Key Lab of Network and System Security(陕西省网络与系统安全重点实验室)