机构
*
Vellore Institute of Technology(维洛雷理工学院)
;
University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)
;
Northwestern University(西北大学)
;
Yale University(耶鲁大学)
;
Algoverse AI Research(Algoverse AI研究)
Comments12 pages, 4 figures, 6 tables. Includes ablation study across Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct on 5 math reasoning benchmarks (GSM8K, MATH500, Minerva, AIME24, Gaokao2023). GPT-4.1 used for structured evaluation of reasoning quality
CommentsPreprint. 16 pages, 2 figures. Live interactive demo: https://huggingface.co/spaces/Squagghy/moxia. Paper artifact and dataset on Zenodo (concept-DOI): 10.5281/zenodo.21906509
Understanding Reasoning from Pretraining to Post-Training
理解从预训练到训练后阶段的推理
Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov
机构
*
New York University(纽约大学)
;
Modal Labs(模态实验室)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Columbia University(哥伦比亚大学)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
在线推理校准:测试时训练使可泛化的符合性大语言模型推理
Cai Zhou, Zekai Wang, Menghua Wu, Qianyu Julie Zhu, Flora C. Shi, Chenyu Wang, Ashia Wilson, Tommi Jaakkola, Stephen Bates
机构
*
Department of Electrical Engineering and Computer Science (MIT EECS)(麻省理工学院电气工程与计算机科学系)
;
Computer Science and Artificial Intelligence Laboratory (MIT CSAIL)(麻省理工学院计算机科学与人工智能实验室)
;
Laboratory for Information and Decision Systems (MIT LIDS)(麻省理工学院信息与决策系统实验室)
;
Computational Science and Engineering (MIT CSE)(麻省理工学院计算科学与工程)
;
Massachusetts Institute of Technology(麻省理工学院)
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning
R$^2$PO:用于大语言模型推理的解耦展开与推理策略
Jingchu Wang, Bingbing Xu, Yige Yuan, Dan Zhang, Bin Xie, Xiaoqian Sun, Huawei Shen
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
National University of Singapore(新加坡国家大学)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
MoReBench:评估语言模型中的程序性和多元道德推理,超越结果
Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Raphaël Millière, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
机构
*
University of Washington(华盛顿大学)
;
New York University(纽约大学)
;
Scale AI
;
Harvard University(哈佛大学)
;
University of Michigan(密歇根大学)
;
UNC Chapel Hill(北卡罗来纳大学教堂山分校)
;
Center for AI Safety(人工智能安全中心)
;
Stanford University(斯坦福大学)
;
MIT(麻省理工学院)
;
University of Oxford(牛津大学)
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
熵梯度反转:迈向大型推理模型的内部机制
Junyao Yang, Chen Qian, Kun Wang, Linfeng Zhang, Quanshi Zhang, Yong Liu, Dongrui Liu
机构
*
National University of Singapore(新加坡国立大学)
;
Renmin University of China(中国人民大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
City University of Hong Kong(香港城市大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
Meta AI
;
Department of Computer Science, University of California, Santa Barbara(加州大学圣芭芭拉分校计算机科学系)
;
Google DeepMind(谷歌DeepMind)
;
Independent Researcher(独立研究者)
HauntAttack: When Attack Follows Reasoning as a Shadow
HauntAttack:当攻击如影随形地跟随推理
Jingyuan Ma, Rui Li, Zheng Li, Junfeng Liu, Heming Xia, Lei Sha, Zhifang Sui
机构
*
State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
;
StepFun
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)
;
Institute of Artificial Intelligence, Beihang University(北航人工智能研究院)