Reasoning Compression with Mixed-Policy Distillation
基于混合策略蒸馏的推理压缩
Han Yang, Mingyan Wu, Bailan He, Zeyu Cao, Sikuan Yan, Kevin Qinghong Lin, Zifeng Ding
机构
*
Technical University of Munich(慕尼黑技术大学)
;
GESIS – Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所)
;
Northeastern University(东北大学)
;
LMU Munich(慕尼黑大学)
;
Siemens AG(西门子股份公司)
;
University of Cambridge(剑桥大学)
;
University of Oxford(牛津大学)
;
Mina AI
机构
*
Qwen Large Model Application Team, Alibaba(阿里巴巴大模型应用团队)
;
Alibaba Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所阿里巴巴分所)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
What should post-training optimize? A test-time scaling law perspective
在测试时间视角下,后训练应优化什么?
Muheng Li, Jian Qian, Wenlong Mou
机构
*
Department of Statistical Sciences, University of Toronto(多伦多大学统计科学系)
;
Department of AI and Data Science, The University of Hong Kong(香港大学人工智能与数据科学系)
机构
*
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
University of Virginia(弗吉尼亚大学)
;
University of Maryland(马里兰大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
镜中的攻击者:通过锚定双策略自我博弈打破安全性中的自我一致性
Gabriele La Malfa, Emanuele La Malfa, Saar Cohen, Jie M. Zhang, Michael Luck, Michael Wooldridge, Elizabeth Black
机构
*
Department of Informatics, King’s College London(伦敦国王学院信息学院)
;
Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
University of Sussex(Sussex大学)
;
Institute for Decentralized AI(去中心化人工智能研究所)
机构
*
School of Electronic, Electrical and Physics, Fujian University of Technology(福建工程学院电子、电气与物理系)
;
School of Humanities, Fujian University of Technology(福建工程学院人文学院)
;
Department of Investigation, Fujian Police College(福建警察学院调查系)
;
Institute of Applied Physics and Materials Engineering, University of Macau(澳门大学应用物理与材料工程研究院)
Echo-LoRA: Parameter-Efficient Fine-Tuning via Cross-Layer Representation Injection
Echo-LoRA:通过跨层表示注入实现参数高效微调
Yihang Peng, Peng Jin, Jie Gong, Xingyuan Chen, Lingjiao Xu, Ning Su, Yan Ran
机构
*
School of Computer Science and Software Engineering, Southwest Petroleum University(西南石油大学计算机科学与软件工程学院)
;
School of Electronic Information and Artificial Intelligence, Leshan Normal University(乐山师范学院电子信息与人工智能学院)
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
计算作为教师:将推理计算转化为无参考监督
Dulhan Jayalath, Shashwat Goel, Thomas Foster, Parag Jain, Suchin Gururangan, Cheng Zhang, Anirudh Goyal, Alan Schelten
机构
*
University of Oxford(牛津大学)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Meta Superintelligence Labs(元宇宙超级智能实验室)
Prompt to Protection: A Comparative Study of Multimodal LLMs in Construction Hazard Recognition
提示到保护:多模态大语言模型在建筑危险识别中的比较研究
Nishi Chaudhary, S M Jamil Uddin, Sathvik Sharath Chandra, Anto Ovid, Alex Albert
机构
*
Department of Construction Management, Colorado State University(科罗拉多州立大学建设管理系)
;
Department of Civil, Construction, and Environmental Engineering, North Carolina State University(北卡罗来纳州立大学土木、建设与环境工程系)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
Self-ReSET: 通过学习自我恢复来应对不安全推理轨迹
Dongcheng Zhang, Yi Zhang, Yuxin Chen, An Zhang, Xiang Wang, Chaochao Lu
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
National University of Singapore(新加坡国立大学)