Entropy-Aware On-Policy Distillation of Language Models
熵感知的在线策略蒸馏语言模型
Woogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei, Yi Zhou, Swanand Ravindra Kadhe, Nathalie Baracaldo, Kimin Lee
机构
*
IBM Research, San Jose, CA, USA(IBM研究院,旧金山,加州,美国)
;
University of Toronto, Ontario, Canada(多伦多大学,安大略,加拿大)
;
Vector Institute, Ontario, Canada(向量研究所,安大略,加拿大)
机构
*
Morgan Stanley(摩根士丹利)
;
Clemson University(克莱姆森大学)
;
Arizona State University(亚利桑那州立大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
University of Notre Dame(圣母大学)
;
University of Arizona(亚利桑那大学)
Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations
基于LLM解释的完备且可靠的神经常识推理
Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth
机构
*
University of Amsterdam(阿姆斯特丹大学)
;
University of Southern California(南加州大学)
;
Rensselaer Polytechnic Institute(拉特格斯理工学院)
;
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
Comments43 pages, 14 tables, 4 figures. Accepted to the 19th Conference on Neurosymbolic Learning and Reasoning (NeSy 2025); to appear Neurosymbolic Artifical Intelligence Special Issue on NeSy 2025 Extended Papers
机构
*
The Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格雷格学院工程学院、核等工程学院)
;
Department of Nuclear Engineering, Hanyang University(汉阳大学核工程系)
;
University of Texas - El Paso(德克萨斯大学埃尔帕索分校)
;
National Center for Supercomputing Applications(国家超级计算应用中心)
;
Department of Applied Mechanics, Indian Institute of Technology Delhi(印度德里理工学院应用力学系)
;
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度德里理工学院亚里人工智能学院)
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
InvEvolve:通过具有性能保证的大语言模型进化白盒库存策略
Chenyu Huang, Jianghao Lin, Zhengyang Tang, Bo Jiang, Ruoqing Jiang, Benyou Wang, Lai Wei
机构
*
Shanghai University of Finance and Economics(上海财经大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Tsinghua University(清华大学)
;
Boston College(波士顿大学)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
自回归与扩散语言模型中的逐步拒绝动态
Eliron Rahimi, Elad Hirshel, Rom Himelstein, Amit LeVi, Avi Mendelson, Chaim Baskin
机构
*
Department of Computer Science, Technion – Israel Institute of Technology(技术学院计算机科学系,以色列技术学院)
;
INSIGHT Lab, School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel(内斯坦实验室,贝内-加隆大学内加尔分校,以色列)
;
Computer Science Department, University of Haifa, Haifa, Israel(海法大学计算机科学系,海法,以色列)
Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach
自动驾驶鲁棒控制:一种智能一般和约束对抗强化学习方法
Junchao Fan, Qi Wei, Ruichen Zhang, Yang Lu, Jianhua Wang, Xiaolin Chang, Bo Ai
机构
*
Beijing Key Laboratory of Security and Privacy in Intelligent Transportation(北京智能交通安全与隐私重点实验室)
;
Beijing Jiaotong University(北京交通大学)
;
College of Computing and Data Science(计算与数据科学学院)
;
Nanyang Technological University(南洋理工大学)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
Taiyuan University of Technology(太原科技大学)
;
School of Electronics and Information Engineering(电子与信息工程学院)
Comments23 pages, 14 figures. Major revision merging the follow-up "Standing Invariants at Scale" into this paper per arXiv moderation: the machine-checked action-safety core preserved across six further releases, six new invariant families with teeth, and real-hardware self-improvement; suite grown from 122 to 563 tests. Software, gate suite, run artifacts, and TLA+ spec are open source
机构
*
Independent Researcher, Seattle, WA, USA(华盛顿州塞勒姆独立研究员)
;
Independent Researcher, New York City, NY, USA(纽约市纽约独立研究员)
;
King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹科学技术大学)
Grounding Large Language Models as Generalizable Policies in Network Control
大语言模型作为网络优化的通用策略
Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang
机构
*
Bytedance(字节跳动)
;
Shenzhen International Graduate School(深圳国际研究生院)
;
Tsinghua University(清华大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Department of Computer Science and Technology(计算机科学与技术系)
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
Goedel-Code-Prover:面向开放状态的最新代码验证的分层证明搜索
Zenan Li, Ziran Yang, Deyuan He, Haoyu Zhao, Andrew Zhao, Shange Tang, Kaiyu Yang, Aarti Gupta, Zhendong Su, Chi Jin
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
Princeton Language and Intelligence(普林斯顿语言与智能实验室)
;
Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
;
MiroMind
Comments17 pages, 3 figures, 6 tables (9-page main text). Ancillary file hidden-automata-rl-code.zip contains reproduction code and the complete per-run data behind every table and figure. v3: author name corrected, title revised, text rewritten for clarity, new robustness checks (nonlinear and off-policy probes) added; results and conclusions unchanged
Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents
编译,然后分页:用于过程性语言模型代理的可执行标准操作程序和能力门控运行时
Chenglin Yu, Li Yin, Qingxin Fan, Ying Yu, RunyangRay Zhong, Ming Li
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Hong Kong(香港大学)
;
Zhejiang Normal University(浙江师范大学)
;
Research Institute for Generative AI, The Hong Kong Polytechnic University(香港理工大学生成式人工智能研究所)