Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
尾感知信息论泛化用于RLHF和SGLD
Huiming Zhang, Binghan Li, Wan Tian, Qiang Sun
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高精尖创新中心)
;
Advanced Institute of Information Technology, Peking University(北京大学信息技术高等研究院)
;
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所)
;
Computer and Mathematical Sciences, Computer Science, and Statistics, University of Toronto(多伦多大学计算机与数学科学、计算机科学和统计学系)
;
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
从表征互补性到双系统:协同VLM和纯视觉骨干网络用于端到端驾驶
Sining Ang, Yuguang Yang, Chenxu Dang, Canyu Chen, Cheng Chi, Haiyan Liu, Xuanyao Mao, Jason Bao, Xuliang, Bingchuan Sun, Yan Wang
机构
*
Department of Automation, University of Science and Technology of China(中国科学技术大学自动化系)
;
School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院)
;
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
National Superior College for Engineers, Beihang University(北京航空航天大学国家级工程师学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Lenovo Group Limited(联想集团有限公司)
;
Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
机构
*
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
School of Software, Beihang University(北京航空航天大学软件学院)
;
Shandong Inspur Intelligent Production Technology Co., Ltd(山东 Inspur 智能生产技术有限公司)
;
Zhongguancun Laboratory(中关村实验室)
CommentsAccepted for publication at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026). \c{opyright} ACM, 2026. This is the author's version of the work. It is posted here by permission of ACM for your personal use. Not for redistribution. The definitive Version of Record will be published by ACM, https://doi.org/10.1145/3832783.3834487