Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
Wasserstein分布鲁棒遗憾优化用于人类反馈的强化学习
Yikai Wang, Shang Liu, Jose Blanchet
机构
*
Department of Statistics and Operations Research, University of North Carolina(统计与运筹学系,北卡罗来纳大学)
;
Imperial Business School, Imperial College London(帝国理工学院伦敦商学院)
;
Department of Management Science and Engineering, Stanford University(管理科学与工程系,斯坦福大学)
Comments171 pages. Formalized in Lean 4 with Mathlib: 240 theorems in the elaborated environment, 141 audited headline results, cold-compiling from a clean checkout with zero custom axioms. Source, theorem-by-theorem contract, and reproducible axiom audit: https://github.com/selfreferencing/TSE_Formal. Companion to Agentic Capital
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
Goedel-Code-Prover:面向开放状态的最新代码验证的分层证明搜索
Zenan Li, Ziran Yang, Deyuan He, Haoyu Zhao, Andrew Zhao, Shange Tang, Kaiyu Yang, Aarti Gupta, Zhendong Su, Chi Jin
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
Princeton Language and Intelligence(普林斯顿语言与智能实验室)
;
Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
;
MiroMind
DiSCo: Diffusion Sequence Copilots for Shared Autonomy
DiSCo:用于共享自主的扩散序列助手
Andy Wang, Xu Yan, Brandon McMahan, Michael Zhou, Yuyang Yuan, Johannes Y. Lee, Ali Shreif, Matthew Li, Zhenghao Peng, Bolei Zhou, Yuchen Cui, Jonathan C. Kao
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
机构
*
SKLP, ICT, CAS & UCAS(SKLP、信息科技研究院、中国科学院及中国科学院大学)
;
University of Aberdeen(阿伯丁大学)
;
University of Leeds(利兹大学)
;
SKLP, ICT, CAS(SKLP、信息科技研究院、中国科学院)
Comments\c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
通道卫士:安全模型无法组合成安全的多智能体系统
Elias Hossain, Md Mehedi Hasan Nipu, Fatema Tuj Johora Faria, Tasfia Nuzhat Ornee, Maleeha Sheikh
机构
*
College of Engineering and Computer Science, University of Central Florida(工程与计算机科学学院,中央佛罗里达大学)
;
Department of Computer Science and Engineering, North South University(计算机科学与工程系,北南大学)
;
Computer Science and Engineering, Ahsanullah University of Science and Technology(计算机科学与工程,阿沙努拉科学与技术大学)
;
Department of Electrical and Computer Engineering, Purdue University Fort Wayne(电气与计算机工程系,普渡大学弗拉特沃恩分校)
Comments21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval
Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings
解释、验证和对齐视觉语言模型嵌入中的语义层次结构
Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika Mütze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann
机构
*
University of Lübeck(吕贝克大学)
;
Technical University of Munich(慕尼黑工业大学)
;
AUMOVIO SE
;
Osnabrück University(奥斯纳布吕克大学)
;
Leibniz Institute for Agricultural Engineering and Bioeconomy(莱布尼茨农业工程与生物经济研究所)
;
University of Hamburg(汉堡大学)
Integrating RCTs, RWD, AI/ML and Statistics: Next-Generation Evidence Synthesis
整合RCTs、RWD、AI/ML和统计学:下一代证据合成
Shu Yang, Margaret Gamalo, Haoda Fu
机构
*
Department of Statistics, North Carolina State University(统计学系,北卡罗来纳州立大学)
;
VP and Statistics Head, Inflammation, Immunology & Specialty Care, Pfizer(副总裁及统计学负责人,炎症、免疫与专科医疗,辉瑞)
;
Head of Exploratory Biostatistics, Amgen(探索性生物统计学负责人,安进)