What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
他们看到了什么?通过视觉-语言-动作模型的视角解读复杂道路场景以实现安全可靠的自动驾驶学习
Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer
机构
*
Birla Institute of Technology and Science, Pilani(贝拉理工科学学院皮拉尼分校)
;
University of California, Merced(加州大学默塞德分校)
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
首先,不伤害:迈向临床安全的大语言模型
David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
机构
*
Harvard Combined Dermatology Program(哈佛联合皮肤科项目)
;
Department of Dermatology, Mass General Brigham(麻省总医院皮肤科)
;
Harvard Medical School(哈佛医学院)
;
Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心)
;
Stanford University(斯坦福大学)
;
Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科)
;
Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科)
;
Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯)
;
Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科)
;
Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科)
;
Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科)
;
Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科)
;
Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科)
;
Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科)
;
Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科)
;
Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心)
;
Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所)
;
Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)
Safety from Honesty in a Disinterested AI Predictor
无兴趣AI预测器中的诚实安全性
Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gavenčiak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
机构
*
LawZero
;
Université de Montréal(蒙特利尔大学)
;
Mila
;
University of California Berkeley(加州大学伯克利分校)
;
McGill University(麦吉尔大学)
;
Arb Research
;
Center for Theoretical Study Charles University in Prague(布拉格查理大学理论研究中心)
;
University of Oxford(牛津大学)
;
Sapien Institute(Sapien研究所)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
正义的尺度:关于大语言模型安全评估的全面调查
Songyang Liu, Chaozhuo Li, Jiameng Qiu, Xi Zhang, Feiran Huang, Litian Zhang, Yiming Hei, Philip S. Yu
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Jinan University(暨南大学)
;
Beihang University(北京航空航天大学)
;
China Academy of Information and Communications Technology(中国信息通信研究院)
机构
*
Brain-inspired Cognitive Intelligence Lab, Institute of Automation, Chinese Academy of Sciences, Beijing, China(脑启发认知智能实验室,自动化研究所,中国科学院,北京,中国)
;
School of Future Technology, University of Chinese Academy of Sciences, China(未来技术学院,中国科学院大学,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, China(人工智能学院,中国科学院大学,中国)
;
Zhongguancun Academy, China(中关村学院,中国)
;
Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐重点实验室)
;
Gaoling School of AI, Renmin University of China(甘露人工智能学院,中国人民大学)
;
Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究院(北京-AISI))
;
School of Humanities, University of Chinese Academy of Sciences, China(人文学院,中国科学院大学,中国)
机构
*
Supervised Program for Alignment Research (SPAR)(对齐研究监督计划)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Yale University(耶鲁大学)
机构
*
Nankai University(南开大学)
;
James Cook University(詹姆斯库克大学)
;
Western Sydney University(西悉尼大学)
;
Beijing University of Technology(北京工业大学)
;
Fuzhou University(福州大学)
;
Nanjing University of Science and Technology(南京理工大学)
;
CSIRO's Data 61(澳大利亚联邦科学与工业研究组织Data61)
;
The University of Adelaide(阿德莱德大学)
Blockchain Infrastructure for Intelligent Cyber--Physical--Social Systems:Post-Quantum Security, Interoperability, and Trustworthy Data Economies in the Era of Embodied AI
面向智能信息-物理-社会系统的区块链基础设施:具身AI时代的后量子安全、互操作性与可信数据经济
Song Guo, Huawei Huang, Dongping Liu, Aoyu Zhang, Luyao Zhang
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Sun Yat-sen University(中山大学)
;
Amazon Web Services(亚马逊网络服务)
;
Duke Kunshan University(杜克昆山大学)
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
因果关系是理解和平衡可信机器学习与基础模型中多个目标的关键
Ruta Binkyte, Ivaxi Sheth, Zhijing Jin, Mohammad Havaei, Bernhard Schölkopf, Mario Fritz
机构
*
CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心)
;
Max Planck Institute for Intelligent Systems, Tübingen(马克斯·普朗克智能系统研究所(图宾根))
;
Google Research(谷歌研究)
;
ETH Zürich(苏黎世联邦理工学院)
;
University of Toronto(多伦多大学)
Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming
Toward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and Alignment
迈向理解自我与他人的AI系统:人类认知多样性与世界模型对齐的多阶段推理框架
Toru Takahashi
机构
*
Human Informatics and Systems Lab, Doshisha University(立命馆大学人机系统实验室)
;
Linked Open Data Initiative, NPO Keio Research Institute at SFC(庆应义塾大学SFC研究所开放数据计划)
;
Stroly Inc(Stroly公司)
Comments87 pages. Revised version with a refined abstract emphasizing disagreement as a late-stage phenomenon, target admissibility, processability, and the methodological abstraction used to compare humans, AI systems, and institutional decision procedures under shared information-theoretic constraints