MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
MoReBench:评估语言模型中的程序性和多元道德推理,超越结果
Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Raphaël Millière, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
机构
*
University of Washington(华盛顿大学)
;
New York University(纽约大学)
;
Scale AI
;
Harvard University(哈佛大学)
;
University of Michigan(密歇根大学)
;
UNC Chapel Hill(北卡罗来纳大学教堂山分校)
;
Center for AI Safety(人工智能安全中心)
;
Stanford University(斯坦福大学)
;
MIT(麻省理工学院)
;
University of Oxford(牛津大学)
Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning
鲁棒的指令遵从:合作多智能体强化学习
Wo Wei Lin, Ethan Rathbun, Enrico Marchesini, Xiang Zhi Tan
机构
*
Department of Computer Sciences, Northeastern University(东北大学计算机科学系)
;
Department of Computer Sciences, Massachusetts Institute of Technology(麻省理工学院计算机科学系)
CU-Multi: A Dataset for Multi-Robot Collaborative Perception
CU-Multi:多机器人协同感知数据集
Doncey Albin, Daniel McGann, Miles Mena, Annika Thomas, Harel Biggie, Xuefei Sun, Steve McGuire, Jonathan P. How, Christoffer Heckman
机构
*
Autonomous Robotics and Perception Group at the University of Colorado Boulder(科罗拉多大学波尔得分校自主机器人与感知组)
;
Robot Perception Lab at Carnegie Mellon University(卡内基梅隆大学机器人感知实验室)
;
Aerospace Controls Laboratory at Massachusetts Institute of Technology(麻省理工学院航空航天控制实验室)
;
Computer Science and Artificial Intelligence Laboratory at Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)
;
Human-Aware Robotic Exploration Lab at University of California Santa Cruz(加州大学圣克ruz分校人感知机器人探索实验室)
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
评估卡:AI评估报告的解释层
Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Max Lamparth, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
机构
*
Hugging Face
;
Stanford University(斯坦福大学)
;
Queen Mary University of London(伦敦玛丽女王大学)
;
University of Copenhagen(哥本哈根大学)
;
Trustible
;
EleutherAI
;
TU Darmstadt(达姆施塔特工业大学)
;
Weizenbaum Institute & Technical University of Munich(魏森鲍姆研究所与慕尼黑工业大学)
;
Harvard University(哈佛大学)
;
The Hebrew University of Jerusalem(耶路撒冷希伯来大学)
;
Iowa State University(爱荷华州立大学)
;
IBM Research(IBM研究院)
;
University of Chicago(芝加哥大学)
;
Independent(独立)
;
Berkeley AI Safety Institute (BASIS)(伯克利人工智能安全研究所)
;
Simula
;
University of Edinburgh(爱丁堡大学)
;
ETH Zurich & ETH AI Center(苏黎世联邦理工学院与ETH AI中心)
;
Oxford Internet Institute(牛津互联网研究所)
;
Amherst College(阿默斯特学院)
;
University of Nebraska(内布拉斯加大学)
;
Syntony Research
;
McGill University(麦吉尔大学)
;
Evals Consensus
;
Israel Institute of Technology(以色列理工学院)
;
IOL.Learn & Zuse Institute Berlin(IOL.Learn与柏林祖泽研究所)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Quebec AI Institute, Université de Montréal(魁北克人工智能研究所,蒙特利尔大学)
;
University of Notre Dame(圣母大学)
;
Georgetown University(乔治城大学)
;
DHBW Stuttgart(斯图加特双元制大学)
;
Massachusetts Institute of Technology(麻省理工学院)
Training Set Augmentation and Biology-Aware Harmonization Improve Radiomic Models for Lung Cancer Prediction in Indeterminate Nodules
训练集增强与生物学感知的谐波化改善不确定肺结节中肺癌预测的影像组学模型
Claire Huchthausen, Menglin Shi, Gabriel L. A. de Sousa, James Larner, Einsley Janowski, Jonathan Colen, Krishni Wijesooriya
机构
*
Department of Radiation Oncology, University of Virginia School of Medicine(弗吉尼亚大学医学院放射肿瘤学系)
;
Department of Physics, University of Virginia(弗吉尼亚大学物理系)
;
Department of Physics, Massachusetts Institute of Technology(麻省理工学院物理系)
;
Department of Biomedical Engineering, Northwestern University(西北大学生物医学工程系)
;
Department of Radiation Oncology, University of Virginia(弗吉尼亚大学放射肿瘤学系)
;
Old Dominion University(旧 Dominion 大学)
机构
*
University of Michigan(密歇根大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
Stanford University(斯坦福大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
Urban Flood Observations: A hand-labeled training and validation dataset of post-flood inundation
城市洪水观测:一个手标注的训练和验证数据集,用于洪水后淹没区域
Rohit Mukherjee, Hannah K. Friedrich, Beth Tellman, Ariful Islam, Zhijie Zhang, Jonathan Giezendanner, Upmanu Lall, Venkataraman Lakshmi
机构
*
Pacific Northwest National Laboratory(太平洋西北国家实验室)
;
University of Arizona(亚利桑那大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Utah State University(犹他州立大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Columbia University(哥伦比亚大学)
;
University of Virginia(弗吉尼亚大学)
TraversalBench: Challenging Paths to Follow for Vision Language Models
TraversalBench: 为视觉语言模型设计的复杂路径挑战测试集
Clara Petrova, Zhuo Chen, Marin Soljačić
机构
*
Massachusetts Institute of Technology, Department of Physics(麻省理工学院物理系)
;
Massachusetts Institute of Technology, Institute for Data, Systems, and Society(麻省理工学院数据、系统与社会研究所)
;
NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用AI研究所)
Learning What's Real: Disentangling Signal and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics
学习真实内容:在多传感器数据中分离信号和测量伪影,应用于天体物理学
Pablo Mercader-Perez, Carolina Cuesta-Lazaro, Daniel Muthukrishna, Jeroen Audenaert, V. Ashley Villar, David W. Hogg, Marc Huertas-Company, William T. Freeman
机构
*
Massachusetts Institute of Technology(麻省理工学院)
;
Flatiron Institute, Simons Foundation(Flatiron研究所,Simons基金会)
;
Institute for Advanced Studies(高级研究 institute)
;
Harvard University(哈佛大学)
;
New York University(纽约大学)
;
Instituto de Astrofísica de Canarias(加那利大天文台)
AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
AlphaOPT: 利用自改进LLM经验库构建优化问题
Minwei Kong, Ao Qu, Xiaotong Guo, Wenbin Ouyang, Chonghe Jiang, Han Zheng, Yining Ma, Dingyi Zhuang, Yuhan Tang, Junyi Li, Shenhao Wang, Haris Koutsopoulos, Hai Wang, Cathy Wu, Jinhua Zhao
机构
*
Singapore-MIT Alliance for Research and Technology(新加坡-麻省理工联合研究技术联盟)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Florida(佛罗里达大学)
;
Northeastern University(东北大学)
;
Singapore Management University(新加坡管理学院)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
MCERF:通过增强检索推进工程文档的多模态大语言模型评估
Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu
机构
*
School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut, Storrs, CT 06269(机械、航空航天与制造工程学院,康涅狄格大学,斯托尔斯,CT 06269)
;
Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA(机械工程系,麻省理工学院,剑桥,MA 02139,美国)
Commentsv2: Minor numerical corrections for Table V. 16 pages, 14 figures, 7 tables. Extended version of paper accepted to 2026 IEEE Intelligent Vehicles Symposium (IV 2026). ScenicRules benchmark available at https://github.com/BerkeleyLearnVerify/ScenicRules