Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
你的AI旅行代理会为你预订斗牛:前沿AI模型中隐含动物福利的代理基准
Jasmine Brazilek, Joel Christoph, Maheep Chaudhary, Oliver Tullio, Carol Kline, Miles Tidmarsh, Arturs Kanepajs
机构
*
Compassion Aligned Machine Learning(同情对齐机器学习)
;
Sentient Futures(感知未来)
;
Harvard Kennedy School(哈佛肯尼迪学院)
;
Appalachian State University Department of Management(阿巴拉契亚州立大学管理系)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
一种评估智能体人工智能自主模型发现的实验设计方法
Hao He, Xueying Liu, Chris J. Kuhlman, Xinwei Deng
机构
*
Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工学院)
;
Department of Statistical Science, Baylor University(统计科学系,贝勒大学)
;
Advanced Research Computing, Virginia Tech(高级研究计算,弗吉尼亚理工学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Pluralis v0.1:迈向用于人工智能风险与可靠性的多元文化、多模态、多语言基准测试
Alicia Parrish, Rajat Shinde, Sanket Badhe, Xinyi Bai, Sree Bhargavi Balija, Hua-Rong Chu, Emilio Ferrara, Armstrong Foundjem, Rajat Ghosh, Aakash Gupta, Xuanli He, Ong Chen Hui, Minji Jung, Madhangi Karimanal, Faiza Khan Khattak, Boryoung Kim, Eugenia Kim, Liliya Lavitas, Seok Min Lim, Victor Lu, Jim Moirangthem, Dhivya Nagasubramanian, Deepak Pandita, Sita Rajagopal, Geetha Raju, Evgeniia Razumovskaia, Aravind Reddy, Federico Ricciuti, Nobin Sarwar, Sungpil Shin, Sunayana Sitaram, Snehal Thorat, Tharindu Cyril Weerasooriya, Jasmijn Bastings, Joachim Baumann, Kongtao Chen, Murali Emani, Mariya Hendriksen, Jiho Jin, Jun Seong Kim, Younghoon Ko, Alicja Kwasniewska, Minjae Lee, Tom Wei-cyuan Lin Kashyap Ramanandula Manjusha, Junho Myung, Junyeong Park, Roma Patel, Shyam Ratan, Sudarsun Santhiappan, Priyanka Suresh, Tuesday, Ksheeraj Sai Vepuri Laura Amortegui-Ordonez, Claire Dennis, Minsuk Kahng, Chris Knotz, Alice Oh, Balaraman Ravindran, Soojung Ryu William Bartholomew, Hiwot Tesfaye, Lora Aroyo
机构
*
Google DeepMind(谷歌DeepMind)
;
University of Alabama in Huntsville(阿拉巴马大学亨茨维尔分校)
;
Google(谷歌)
;
University of Missouri Columbia(密苏里大学哥伦比亚分校)
;
Chunghwa Telecom Laboratories(春木电信实验室)
;
University of Southern California(南加州大学)
;
Polytechnique Montreal(蒙特利尔理工学院)
;
Nutanix
;
ThinkEvolve Labs(ThinkEvolve实验室)
;
UCL(伦敦大学学院)
;
Infocomm Media Development Authority(信息通信媒体发展局)
;
Monark Health(Monark健康)
;
Seoul National University(首尔国立大学)
;
Microsoft(微软)
;
Centre for Responsible AI (CeRAI), Wadhwani School of Data Science and AI (WSAI), Indian Institute of Technology Madras(负责任人工智能中心(CeRAI)、瓦达威人工智能学校(WSAI)、印度理工学院马德拉斯分校)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
Microsoft Research India(微软印度研究院)
;
Stanford University(斯坦福大学)
;
Argonne National Laboratory(阿贡国家实验室)
;
University of Oxford(牛津大学)
;
KAIST(韩国科学技术院)
;
Yonsei University(延世大学)
;
Amazon(亚马逊)
;
UIUC(伊利诺伊大学香槟分校)
;
Rochester Institute of Technology(罗切斯特理工学院)
;
Xenoscube Inc.(Xenoscube公司)
;
Korea AI Safety Institute (K-AISI)(韩国人工智能安全研究所(K-AISI))
;
MLCommons
;
CommonGround
;
Artifex Labs(Artifex实验室)
A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid
通过患者模拟推进AI可信度:对抗抑郁药物选择的对话代理风险评估
Md Tanvir Rouf Shawon, Mohammad Sabik Irbaz, Hadeel R. A. Elyazori, Keerti Reddy Resapu, Yili Lin, Vladimir Franzuela Cardenas, K. Pierre Eklou, Farrokh Alemi, Kevin Lybarger
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
解码多模态思维:通过多模态对齐和自适应路由实现可泛化的脑到文本翻译
Chunyu Ye, Yunhao Zhang, Jingyuan Sun, Chong Li, Yang Zhao, Shaonan Wang
机构
*
State Key Laboratory of Multimodal Artificial Intelligence System, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系)
;
Department of Language Science and Technology, Hong Kong Polytechnic University(香港理工大学语言科学与技术系)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
StepShield: When, Not Whether to Intervene on Rogue Agents
StepShield:何时干预流氓代理,而非是否干预
Gloria Felicia, Zitha Sasindran, Jinfeng He, Michael Eniolade, Hemant Kumar, Milan Hussain Angati
机构
*
University of Virginia(弗吉尼亚大学)
;
Indian Institute of Science, Bangalore(印度科学研究院,班加罗尔)
;
Cornell University(康奈尔大学)
;
University of the Cumberlands(库姆伯兰兹大学)
;
University of Arizona(亚利桑那大学)
;
California State University, Northridge(加州州立大学,北岭分校)
Comments7 pages. To be published in the proceedings of 41st International Conference on Automated Software Engineering (ASE '26), October 12-16, 2026, Munich, Germany (Industry Showcase Track)
Quantifying Retriever-Generator Alignment in RAG with Local Explanations
用局部解释量化检索生成模型(RAG)中的检索器-生成器对齐
Korbinian Randl, Guido Rocchietti, Aron Henriksson, Ziawasch Abedjan, Tony Lindgren, John Pavlopoulos
机构
*
Department of Computer and Systems Sciences, Stockholm University(斯德哥尔摩大学计算机与系统科学系)
;
BIFOLD, Technische Universität Berlin(柏林技术大学BIFOLD)
;
Department of Informatics, Athens University of Economics and Business(雅典经济与商业大学信息系)
;
Archimedes, Athena Research Centre(雅典研究中心Archimedes)
机构
*
China University of Petroleum-Beijing at Karamay(中国石油大学(北京)克拉玛依校区)
;
Guizhou University(贵州大学)
;
Shenzhen Research Institute of Big Data(深圳大数据研究院)
;
Shanghai Jiao Tong University(上海交通大学)
Comments21 pages, 7 figures, 5 tables. Preprint. An earlier version of this work was presented at the 105th Annual Meeting of the Transportation Research Board (TRB), January 2026
机构
*
Mininglamp Technology(旷视科技)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所综合技术集成中心)
Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective
从车道感知角度评估自动驾驶对环境错觉的鲁棒性:基准测试
Tianyuan Zhang, Xianglong Liu, Aishan Liu, Lu Wang, Yitong Zhang, Peng Yue, Mingchuan Zhang, Siyuan Liang, Dacheng Tao
机构
*
SKLCCSE, the School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院软件安全技术与工程北京市重点实验室)
;
the School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络空间科学与技术学院)
;
Henan University of Science and Technology(河南科技大学)
;
the School of Computing, National University of Singapore(新加坡国立大学计算学院)
;
College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)