机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
;
The 63rd Research Institute, National University of Defense Technology, Nanjing(国防科技大学第六三研究所,南京)
Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving
超越自我博弈与规模:面向自动驾驶泛化的行为基准
Aron Distelzweig, Faris Janjoš, Andreas Look, Anna Rothenhäusler, Daniel Jost, Oliver Scheel, Raghu Rajan, Daphne Cornelisse, Eugene Vinitsky, Joschka Boedecker
机构
*
University of Freiburg(弗赖堡大学)
;
Bosch Center for Artificial Intelligence(博世人工智能中心)
;
Coburg University of Applied Sciences(科堡应用科学大学)
;
New York University(纽约大学)
Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei, Junwen Miao, Huaisong Zhang, Songcheng Cai, Yubo Wang, Dongfu Jiang, Yuyu Zhang, Ping Nie, Wenhu Chen, Changqian Yu, Kelsey R. Allen
机构
*
University of British Columbia(不列颠哥伦比亚大学)
;
Vector Institute(向量研究所)
;
Kolors Team, Kuaishou Technology(快手团队)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Waterloo(滑铁卢大学)
;
Etude AI
;
Tsinghua University(清华大学)
;
Georgia Institute of Technology(佐治亚理工学院)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
AgentCollabBench: 评估好代理为何会成为差合作者
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia, Tanzila Khan, Kainat Raisa Hossain, Nehaa Shri, Shubhrangshu Debsarkar, Humayra Tasnim, Gour Gupal Talukder Shawon, Debjoty Mitra, Sumaiya Ahmed Rani, Al Jami Islam Anik, Al Nafeu Khan
机构
*
University of Utah(犹他大学)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
University of Dhaka(达卡大学)
;
Vellore Institute of Technology(韦洛雷理工学院)
;
University of Virginia(弗吉尼亚大学)
;
Rajshahi University of Engineering and Technology(拉贾加赫尔工程与技术大学)
;
Shahjalal University of Science and Technology(沙赫jalal科学与技术大学)
;
BRAC University(BRAC大学)
;
Islamic University of Technology(伊斯兰技术大学)
;
Comilla University(科摩拉大学)
机构
*
Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳瓦分校计算机科学与工程系)
;
School of Computer Engineering, KIIT Deemed to be University, Bhubaneswar, India(比哈尔邦布尔萨大学计算机工程学院)
Speech-based Psychological Crisis Assessment using LLMs
基于语音的心理危机评估使用大语言模型
Terumi Chiba, Yang Luo, Ziyun Cui, Yongsheng Tong, Chao Zhang
机构
*
Tsinghua University(清华大学)
;
Peking University Huilongguan Clinical Medical School(北京大学回龙guan临床医学院)
;
WHO Collaborating Centre for Research and Training in Suicide Prevention(世界卫生组织自杀预防研究与培训协作中心)
NaiAD: Initiate Data-Driven Research for LLM Advertising
NaiAD:启动基于数据的研究以进行大语言模型广告
Yihang Zhang, Zimeng Huang, Ren Zhai, Yipeng Kang, Tonghan Wang
机构
*
Tsinghua University(清华大学)
;
College of AI(人工智能学院)
;
Department of Literature, Arts and Communication(文学、艺术与传播系)
;
Anhui International Studies University(安徽国际关系大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
Towards Conversational Medical AI with Eyes, Ears and a Voice
面向有眼睛、耳朵和声音的对话式医疗AI
Meet Shah, Jason Gusdorf, Anil Palepu, Chunjong Park, Jack W. O'Sullivan, Vishnu Ravi, Tim Strother, Pavel Dubov, Aliya Rysbek, Toshiyuki Fukuzawa, Yana Lunts, Jan Freyberg, Michael B. Chang, Aniruddh Raghu, David Stutz, Devora Berlowitz, Eliseo Papa, Taylan Cemgil, JD Velasquez, Jack Chen, Arthur Chen, Doug Fritz, Charlie Taylor, Katya Tregubova, Jing Rong Lim, Richard Green, Sara Mahdavi, Mahvish Nagda, Jihyeon Lee, Craig Schiff, Liviu Panait, Sukhdeep Singh, Valentin Liévin, David G. T. Barrett, Hannah Gladman, Anna Cupani, Francesca Pietra, Uchechi Okereke, Katherine Tong, Clemens Meyer, Erwan Rolland, Mili Sanwalka, Michael D. Howell, Shixiang Shane Gu, Bibo Xu, Euan A. Ashley, S. M. Ali Eslami, Gregory Wayne, Pushmeet Kohli, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam, Ryutaro Tanno
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究)
;
Beth Israel Deaconess Medical Center, Harvard Medical School(贝塞斯达医院, 哈佛医学院)
;
Stanford University(斯坦福大学)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
生成无泄漏的基准以评估鲁棒的RAG
Jiayi Liu, Jiaxing Zhang, Bowen Jin, Jennifer Neville
机构
*
Department of Computer Science, Purdue University(普渡大学计算机科学系)
;
New Jersey Institute of Technology(新泽西理工学院)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Microsoft Research(微软研究院)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
基准是否低估了大语言模型的性能?通过大语言模型优先的人类仲裁评估来评估幻觉检测
I. F. Atasoy, B. Mutlu, E. A. Sezer, A. Wahdan
机构
*
Department of Computer Engineering, Hacettepe University(哈切塔佩大学计算机工程系)
;
Department of Computer Engineering, Ankara University(安卡拉大学计算机工程系)
;
Zephlen AI and Information Technologies Inc.(泽夫伦人工智能与信息技术公司)
Commentshis is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: https://doi.org/10.1109/SUMMA68668.2025.11302248
Journal ref2025 7th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA), Lipetsk, Russian Federation, 2025, pp. 799-804